AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

UniSpine-GS: An Efficient Physics-Aware Gaussian Framework for Cross-Modality Multi-view Spine Image Synthesis

DGX agent

arXiv:2607.04923v1 Announce Type: new Abstract: The diagnosis of spinal diseases is often assisted by 3D imaging techniques in clinical practice. However, precise 3D spinal assessment is limited by th

researcharxiv-cs-cv
7 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

DGX agent

arXiv:2607.05133v1 Announce Type: new Abstract: World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as den

agentsarxiv-cs-cv
7 Jul 2026
Model Releases

UniVideo: Unified Understanding, Generation, and Editing for Videos

DGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Unsupervised Detection of Underground Tunnels in Ground-Penetrating Radar Using Depth-Restricted Reconstruction Scoring

DGX agent

arXiv:2607.04882v1 Announce Type: new Abstract: Clandestine tunneling beneath oil and gas pipelines enables fuel theft, smuggling, and sabotage, yet conventional monitoring detects damage only after a

researcharxiv-cs-cv
7 Jul 2026
Research

Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images

DGX agent

arXiv:2607.05006v1 Announce Type: new Abstract: While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images re

researcharxiv-cs-cv
7 Jul 2026
Research

USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning

DGX agent

arXiv:2607.03900v1 Announce Type: new Abstract: Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.g., CLIP) on downstream tasks. A

researcharxiv-cs-cv
7 Jul 2026
Agents

Utonia: Toward One Encoder for All Point Clouds

DGX agent

arXiv:2603.03283v2 Announce Type: replace Abstract: We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we pres

agentsarxiv-cs-cv
7 Jul 2026
Tutorials

VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation

DGX agent

arXiv:2603.18797v2 Announce Type: replace Abstract: Spatial graphs provide a lightweight and elegant representation of curvilinear anatomical structures such as blood vessels, lung airways, and neuron

tutorialsarxiv-cs-cv
7 Jul 2026
Research

Video Generation Models Are Inherent Lighting Estimators

DGX agent

arXiv:2607.04674v1 Announce Type: new Abstract: Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video gener

researcharxiv-cs-cv
7 Jul 2026
Research

Vidu S1: A Real-Time Interactive Video Generation Model

DGX agent

arXiv:2607.03118v1 Announce Type: new Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation

researcharxiv-cs-cv
7 Jul 2026
Safety

Virtual Category-Guided Continual Generalized Category Discovery

DGX agent

arXiv:2607.04984v1 Announce Type: new Abstract: Continual Generalized Category Discovery (C-GCD) aims to incrementally identify novel categories from sequential unlabeled data while preserving recogni

safetyarxiv-cs-cv
7 Jul 2026
Safety

Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics

DGX agent

arXiv:2607.03589v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal seq

safetyarxiv-cs-cv
7 Jul 2026
Research

Vision Pretraining for Dense Spatial Perception

DGX agent

arXiv:2607.05247v1 Announce Type: new Abstract: Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable represe

researcharxiv-cs-cv
7 Jul 2026
Model Releases

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

DGX agent

arXiv:2511.20272v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

DGX agent

arXiv:2607.05180v1 Announce Type: cross Abstract: Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the ve

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

DGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

model-releasesarxiv-cs-cv
7 Jul 2026
Research

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

DGX agent

arXiv:2607.02707v1 Announce Type: new Abstract: Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provid

researcharxiv-cs-cv
7 Jul 2026
Applications

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

DGX agent

arXiv:2606.14048v2 Announce Type: replace Abstract: World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing

applicationsarxiv-cs-cv
7 Jul 2026
Model Releases

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

DGX agent

arXiv:2607.04731v1 Announce Type: new Abstract: Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising pr

model-releasesarxiv-cs-cv
7 Jul 2026
Research

When Does Resolution Help a Frozen Backbone? Global Attention at Resolution Predicts Scalable Adaptation for Camouflaged and Marine Animal Segmentation

DGX agent

arXiv:2607.02708v1 Announce Type: new Abstract: Adapting frozen vision foundation models to fine-grained segmentation now largely depends on backbone selection. Whether the backbone applies global att

researcharxiv-cs-cv
7 Jul 2026
Research

When Geometry Aligns: Dihedral Hidden-State Transformations in UNet, ViT, and DiT Architectures

DGX agent

arXiv:2607.03580v1 Announce Type: cross Abstract: Diffusion architectures now encompass convolutional UNets as well as transformer-based designs such as Diffusion Transformers (DiTs), inspired by Visi

researcharxiv-cs-cv
7 Jul 2026
Research

WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images

DGX agent

arXiv:2607.05347v1 Announce Type: new Abstract: While feedforward 3D reconstruction excels at efficient novel view synthesis, it typically falters when faced with scenes under varying illumination. To

researcharxiv-cs-cv
7 Jul 2026
Model Releases

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

DGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

model-releasesarxiv-cs-cv
7 Jul 2026
Research

WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion

DGX agent

arXiv:2603.22972v3 Announce Type: replace Abstract: Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -vi

researcharxiv-cs-cv
7 Jul 2026
Model Releases

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

DGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

model-releasesarxiv-cs-cv
7 Jul 2026
Research

3D Point World Models: Point Completion Enables More Accurate Dynamics Learning

DGX agent

arXiv:2607.00148v1 Announce Type: cross Abstract: Learning predictive models of the world enables robotic control through planning, potentially allowing robots to improvise solutions on new tasks. How

researcharxiv-cs-cv
2 Jul 2026
Applications

A Synthetic-Driven Vision System for Assembly Step Recognition

DGX agent

arXiv:2607.00129v1 Announce Type: new Abstract: Quality control in industrial assembly is essential, and real-time monitoring of the assembly process is crucial for preventing costly defects and ensur

applicationsarxiv-cs-cv
2 Jul 2026
Local Ai

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

DGX agent

arXiv:2607.00678v1 Announce Type: new Abstract: Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typi

local-aiarxiv-cs-cv
2 Jul 2026
Safety

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers

DGX agent

arXiv:2607.00580v1 Announce Type: new Abstract: Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial

safetyarxiv-cs-cv
2 Jul 2026
Research

Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction

DGX agent

arXiv:2603.06852v2 Announce Type: replace Abstract: Sparse-view computed tomography (CT) is critical for reducing radiation exposure to patients. Recent advances in radiative 3D Gaussian Splatting (3D

researcharxiv-cs-cv
2 Jul 2026
Research

AEGIS: A Multi-Task Joint-Embedding Predictive Architecture for Mammography

DGX agent

arXiv:2607.00277v1 Announce Type: new Abstract: We present Aegis, a joint-embedding predictive architecture for breast cancer detection and density assessment in mammography. We train three Vision Tra

researcharxiv-cs-cv
2 Jul 2026
Model Releases

AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware

DGX agent

arXiv:2602.16249v2 Announce Type: replace Abstract: Self-supervised pretraining has transformed computer vision by enabling data-efficient fine-tuning, yet high-resolution pretraining typically requir

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

DGX agent

arXiv:2506.12009v2 Announce Type: replace Abstract: Affordance grounding aims to localize where to interact with an object, a fundamental capability for embodied agents. Yet progress is bottlenecked b

researcharxiv-cs-cv
2 Jul 2026
Model Releases

AnF-DiffPET: Anatomy- and Frequency-Guided Diffusion for PET/CT Denoising

DGX agent

arXiv:2607.00509v1 Announce Type: new Abstract: Positron emission tomography (PET) provides essential functional information for disease assessment, however reducing injected activity or acquisition t

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval

DGX agent

arXiv:2607.00379v1 Announce Type: cross Abstract: Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual se

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

DGX agent

arXiv:2607.00726v1 Announce Type: new Abstract: Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for

model-releasesarxiv-cs-cv
2 Jul 2026
Research

AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution

DGX agent

arXiv:2607.00987v1 Announce Type: new Abstract: Diffusion models have significantly advanced video super-resolution (VSR) but remain largely constrained to fixed upsampling scales. Conversely, while c

researcharxiv-cs-cv
2 Jul 2026
Research

Beyond Pixel Overlap: A Framework for Decomposing Segmentation Evaluation Metrics

DGX agent

arXiv:2607.00886v1 Announce Type: new Abstract: Evaluation metrics are central to binary target segmentation because they determine how progress is measured, compared, and interpreted. In this paper,

researcharxiv-cs-cv
2 Jul 2026
Safety

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

DGX agent

arXiv:2607.00573v1 Announce Type: new Abstract: Diffusion MRI probes brain microstructure with particular sensitivity to early cerebrovascular and neurodegenerative changes. Neurite Orientation Disper

safetyarxiv-cs-cv
2 Jul 2026
Safety

Caption Bottleneck Models

DGX agent

arXiv:2607.00578v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an

safetyarxiv-cs-cv
2 Jul 2026
Safety

ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction

DGX agent

arXiv:2607.00798v1 Announce Type: new Abstract: Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment

safetyarxiv-cs-cv
2 Jul 2026
Research

Closed-loop coupling of personalised and foundation models for real-time treatment guidance with MRI

DGX agent

arXiv:2607.00500v1 Announce Type: cross Abstract: Image-guided therapies, including radiotherapy, biopsy and deep brain stimulation, rely on real-time targeting of anatomical structures. However, in t

researcharxiv-cs-cv
2 Jul 2026
Hardware

Condensing Large-Scale Datasets Directly with Minimal Information Loss

DGX agent

arXiv:2607.00916v1 Announce Type: new Abstract: Recent advancements in scaling dataset distillation rely heavily on decoupled information extraction pipelines, comprising SQUEEZE, RECOVER, and RELABEL

hardwarearxiv-cs-cv
2 Jul 2026
Safety

Continuous Speculative Decoding for Autoregressive Image Generation

DGX agent

arXiv:2411.11925v3 Announce Type: replace Abstract: Continuous visual autoregressive (AR) models have demonstrated promising performance in image generation, but their inherently sequential nature res

safetyarxiv-cs-cv
2 Jul 2026
Tutorials

CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

DGX agent

arXiv:2607.00321v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D models of highly articulated animals, such as dogs, from a single in-the-wild image remains a formidable challenge. In t

tutorialsarxiv-cs-cv
2 Jul 2026
Model Releases

CPDDNet: Color-Polarization Denoising and Demosaicking Network

DGX agent

arXiv:2607.01100v1 Announce Type: new Abstract: Color-polarization imaging using a color-polarization filter array (CPFA) sensor captures both texture (color intensity) and physical (polarization) inf

model-releasesarxiv-cs-cv
2 Jul 2026
Research

DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

DGX agent

arXiv:2607.00672v1 Announce Type: new Abstract: Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing met

researcharxiv-cs-cv
2 Jul 2026
Safety

Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection

DGX agent

arXiv:2607.00948v1 Announce Type: new Abstract: The visual quality of AI-generated videos has improved drastically in recent years, making it increasingly difficult for humans to distinguish between r

safetyarxiv-cs-cv
2 Jul 2026
← Previous
1…6970717273…263
Next →