AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

ApplicationsDGX agent

arXiv:2602.06575v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding

TimeThink: Reasoning with Time for Video LLMs

ResearchDGX agent

arXiv:2607.05089v1 Announce Type: new Abstract: Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Vi

TiROD: Tiny Robotics Dataset and Benchmark for Continual Object Detection

Model ReleasesDGX agent

arXiv:2409.16215v4 Announce Type: replace-cross Abstract: Detecting objects with visual sensors is crucial for numerous mobile robotics applications, from autonomous navigation to inspection. However,


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications

TutorialsDGX agent

arXiv:2502.12096v5 Announce Type: replace-cross Abstract: In this paper, we introduce token communications (TokCom), a large model-driven framework to leverage cross-modal context information in gener

Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

Model ReleasesDGX agent

arXiv:2607.04199v1 Announce Type: new Abstract: The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

Model ReleasesDGX agent

arXiv:2605.25706v2 Announce Type: replace Abstract: Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in v

Towards Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion

Model ReleasesDGX agent

arXiv:2601.15829v2 Announce Type: replace Abstract: Recent years have witnessed the remarkable success of deep learning in remote sensing image interpretation, driven by the availability of large-scal

Towards Standardized Light Field Quality Assessment: Hybrid Subjective Benchmarking and Objective Metric Evaluation

Model ReleasesDGX agent

arXiv:2607.03494v1 Announce Type: new Abstract: Benchmarking immersive media coding solutions, especially in the standardization context, requires reliable and reproducible subjective quality assessme

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

Local AiDGX agent

arXiv:2607.02798v1 Announce Type: new Abstract: Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditi

Trajectory-Anchor Optimization for Overconfident Thermal Visual Place Recognition: Zero-Leakage OOD Auditing and Kidnapped-Robot Recovery

SafetyDGX agent

arXiv:2607.04745v1 Announce Type: cross Abstract: Modern thermal visual place recognition (TIR-VPR) frontends based on foundation models achieve remarkable closed-set retrieval but suffer from an over

Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography

Model ReleasesDGX agent

arXiv:2607.03740v1 Announce Type: cross Abstract: Microbial keratitis requires rapid pathogen identification to guide treatment, but culture- and PCR-based diagnostics are slow and resource-intensive.

TRISTAR: Triple-Signal Stair Recognition and Vision-Only Indoor Navigation for Search-and-Rescue Micro-UAVs

AgentsDGX agent

arXiv:2607.03818v1 Announce Type: new Abstract: Indoor search-and-rescue (SAR) operations often require rapid situational awareness where GNSS signals are unavailable and human access is difficult or

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

ResearchDGX agent

arXiv:2607.04484v1 Announce Type: new Abstract: Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal rea

TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection

ResearchDGX agent

arXiv:2607.04684v1 Announce Type: new Abstract: Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge i

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

SafetyDGX agent

arXiv:2607.02569v1 Announce Type: new Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

Model ReleasesDGX agent

arXiv:2603.24690v2 Announce Type: replace Abstract: In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to exampl

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

Local AiDGX agent

arXiv:2607.04498v1 Announce Type: new Abstract: With the proliferation of AI-generated content, sophisticated multimedia manipulation has raised critical concerns about malicious applications such as

UniSpine-GS: An Efficient Physics-Aware Gaussian Framework for Cross-Modality Multi-view Spine Image Synthesis

ResearchDGX agent

arXiv:2607.04923v1 Announce Type: new Abstract: The diagnosis of spinal diseases is often assisted by 3D imaging techniques in clinical practice. However, precise 3D spinal assessment is limited by th

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

AgentsDGX agent

arXiv:2607.05133v1 Announce Type: new Abstract: World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as den

UniVideo: Unified Understanding, Generation, and Editing for Videos

Model ReleasesDGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

Unsupervised Detection of Underground Tunnels in Ground-Penetrating Radar Using Depth-Restricted Reconstruction Scoring

ResearchDGX agent

arXiv:2607.04882v1 Announce Type: new Abstract: Clandestine tunneling beneath oil and gas pipelines enables fuel theft, smuggling, and sabotage, yet conventional monitoring detects damage only after a

Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images

ResearchDGX agent

arXiv:2607.05006v1 Announce Type: new Abstract: While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images re

USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning

ResearchDGX agent

arXiv:2607.03900v1 Announce Type: new Abstract: Test-time adaptation (TTA) has emerged as a popular paradigm for improving the performance of vision-language models (e.g., CLIP) on downstream tasks. A

Utonia: Toward One Encoder for All Point Clouds

AgentsDGX agent

arXiv:2603.03283v2 Announce Type: replace Abstract: We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we pres

VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation

TutorialsDGX agent

arXiv:2603.18797v2 Announce Type: replace Abstract: Spatial graphs provide a lightweight and elegant representation of curvilinear anatomical structures such as blood vessels, lung airways, and neuron

Video Generation Models Are Inherent Lighting Estimators

ResearchDGX agent

arXiv:2607.04674v1 Announce Type: new Abstract: Recovering dynamic environment maps from a single in-the-wild video is crucial for photorealistic rendering, yet remains a challenge. Recent video gener

Vidu S1: A Real-Time Interactive Video Generation Model

ResearchDGX agent

arXiv:2607.03118v1 Announce Type: new Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation

Virtual Category-Guided Continual Generalized Category Discovery

SafetyDGX agent

arXiv:2607.04984v1 Announce Type: new Abstract: Continual Generalized Category Discovery (C-GCD) aims to incrementally identify novel categories from sequential unlabeled data while preserving recogni

Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics

SafetyDGX agent

arXiv:2607.03589v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal seq

Vision Pretraining for Dense Spatial Perception

ResearchDGX agent

arXiv:2607.05247v1 Announce Type: new Abstract: Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable represe

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2511.20272v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

SafetyDGX agent

arXiv:2607.05180v1 Announce Type: cross Abstract: Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the ve

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Model ReleasesDGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

ResearchDGX agent

arXiv:2607.02707v1 Announce Type: new Abstract: Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provid

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

ApplicationsDGX agent

arXiv:2606.14048v2 Announce Type: replace Abstract: World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

Model ReleasesDGX agent

arXiv:2607.04731v1 Announce Type: new Abstract: Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising pr

When Does Resolution Help a Frozen Backbone? Global Attention at Resolution Predicts Scalable Adaptation for Camouflaged and Marine Animal Segmentation

ResearchDGX agent

arXiv:2607.02708v1 Announce Type: new Abstract: Adapting frozen vision foundation models to fine-grained segmentation now largely depends on backbone selection. Whether the backbone applies global att

When Geometry Aligns: Dihedral Hidden-State Transformations in UNet, ViT, and DiT Architectures

ResearchDGX agent

arXiv:2607.03580v1 Announce Type: cross Abstract: Diffusion architectures now encompass convolutional UNets as well as transformer-based designs such as Diffusion Transformers (DiTs), inspired by Visi

WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images

ResearchDGX agent

arXiv:2607.05347v1 Announce Type: new Abstract: While feedforward 3D reconstruction excels at efficient novel view synthesis, it typically falters when faced with scenes under varying illumination. To

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Model ReleasesDGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion

ResearchDGX agent

arXiv:2603.22972v3 Announce Type: replace Abstract: Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -vi

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

Model ReleasesDGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

2 Jul 2026

3D Point World Models: Point Completion Enables More Accurate Dynamics Learning

ResearchDGX agent

arXiv:2607.00148v1 Announce Type: cross Abstract: Learning predictive models of the world enables robotic control through planning, potentially allowing robots to improvise solutions on new tasks. How

A Synthetic-Driven Vision System for Assembly Step Recognition

ApplicationsDGX agent

arXiv:2607.00129v1 Announce Type: new Abstract: Quality control in industrial assembly is essential, and real-time monitoring of the assembly process is crucial for preventing costly defects and ensur

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Local AiDGX agent

arXiv:2607.00678v1 Announce Type: new Abstract: Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typi

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers

SafetyDGX agent

arXiv:2607.00580v1 Announce Type: new Abstract: Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial

Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction

ResearchDGX agent

arXiv:2603.06852v2 Announce Type: replace Abstract: Sparse-view computed tomography (CT) is critical for reducing radiation exposure to patients. Recent advances in radiative 3D Gaussian Splatting (3D

AEGIS: A Multi-Task Joint-Embedding Predictive Architecture for Mammography

ResearchDGX agent

arXiv:2607.00277v1 Announce Type: new Abstract: We present Aegis, a joint-embedding predictive architecture for breast cancer detection and density assessment in mammography. We train three Vision Tra

AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware

Model ReleasesDGX agent

arXiv:2602.16249v2 Announce Type: replace Abstract: Self-supervised pretraining has transformed computer vision by enabling data-efficient fine-tuning, yet high-resolution pretraining typically requir

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

ResearchDGX agent

arXiv:2506.12009v2 Announce Type: replace Abstract: Affordance grounding aims to localize where to interact with an object, a fundamental capability for embodied agents. Yet progress is bottlenecked b

AnF-DiffPET: Anatomy- and Frequency-Guided Diffusion for PET/CT Denoising

Model ReleasesDGX agent

arXiv:2607.00509v1 Announce Type: new Abstract: Positron emission tomography (PET) provides essential functional information for disease assessment, however reducing injected activity or acquisition t

Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval

SafetyDGX agent

arXiv:2607.00379v1 Announce Type: cross Abstract: Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual se

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Model ReleasesDGX agent

arXiv:2607.00726v1 Announce Type: new Abstract: Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for

AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution

ResearchDGX agent

arXiv:2607.00987v1 Announce Type: new Abstract: Diffusion models have significantly advanced video super-resolution (VSR) but remain largely constrained to fixed upsampling scales. Conversely, while c

Beyond Pixel Overlap: A Framework for Decomposing Segmentation Evaluation Metrics

ResearchDGX agent

arXiv:2607.00886v1 Announce Type: new Abstract: Evaluation metrics are central to binary target segmentation because they determine how progress is measured, compared, and interpreted. In this paper,

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

SafetyDGX agent

arXiv:2607.00573v1 Announce Type: new Abstract: Diffusion MRI probes brain microstructure with particular sensitivity to early cerebrovascular and neurodegenerative changes. Neurite Orientation Disper

Caption Bottleneck Models

SafetyDGX agent

arXiv:2607.00578v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an

ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction

SafetyDGX agent

arXiv:2607.00798v1 Announce Type: new Abstract: Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment

Closed-loop coupling of personalised and foundation models for real-time treatment guidance with MRI

ResearchDGX agent

arXiv:2607.00500v1 Announce Type: cross Abstract: Image-guided therapies, including radiotherapy, biopsy and deep brain stimulation, rely on real-time targeting of anatomical structures. However, in t

Condensing Large-Scale Datasets Directly with Minimal Information Loss

HardwareDGX agent

arXiv:2607.00916v1 Announce Type: new Abstract: Recent advancements in scaling dataset distillation rely heavily on decoupled information extraction pipelines, comprising SQUEEZE, RECOVER, and RELABEL

← Previous
1…5354555657…209
Next →