AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
2 Jun 2026

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

ResearchDGX agent

arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from s

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models

AgentsDGX agent

arXiv:2606.02580v1 Announce Type: new Abstract: Inverse graphics is a longstanding and highly underconstrained problem that seeks to reconstruct images as editable 3D scenes which can be rendered, rel

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.00592v1 Announce Type: new Abstract: Effective visual communication stems from the harmony of multiple design principles, such as readability, contrast, alignment, overlap, and coherence, w

TIDES: Time-Derivative Event Simulation via Deformable Reconstruction

TutorialsDGX agent

arXiv:2606.02058v1 Announce Type: new Abstract: Event cameras emit asynchronous events in response to environmental appearance changes. The scarcity of real-world event datasets makes simulation essen

TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning

Model ReleasesDGX agent

arXiv:2606.01591v1 Announce Type: new Abstract: The TimeLogic Challenge evaluates formal temporal-logic reasoning over video - 16 operators (before, after, until, since, always, co-occur, ordering, ..

ToolFG: Towards Well-Grounded Fine-Grained Image Classification

SafetyDGX agent

arXiv:2606.02518v1 Announce Type: new Abstract: Fine-grained image classification (FGIC) has broad applications and has attracted significant research attention. In this paper, we explore a novel para

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

Model ReleasesDGX agent

arXiv:2509.16635v2 Announce Type: replace Abstract: In real applications, person re-identification (ReID) is expected to retrieve the target person at any time, including both daytime and nighttime, r

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends

AgentsDGX agent

arXiv:2606.01164v1 Announce Type: new Abstract: With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, bene

Towards Sparse Video Understanding and Reasoning

AgentsDGX agent

arXiv:2602.13602v2 Announce Type: replace Abstract: We present revise (nderline{Re}asoning with nderline{Vi}deo nderline{S}parsity), a multi-round agent for video question answering (VQA). Instead of

Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

ResearchDGX agent

arXiv:2606.02321v1 Announce Type: new Abstract: Recent advances in large vision-language models have expanded video retrieval from simple text-based search to more flexible scenarios, where users may

Training-Free Continuous Bitrate Control for Scalable Image Coding for Humans and Machines

ApplicationsDGX agent

arXiv:2606.00158v1 Announce Type: cross Abstract: Continuous variable-rate compression is highly demanded in real-world applications, but remains underexplored in scalable image coding for humans and

Training-Free Coverless Multi-Image Steganography with Access Control

ResearchDGX agent

arXiv:2603.09390v2 Announce Type: replace Abstract: Coverless Image Steganography (CIS) hides information without explicitly modifying a cover image, providing strong imperceptibility and inherent rob

Training-free image inversion for one-step diffusion models

SafetyDGX agent

arXiv:2606.01380v1 Announce Type: new Abstract: In this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models,addressing key challenges in real image inver

Training-Free Object-Agnostic Jam Detection in Fulfillment Centers

ResearchDGX agent

arXiv:2606.00321v1 Announce Type: new Abstract: In fulfillment centers, diverse objects move continuously from inbound to outbound operations and can become jammed due to excessive conveyor friction,

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

SafetyDGX agent

arXiv:2606.02350v1 Announce Type: new Abstract: Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior wor

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval

SafetyDGX agent

arXiv:2606.01615v1 Announce Type: new Abstract: Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear

Two Datasets Are Better Than One: Method of Double Moments for 3-D Reconstruction in Cryo-EM

ResearchDGX agent

arXiv:2511.07438v3 Announce Type: replace Abstract: Cryo-electron microscopy (cryo-EM) is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic

Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking From Sparse Inertial Sensors and Ranging-Based Between-Sensor Distances

SafetyDGX agent

arXiv:2606.02153v1 Announce Type: new Abstract: Methods using inertial measurement units (IMUs) provide a wearable alternative to camera-based motion capture. To mitigate drift from inertial signals,

Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning

AgentsDGX agent

arXiv:2606.01935v1 Announce Type: new Abstract: Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tok

Unified Semantic Transformer for 3D Scene Understanding

ApplicationsDGX agent

arXiv:2512.14364v3 Announce Type: replace Abstract: Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, exist

UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization

ResearchDGX agent

arXiv:2606.00351v1 Announce Type: new Abstract: Personalized visual understanding has advanced significantly, yet existing approaches struggle to localize and extract specific concepts when input imag

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

AgentsDGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations

ResearchDGX agent

arXiv:2510.13774v2 Announce Type: replace-cross Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.

VEDAL: Variational Error-Driven Asynchronous Learning for 3D Gaussian Splatting Pruning

ResearchDGX agent

arXiv:2606.02346v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) achieves remarkable novel view synthesis quality with real-time rendering, yet suffers from excessive memory consumption du

VICR: Visual In-Context Restoration for Real-World Image Super-Resolution

Local AiDGX agent

arXiv:2606.00704v1 Announce Type: new Abstract: Real-world image super-resolution (Real-ISR) requires balancing structural fidelity to degraded observations with realistic detail synthesis. However, e

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

AgentsDGX agent

arXiv:2602.04094v2 Announce Type: replace Abstract: Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints an

Visible Light Positioning With Lame Curve LEDs: A Generic Approach for Camera Pose Estimation

ResearchDGX agent

arXiv:2602.01577v3 Announce Type: replace-cross Abstract: Camera-based visible light positioning (VLP) is a promising technique for accurate and low-cost indoor camera pose estimation (CPE). To reduce

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

Model ReleasesDGX agent

arXiv:2606.02273v1 Announce Type: new Abstract: Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on g

VISReg: Variance-Invariance-Sketching Regularization for JEPA training

ResearchDGX agent

arXiv:2606.02572v1 Announce Type: new Abstract: Self-supervised learning methods prevent embedding collapse via modeling heuristics or explicit regularization of the embedding space. Among the latter,

Visualizing definitional divergence in high-dimensional data by manifold alignment: Application to 3D right ventricular strain computations

SafetyDGX agent

arXiv:2501.12178v2 Announce Type: replace Abstract: Medical imaging studies often rely on a single sample per subject, assuming it is representative of their physiological traits. However, variations

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

ResearchDGX agent

arXiv:2606.02564v1 Announce Type: new Abstract: The recent 'Reasoning with Video' paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reaso

WALL-WM: Carving World Action Modeling at the Event Joints

ApplicationsDGX agent

arXiv:2606.01955v1 Announce Type: cross Abstract: WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining

Wavelet-Fusion Diffusion Model for Multimodal Brain MRI Synthesis with Modality and Metadata Conditioning

SafetyDGX agent

arXiv:2606.00689v1 Announce Type: new Abstract: Multimodal MRI provides complementary information for neuroimaging analysis, where different imaging modalities capture distinct anatomical, tissue, and

WebSpline: Structure-Informed Splines for Real-Time 3D Gaussians from Monocular Videos

Local AiDGX agent

arXiv:2606.02096v1 Announce Type: new Abstract: Dynamic scene reconstruction from monocular videos remains highly challenging, as existing methods often struggle to balance global structural coherence

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs

SafetyDGX agent

arXiv:2606.01624v1 Announce Type: new Abstract: Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet veri

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

Model ReleasesDGX agent

arXiv:2606.01247v1 Announce Type: new Abstract: Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has la

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation

ResearchDGX agent

arXiv:2606.00310v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models deliver high-quality image generation but suffer from significant inference latency at high resolutions. Recent accel

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata

ApplicationsDGX agent

arXiv:2602.12819v2 Announce Type: replace-cross Abstract: In this paper, we present WISE, an open-source audiovisual search engine which integrates a range of multimodal retrieval capabilities into a

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

Model ReleasesDGX agent

arXiv:2603.06331v2 Announce Type: replace Abstract: Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactiv

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

Model ReleasesDGX agent

arXiv:2512.10958v2 Announce Type: replace Abstract: Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fa

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

SafetyDGX agent

arXiv:2605.24892v2 Announce Type: replace Abstract: Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and gen

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

Model ReleasesDGX agent

arXiv:2606.02482v1 Announce Type: new Abstract: While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and mul

XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation

AgentsDGX agent

arXiv:2510.13565v2 Announce Type: replace Abstract: Depth estimation remains central to autonomous driving, and radar-camera fusion offers robustness in adverse conditions by providing complementary g

Zero-Shot Multi-Animal Tracking in the Wild

ResearchDGX agent

arXiv:2511.02591v2 Announce Type: replace Abstract: Multi-animal tracking is crucial for understanding animal ecology and behavior, yet remains challenging due to variations in habitat, motion pattern

1 Jun 2026

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

Model ReleasesDGX agent

arXiv:2605.30469v1 Announce Type: cross Abstract: 3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural

A Context-Aware Middleware for Medical Image Based Reports: An approach based on image feature extraction and association rules

ResearchDGX agent

arXiv:2605.30699v1 Announce Type: cross Abstract: This work proposes a context-aware middleware for medical workflow organization and efficiency improvement. In hospitals, laboratories and teleradiolo

A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss

Model ReleasesDGX agent

arXiv:2509.10114v2 Announce Type: replace Abstract: Face image quality assessment (FIQA) plays a critical role in face recognition and verification systems, especially in uncontrolled, real-world envi

A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications

ResearchDGX agent

arXiv:2601.22202v2 Announce Type: replace-cross Abstract: Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw d

A Unifying View of Variational Generative Wasserstein Flows

ResearchDGX agent

arXiv:2605.31369v1 Announce Type: cross Abstract: Many modern generative models can be viewed as minimizing divergences between probability distributions, yet they rely on different algorithmic and ge

AdvScene: Rethinking Adversarial Patch Evaluation Through Scene Robustness

ApplicationsDGX agent

arXiv:2605.30578v1 Announce Type: cross Abstract: Adversarial patches are physical patterns attached to real objects to mislead AI vision systems. Their real-world risk is not determined by a single s

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

SafetyDGX agent

arXiv:2605.30742v1 Announce Type: new Abstract: This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topi

Astra: a generalizable report generation foundation model for 3D computed tomography

ApplicationsDGX agent

arXiv:2605.31437v1 Announce Type: new Abstract: CT interpretation requires radiologists to review hundreds of volumetric slices per examination, making reporting time-consuming and highly expertise-de

Authentication of Copy Detection Patterns via Cross-Camera Dual-Synthetic Referencing

ResearchDGX agent

arXiv:2605.31292v1 Announce Type: new Abstract: Copy Detection Patterns (CDPs) are structures printed on physical objects to enable cost-effective authentication. Verification is achieved by comparing

Automated Prediction of Postoperative Pancreatic Fistula Using Preoperative Computed Tomography

Model ReleasesDGX agent

arXiv:2605.31539v1 Announce Type: new Abstract: Postoperative pancreatic fistula (POPF) is a serious complication after pancreatic resection, increasing morbidity, hospital stay, and healthcare costs.

BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation

ResearchDGX agent

arXiv:2511.19394v2 Announce Type: replace Abstract: Segmenting small lesions in medical images remains notoriously difficult. Most prior work tackles this challenge by either designing better architec

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

ResearchDGX agent

arXiv:2605.31246v1 Announce Type: cross Abstract: Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing a

Benchmarking Single-Step Inpainting Methods for Multi-Object 3D Gaussian Splatting Scenes

ApplicationsDGX agent

arXiv:2605.30987v1 Announce Type: new Abstract: The tasks of object removal and inpainting 3D Gaussian Splatting (3DGS) scenes face challenges such as 3D consistency across camera views. In comparing

Beyond Accuracy: Evaluating Efficiency, Robustness and Explainability in Deep Learning for Malaria Diagnosis

Local AiDGX agent

arXiv:2605.30734v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in sub-Saharan Africa, where scarce diagnostic infrastructure makes timely, accurate diagnosis particular

BIAS-ID: A Framework for Analyzing Transformation Biases in AI-Generated Image Detectors

SafetyDGX agent

arXiv:2605.31153v1 Announce Type: new Abstract: Given the surge of harmful AI-generated imagery online, reliably distinguishing authentic images from generated ones has become an urgent research topic

BiSegMamba: Efficient Bidirectional Tri-Oriented Mamba for 3D Medical Image Segmentation

SafetyDGX agent

arXiv:2605.30972v1 Announce Type: new Abstract: Accurate 3D medical image segmentation requires both long-range volumetric context and fine boundary preservation. CNN-based methods have limited global

← Previous
1…104105106107108…211
Next →