AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration

ApplicationsDGX agent

arXiv:2606.21240v1 Announce Type: cross Abstract: Training datasets have tremendous proprietary value and are vulnerable to unauthorized copying. Existing defenses mainly focus on tracking individual

Discovering Latent Groups for Robust Classification

ResearchDGX agent

arXiv:2606.23609v1 Announce Type: cross Abstract: Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately on underrepresented subgroups. E

DivCon-NeRF: Diverse and Consistent Ray Augmentation for Few-Shot NeRF

ResearchDGX agent

arXiv:2503.12947v3 Announce Type: replace Abstract: Neural Radiance Field (NeRF) has shown remarkable performance in novel view synthesis but requires numerous multi-view images, limiting its practica


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy

Model ReleasesDGX agent

arXiv:2509.17901v4 Announce Type: replace Abstract: Speech and audio encoders developed over years of community effort are routinely excluded from video understanding pipelines, not because they fail,

DR-Mamba: Automatic Inference-Time Domain Adaptation for Document Image Binarization via Sample-Conditioned Detail-Background Suppression

Model ReleasesDGX agent

arXiv:2606.22625v1 Announce Type: new Abstract: Degraded document image binarization is sensitive to domain shifts caused by paper aging, bleed-through, stains, shadows, and uneven illumination, and t

DreamUV: Unwrap Artist-like UV by End-to-End Flow Matching

ApplicationsDGX agent

arXiv:2606.22445v1 Announce Type: new Abstract: UV parameterization is a fundamental step in 3D content creation, yet producing production-ready UV layouts remains challenging due to the gap between g

DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction

Model ReleasesDGX agent

arXiv:2606.23031v1 Announce Type: new Abstract: Reconstructing dynamic urban scenes remains challenging due to the unbounded nature of driving environments and the presence of multiple dynamic objects

Dual-Stream EEG Decoding for 3D Visual Perception

ResearchDGX agent

arXiv:2606.22182v1 Announce Type: new Abstract: This paper explores a novel brain decoding model for 3D shape perception through a dual pathway architecture mirroring biological vision. Our bio-inspir

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

Model ReleasesDGX agent

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that prov

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

ResearchDGX agent

arXiv:2512.10324v2 Announce Type: replace Abstract: Audio-Visual Large Language Models (AV-LLMs) face prohibitive computational costs of processing massive, redundant audio-visual tokens. Existing uni

Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT-Quantization Embedding

Local AiDGX agent

arXiv:2606.22285v1 Announce Type: new Abstract: Localizing document tampering is extremely challenging, as manipulations are crafted to appear visually consistent and often leave only subtle traces th

Efficient Traffic State Prediction With Dynamic Joint Spatio-Temporal Relation Inference

Local AiDGX agent

arXiv:2504.08061v2 Announce Type: replace Abstract: Traffic prediction is difficult due to the complex interplay of temporal evolution, spatial interactions, and delayed spatio-temporal propagation ov

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Model ReleasesDGX agent

arXiv:2510.26113v2 Announce Type: replace Abstract: Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introd

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion

Model ReleasesDGX agent

arXiv:2606.20924v1 Announce Type: new Abstract: In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image s

EmbodiedUS-FS: Fast Slow Intelligence for Ultrasound Robotics

SafetyDGX agent

arXiv:2606.22319v1 Announce Type: cross Abstract: Robotic ultrasound scanning in real clinical environments requires both high-level clinical workflow reasoning and low-level closed-loop execution. Ph

Enhancing IMU-Based Online Handwriting Recognition via Contrastive Learning with Zero Inference Overhead

SafetyDGX agent

arXiv:2602.07049v2 Announce Type: replace Abstract: Online handwriting recognition using inertial measurement units opens up handwriting on paper as input for digital devices. Doing it on edge hardwar

Enhancing Road Safety: An IoT-Based Accident Detection and Prevention Mechanism

SafetyDGX agent

arXiv:2606.22381v1 Announce Type: cross Abstract: Road traffic accidents remain a critical global crisis, consistently serving as a primary driver of preventable mortality and severe injury. These inc

Enlight: Fast Low-Light Image Enhancement via Multi-Objective Optimization and Shadow-Aware Refinement

Local AiDGX agent

arXiv:2606.21674v1 Announce Type: new Abstract: We present ENLIGHT, a fast and training free framework for low-light image enhancement based on direct optimization of a perceptual objective. Unlike de

EnTrust: Modeling Inter-Modal Conflict for Trustworthy Multimodal Medical Image Analysis

Local AiDGX agent

arXiv:2606.21384v1 Announce Type: new Abstract: Multimodal medical imaging fuses complementary anatomical and functional information, yet modalities frequently disagree in pathologically heterogeneous

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

Model ReleasesDGX agent

arXiv:2606.22948v1 Announce Type: cross Abstract: As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

Model ReleasesDGX agent

arXiv:2606.20177v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to compre

Evaluating self-supervised echocardiographic representations across downstream extraction strategies for left-ventricular segmentation and ejection fraction estimation

ResearchDGX agent

arXiv:2606.22943v1 Announce Type: new Abstract: Self-supervised learning (SSL) is increasingly used in medical imaging to reduce annotation requirements, but representation quality is often judged usi

Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images

Model ReleasesDGX agent

arXiv:2606.20723v1 Announce Type: new Abstract: Chronic wound assessment remains a clinically challenging task that requires accurate interpretation of wound morphology, tissue composition, vascular c

Evo-RAD: Navigating Rare Retinal Disease Diagnosis via Self-Evolving Agentic Retrieval

Model ReleasesDGX agent

arXiv:2606.22955v1 Announce Type: new Abstract: Large-scale pretrained foundation models have revolutionized general medical screening, but often falter on rare diseases because such conditions are un

Expert Consensus on Criteria for the Automated Assessment of Laparoscopic Camera Navigation

SafetyDGX agent

arXiv:2606.23131v1 Announce Type: new Abstract: Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are tim

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

SafetyDGX agent

arXiv:2606.21197v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioni

Failure Analysis in Transition: An Industry Survey of Challenges, Priorities, and Standardization Needs in Advanced Packaging and Heterogeneous Integration

ResearchDGX agent

arXiv:2606.22149v1 Announce Type: cross Abstract: Failure analysis is being reshaped by heterogeneous integration, chiplet-based architectures, hybrid bonding, backside technologies, & increasingly bu

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

ResearchDGX agent

arXiv:2606.23354v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their 'black-box' nature hinders deplo

FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction

ResearchDGX agent

arXiv:2508.18031v2 Announce Type: replace Abstract: Craniofacial reconstruction in forensics is one of the processes to identify victims of crime and natural disasters. Identifying an individual from

FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs

ResearchDGX agent

arXiv:2606.22875v1 Announce Type: new Abstract: Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ability to combine the powerful gen

Feed-forward Motion In-betweening for Any 4D

ResearchDGX agent

arXiv:2606.22131v1 Announce Type: new Abstract: 4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., ani

FetSelect: Task-Specific Architectures and Self-Supervised Learning for Automated Fetal Ultrasound Frame Selection

ResearchDGX agent

arXiv:2606.22487v1 Announce Type: new Abstract: Automated frame selection for fetal biometry remains under addressed, with most prior work targeting generic quality assessment or downstream measuremen

Few-Shot Hyperspectral Aphid Detection via FastGAN Synthetic Data Generation, Transformer-Based Classification and Explainable AI

ResearchDGX agent

arXiv:2606.21267v1 Announce Type: new Abstract: Early detection of aphid infestation in crops is essential for preventing yield loss and reducing unnecessary pesticide use. Hyperspectral imaging combi

Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution

ResearchDGX agent

arXiv:2606.21910v1 Announce Type: new Abstract: Arbitrary-scale image super-resolution (ASISR) aims to reconstruct high-resolution images from low-resolution inputs over a continuous range of upscalin

Fine-grained Human Motion Understanding with Language Models

ResearchDGX agent

arXiv:2606.20888v1 Announce Type: new Abstract: In this work, we propose methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal pos

FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

Local AiDGX agent

arXiv:2606.21373v1 Announce Type: new Abstract: Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training

Flow6D: Discrete-to-Continuous Flow Matching for Efficient and Accurate Category-Level 6D Pose Estimation

Local AiDGX agent

arXiv:2606.23293v1 Announce Type: new Abstract: 6D pose estimation is a key task in computer vision and embodied AI, widely used in robotic manipulation, augmented reality, etc. Existing methods direc

FlowCodec: One-Step Flow Prior for Generative Image Compression

Model ReleasesDGX agent

arXiv:2606.21030v1 Announce Type: cross Abstract: Diffusion-based image compression methods, leveraging powerful generative priors, have demonstrated remarkable perceptual quality at ultra-low bitrate

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation

ApplicationsDGX agent

arXiv:2606.22424v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

SafetyDGX agent

arXiv:2606.20867v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-sho

Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection

Model ReleasesDGX agent

arXiv:2606.22378v1 Announce Type: new Abstract: Event cameras enable high-frequency visual perception with microsecond latency, offering advantages for dynamic scenes. However, event-based small objec

ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

ResearchDGX agent

arXiv:2602.06226v2 Announce Type: replace Abstract: The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reco

From Convolution to Transformer: A Comparative Study of U-Net Variants for Brain Tumor and Retinal Vessel Segmentation

Model ReleasesDGX agent

arXiv:2606.22168v1 Announce Type: new Abstract: Medical image segmentation plays an important role in computer aided diagnosis, treatment planning, and disease monitoring. U-Net has been widely used f

From Driving Videos to Simulatable Scenarios

SafetyDGX agent

arXiv:2606.21993v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) face driving scenarios ranging from routine traffic to rare events. To assess safety it is crucial to reproduce these scenar

From Gradient Clipping to Structural Refinement: Improving DPSGD for Medical Image Segmentation

SafetyDGX agent

arXiv:2606.21763v1 Announce Type: new Abstract: Medical image segmentation is widely used for disease detection but relies on sensitive data, raising privacy concerns as trained models can leak inform

From Point Estimates to Distributions: GMM Pooling for MIL in Preterm Birth Prediction

Model ReleasesDGX agent

arXiv:2606.23005v1 Announce Type: new Abstract: Preterm birth (PTB) prediction can enable targeted surveillance and timely intervention, yet most ultrasound-based models use a single selected transvag

From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation

SafetyDGX agent

arXiv:2606.23486v1 Announce Type: new Abstract: Curvilinear object segmentation, including vessels and cracks, is challenging due to extreme spatial sparsity and topological fragility, where small loc

From Uncertainty to Stability and Fidelity: Guiding Sparse-View 3D Gaussian Splatting with Fisher Information

TutorialsDGX agent

arXiv:2606.20842v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising technique for novel view synthesis. However, 3DGS requires dense input views to achieve high-qua

Full-Body Golf Swing Kinematic Reconstruction From a Smartwatch IMU

TutorialsDGX agent

arXiv:2606.22876v1 Announce Type: new Abstract: Quantitative measurement of the golf swing is critical for evaluating technique and enabling individualized feedback. However, existing methods are impr

Fursee: Hybrid YOLO-DINOv3 Framework for Fursuit Identity Retrieval and Clustering

Model ReleasesDGX agent

arXiv:2606.22872v1 Announce Type: new Abstract: Global furry conventions produce massive fursuit photographs, while manual sorting brings heavy labor costs and calls for automatic identity retrieval a

G-MASt3R-SfM: Graph-based View Pruning and Multi-stage Optimization for Robust SfM

ResearchDGX agent

arXiv:2606.22856v1 Announce Type: new Abstract: Structure from Motion (SfM) is essential for multi-view 3D reconstruction, however, its accuracy heavily relies on the accuracy of image matching. While

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

TutorialsDGX agent

arXiv:2606.22416v1 Announce Type: new Abstract: We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video genera

Generating Fearful Images: Investigating Potential Emotional Biases in Image-Generation Models

Model ReleasesDGX agent

arXiv:2411.05985v3 Announce Type: replace-cross Abstract: This paper examines potential biases and inconsistencies in the emotions evoked by images produced by generative artificial intelligence (AI)

Generative Relightable Avatars

ResearchDGX agent

arXiv:2606.22718v1 Announce Type: new Abstract: We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-

GenTrack: A New Generation of Multi-Object Tracking

AgentsDGX agent

arXiv:2510.24399v3 Announce Type: replace Abstract: This paper introduces a novel multi-object tracking (MOT) method, dubbed GenTrack, whose main contributions include: first-a hybrid tracking approac

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

Model ReleasesDGX agent

arXiv:2606.23669v1 Announce Type: new Abstract: Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generi

GEOPHYS: The Geometry of Physical Plausibility

Model ReleasesDGX agent

arXiv:2606.20707v1 Announce Type: new Abstract: While humans can identify physically implausible events within milliseconds, machine learning approaches addressing the same problem are extremely slow

GIM-ENDO: A Multimodal Endoscopic Image and Video Dataset for Gastric Intestinal Metaplasia Morphology and Pathology

ResearchDGX agent

arXiv:2606.20919v1 Announce Type: new Abstract: Gastric intestinal metaplasia (GIM) is a precursor lesion to gastric dysplasia and adenocarcinoma whose early detection is crucial for intervening in th

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

ResearchDGX agent

arXiv:2606.20891v1 Announce Type: new Abstract: Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat separately. Point-track-conditione

Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding

Model ReleasesDGX agent

arXiv:2606.22409v1 Announce Type: new Abstract: Robots operating in everyday environments must understand fine-grained human actions, intentions, and contextual cues from broad views where people occu

← Previous
1…7778798081…211
Next →