AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
21 May 2026

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models

ResearchDGX agent

arXiv:2603.14184v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) often suffer from perceptual impairments under extended reasoning modes, particularly in visual question an

Deformba: Vision State Space Model with Adaptive State Fusion

ResearchDGX agent

arXiv:2605.21308v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as a powerful and efficient alternative to Transformers, demonstrating linear-time complexity and exceptional seq

Deltaynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.20576v1 Announce Type: new Abstract: Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

ResearchDGX agent

arXiv:2605.20766v1 Announce Type: new Abstract: Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two

Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration

ResearchDGX agent

arXiv:2605.21381v1 Announce Type: new Abstract: Recent advances in Image Restoration (IR) have been largely driven by generative methods such as Diffusion Models and Flow Matching, which excel in synt

Disentangling Sampling from Training Budget in Class-Imbalanced CT Body Composition Segmentation

ResearchDGX agent

arXiv:2605.20405v1 Announce Type: cross Abstract: Class imbalance is a fundamental challenge in medical image segmentation, where frequent classes typically dominate training at the expense of rare cl

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

SafetyDGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?

Model ReleasesDGX agent

arXiv:2605.20448v1 Announce Type: new Abstract: Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-c

DrawMotion: Generating 3D Human Motions by Freehand Drawing

ResearchDGX agent

arXiv:2605.20955v1 Announce Type: new Abstract: Text-to-motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey t

DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions

Model ReleasesDGX agent

arXiv:2605.21273v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) commonly introduce natural-language reasoning as an intermediate interface for end-to-end planning,

Dynamic Video Generation: Shaping Video Generation Across Time and Space

Model ReleasesDGX agent

arXiv:2605.21042v1 Announce Type: new Abstract: Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due t

DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation

Model ReleasesDGX agent

arXiv:2605.21028v1 Announce Type: new Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity w

E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

Local AiDGX agent

arXiv:2605.20717v1 Announce Type: cross Abstract: This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edg

E2GS: Event Enhanced Gaussian Splatting

ApplicationsDGX agent

arXiv:2406.14978v2 Announce Type: replace Abstract: Event cameras, known for their high dynamic range, absence of motion blur, and low energy usage, have recently found a wide range of applications th

Early High-Frequency Injection for Geometry-Sensitive OOD Detection

ResearchDGX agent

arXiv:2605.20728v1 Announce Type: new Abstract: Post-hoc OOD detectors score logits or features after training, so their success depends on the geometry already encoded in the representation. We revis

Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

ResearchDGX agent

arXiv:2605.20254v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, bec

Efficient training for compact compression models via sequential distillation

ResearchDGX agent

arXiv:2601.05639v2 Announce Type: replace Abstract: Deep learning models for image compression often face practical limitations in hardware-constrained applications. Although these models achieve high

ELEMENT: Multi-Modal Retinal Vessel Segmentation Based on a Coupled Region Growing and Machine Learning Approach

ResearchDGX agent

arXiv:2605.20458v1 Announce Type: new Abstract: Vascular structures in the retina contain important information for the detection and analysis of ocular diseases, including age-related macular degener

End-to-End Unmixing with Material Prompts for Hyperspectral Object Tracking

Local AiDGX agent

arXiv:2605.20569v1 Announce Type: new Abstract: Hyperspectral imagery encodes rich material properties that can improve tracking robustness under appearance ambiguity, illumination change, and backgro

End2Reg: Learning Task-Specific Segmentation for Markerless Registration in Spine Surgery

ResearchDGX agent

arXiv:2512.13402v2 Announce Type: replace Abstract: Intraoperative navigation in spine surgery demands millimeter-level accuracy. Currently, this is achieved through radiation-intensive intraoperative

EPC-3D-Diff: Equivariant Physics Consistent Conditional 3D Latent Diffusion for CBCT to CT Synthesis

ResearchDGX agent

arXiv:2605.20470v1 Announce Type: new Abstract: Cone-beam CT (CBCT) is routinely acquired during radiotherapy for patient setup, but its quantitative reliability is degraded by scatter, noise, and rec

Exploring Deep Learning and Ultra-Widefield Imaging for Diabetic Retinopathy and Macular Edema

Model ReleasesDGX agent

arXiv:2603.08235v2 Announce Type: replace Abstract: Diabetic retinopathy (DR) and diabetic macular edema (DME) are leading causes of preventable blindness among working-age adults. Traditional approac

Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning

ResearchDGX agent

arXiv:2605.20551v1 Announce Type: new Abstract: Visual Place Recognition (VPR) aims to match a query image to reference images of the same place in a large-scale database. Recent state-of-the-art meth

FedCritic: Serverless Federated Critic Learning-based Resource Allocation for Multi-Cell OFDMA in 6G

Model ReleasesDGX agent

arXiv:2605.21418v1 Announce Type: cross Abstract: In sixth-generation (6G) ultra-dense networks, aggressive frequency reuse amplifies inter-cell interference (ICI), making multi-cell orthogonal freque

Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy

Model ReleasesDGX agent

arXiv:2605.20965v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have shown remarkable performance on a wide range of vision-language tasks. Despite this progress, they are still p

Findings of the Counter Turing Test: AI-Generated Image Detection

ApplicationsDGX agent

arXiv:2605.20787v1 Announce Type: new Abstract: The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of s

FineVision: Open Data Is All You Need

SafetyDGX agent

arXiv:2510.17269v2 Announce Type: replace Abstract: The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduc

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

SafetyDGX agent

arXiv:2605.20910v1 Announce Type: new Abstract: Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free app

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models

ResearchDGX agent

arXiv:2605.20950v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a bottleneck of prohibitive computational costs arising from massive visual token sequences during inference. Existin

Free-Grained Hierarchical Visual Recognition

Model ReleasesDGX agent

arXiv:2510.14737v3 Announce Type: replace Abstract: Hierarchical image recognition seeks to predict class labels along a semantic taxonomy, from broad categories to specific ones, typically under the

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

ApplicationsDGX agent

arXiv:2605.20892v1 Announce Type: new Abstract: Fine-grained fruit classification is a critical yet challenging task in agricultural computer vision, primarily hindered by a severe shortage of high-qu

FTerViT: Fully Ternary Vision Transformer

Local AiDGX agent

arXiv:2605.21171v1 Announce Type: new Abstract: Ternary Vision Transformers offer substantial model compression, however state-of-the-art methods only ternarize the encoder layers, leaving patch embed

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation

Model ReleasesDGX agent

arXiv:2605.20316v1 Announce Type: new Abstract: Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified visi

FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction

ResearchDGX agent

arXiv:2605.20287v1 Announce Type: cross Abstract: Standard cells form the building blocks of digital circuits, so their delay and power critically influence chip-level performance; yet characterizatio

GAMR: Geometric-Aware Manifold Regularization with Virtual Outlier Synthesis for Learning with Noisy Labels

SafetyDGX agent

arXiv:2605.20727v1 Announce Type: new Abstract: Deep neural networks (DNNs) experience significant performance degradation when processing noisy labels, primarily due to overfitting on mislabeled data

Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation

ResearchDGX agent

arXiv:2605.20651v1 Announce Type: new Abstract: Existing deep learning frameworks for Optical Coherence Tomography Angiography (OCTA) vessel segmentation are largely derived from the U-Net architectur

Generation of Heterogeneous PET Images from Uniform Organ Activity Maps Using a Pretrained Domain-Adapted Diffusion Model

ResearchDGX agent

arXiv:2605.20267v1 Announce Type: new Abstract: Synthetic PET images are valuable for quantitative imaging workflow development, scalable virtual imaging trials, and deep learning model training, but

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

ResearchDGX agent

arXiv:2605.20476v1 Announce Type: new Abstract: Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are cont

Grounding Driving VLA via Inverse Kinematics

ResearchDGX agent

arXiv:2605.21061v1 Announce Type: new Abstract: Existing Driving VLAs predict trajectories while largely ignoring their visual tokens -- a phenomenon we trace not to insufficient training but to a str

GSA-YOLO: A High-Efficiency Framework via Structured Sparsity and Adaptive Knowledge Distillation for Real-Time X-ray Security Inspection

ResearchDGX agent

arXiv:2605.20669v1 Announce Type: new Abstract: X-ray security inspection requires accurate real-time detection of prohibited items, but existing models often struggle to balance the challenges of sev

HADS-Net:A Hybrid Attention-Augmented Dual-Stream Network with Physics-Informed Augmentation for Breast Ultrasound Image Classification

ResearchDGX agent

arXiv:2605.20536v1 Announce Type: new Abstract: Accurate classification of breast ultrasound images into benign, malignant, and normal categories is a critical clinical task complicated by speckle noi

HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

Model ReleasesDGX agent

arXiv:2605.20469v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible b

HAPS: Rethinking Image Similarity for Virtual Staining

ResearchDGX agent

arXiv:2605.20362v1 Announce Type: new Abstract: Virtual staining of histopathology images (e.g., H&E-IHC) is an emerging tool in digital pathology, enabling faster and cheaper workflows by synthesizin

HDMoE: A Hierarchical Decoupling-Fusion Mixture-of-Experts Framework for Multimodal Cancer Survival Prediction

Local AiDGX agent

arXiv:2605.20891v1 Announce Type: new Abstract: Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (eg Whole Slide Images (WSIs) and Gen

Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation

TutorialsDGX agent

arXiv:2605.20600v1 Announce Type: new Abstract: Autoregressive (AR) visual generation has achieved remarkable performance but suffers from high memory usage and low throughput, as it requires caching

Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label

ApplicationsDGX agent

arXiv:2605.20725v1 Announce Type: new Abstract: Learning with noisy labels in multimedia classification often combines external annotations and model predictions into a single reliability weight, even

How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study

AgentsDGX agent

arXiv:2604.06750v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poo

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

ResearchDGX agent

arXiv:2605.20388v1 Announce Type: new Abstract: Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is

Hybrid Machine Learning Model for Forest Height Estimation from TanDEM-X and Landsat Data

ResearchDGX agent

arXiv:2605.20997v1 Announce Type: new Abstract: Integrating machine learning (ML) with physical models (PM) has emerged as a promising way of retrieving geophysical parameters from remote sensing data

HyDAR-Pano3D: A Hybrid Disentangled Anatomical Recovery Framework for Panoramic-to-3D Reconstruction

TutorialsDGX agent

arXiv:2605.20827v1 Announce Type: new Abstract: Panoramic radiograph (PR) is fundamentally used in routine dental care, but it inherently provides only a two-dimensional (2D) projection of complex thr

Hyper-V2X: Hypernetworks for Estimating Epistemic and Aleatoric Uncertainty in Cooperative Bird's-Eye-View Semantic Segmentation

Model ReleasesDGX agent

arXiv:2605.21309v1 Announce Type: new Abstract: Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental repre

HyperBones: Realtime Bone-driven Neural Garment Simulation with Hypernetwork Conditioning

HardwareDGX agent

arXiv:2605.20460v1 Announce Type: cross Abstract: Recent advances in garment simulation have brought high-quality results closer to real-time performance. Physics-based simulators can produce accurate

Improving 3D Gaussian Splatting Compression by Scene-Adaptive Lattice Vector Quantization

ResearchDGX agent

arXiv:2509.13482v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) is rapidly gaining popularity for its photorealistic rendering quality and real-time performance, but it generates mass

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools

Local AiDGX agent

arXiv:2605.20682v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown remarkable capability in bridging visual perception and textual reasoning, enabling zero-shot unders

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

TutorialsDGX agent

arXiv:2605.21431v1 Announce Type: new Abstract: Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant stri

JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026

Model ReleasesDGX agent

arXiv:2605.20904v1 Announce Type: new Abstract: We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representat

JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA

Model ReleasesDGX agent

arXiv:2605.20284v1 Announce Type: new Abstract: Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, pa

LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

Model ReleasesDGX agent

arXiv:2506.16950v2 Announce Type: replace Abstract: Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from

Latent Dynamics for Full Body Avatar Animation

ResearchDGX agent

arXiv:2605.21478v1 Announce Type: new Abstract: Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic ele

Latent Space Guided Scenario Sampling for Multimodal Segmentation Under Missing Modalities

ApplicationsDGX agent

arXiv:2605.20372v1 Announce Type: new Abstract: Multimodal semantic segmentation benefits remote sensing analysis by combining complementary information from different sensor modalities. In real-world

← Previous
1…120121122123124…211
Next →