AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

DGX agent

arXiv:2605.20278v1 Announce Type: cross Abstract: Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the

model-releasesarxiv-cs-cv
21 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training

DGX agent

arXiv:2605.21372v1 Announce Type: new Abstract: Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world dri

agentsarxiv-cs-cv
21 May 2026
Model Releases

Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection

DGX agent

arXiv:2605.20301v1 Announce Type: new Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion ofte

model-releasesarxiv-cs-cv
21 May 2026
Applications

Comparative Analysis of Military Detection Using Drone Imagery Across Multiple Visual Spectrums

DGX agent

arXiv:2605.21157v1 Announce Type: new Abstract: In modern warfare, drones are becoming an essential part of intelligence gathering and carrying out precise attacks in different kinds of hostile enviro

applicationsarxiv-cs-cv
21 May 2026
Safety

Comparative Evaluation of Deep Learning Models for Fake Image Detection

DGX agent

arXiv:2605.20971v1 Announce Type: new Abstract: The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of

safetyarxiv-cs-cv
21 May 2026
Research

ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning

DGX agent

arXiv:2605.20385v1 Announce Type: new Abstract: Recent progress in promptable segmentation has shifted visual perception from object-level localization toward concept-level understanding. However, the

researcharxiv-cs-cv
21 May 2026
Research

Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards

DGX agent

arXiv:2605.20758v1 Announce Type: cross Abstract: Inference-time guided sampling steers state-of-the-art diffusion and flow models without fine-tuning by interpreting the generation process as a contr

researcharxiv-cs-cv
21 May 2026
Model Releases

Continual Segmentation under Joint Nonstationarity

DGX agent

arXiv:2605.20538v1 Announce Type: new Abstract: Evolving data streams induce joint nonstationarity in continual semantic segmentation, where semantic classes, input distributions, and supervision avai

model-releasesarxiv-cs-cv
21 May 2026
Research

DAMA: Disentangled Body-Anchored Gaussians for Controllable Multi-Layered Avatars

DGX agent

arXiv:2605.21001v1 Announce Type: new Abstract: Existing 3D clothed avatar reconstruction methods achieve high visual fidelity but ignore geometric structure and physical plausibility. They either mod

researcharxiv-cs-cv
21 May 2026
Model Releases

DarkShake-DVS: Event-based Human Action Recognition under Low-light andShaking Camera Conditions

DGX agent

arXiv:2605.20680v1 Announce Type: new Abstract: Human Action Recognition (HAR) is a fundamental computer vision task with diverse real-world applications. Practical deployments often involve low-light

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction

DGX agent

arXiv:2605.20807v1 Announce Type: new Abstract: Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods

model-releasesarxiv-cs-cv
21 May 2026
Safety

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations

DGX agent

arXiv:2605.20732v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) often exploit spurious correlations in datasets, learning superficially predictive yet causally irrelevant features

safetyarxiv-cs-cv
21 May 2026
Research

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models

DGX agent

arXiv:2603.14184v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) often suffer from perceptual impairments under extended reasoning modes, particularly in visual question an

researcharxiv-cs-cv
21 May 2026
Research

Deformba: Vision State Space Model with Adaptive State Fusion

DGX agent

arXiv:2605.21308v1 Announce Type: new Abstract: State Space Models (SSMs) have emerged as a powerful and efficient alternative to Transformers, demonstrating linear-time complexity and exceptional seq

researcharxiv-cs-cv
21 May 2026
Model Releases

Deltaynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

DGX agent

arXiv:2605.20576v1 Announce Type: new Abstract: Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing

model-releasesarxiv-cs-cv
21 May 2026
Research

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

DGX agent

arXiv:2605.20766v1 Announce Type: new Abstract: Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two

researcharxiv-cs-cv
21 May 2026
Research

Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration

DGX agent

arXiv:2605.21381v1 Announce Type: new Abstract: Recent advances in Image Restoration (IR) have been largely driven by generative methods such as Diffusion Models and Flow Matching, which excel in synt

researcharxiv-cs-cv
21 May 2026
Research

Disentangling Sampling from Training Budget in Class-Imbalanced CT Body Composition Segmentation

DGX agent

arXiv:2605.20405v1 Announce Type: cross Abstract: Class imbalance is a fundamental challenge in medical image segmentation, where frequent classes typically dominate training at the expense of rare cl

researcharxiv-cs-cv
21 May 2026
Safety

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

DGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

safetyarxiv-cs-cv
21 May 2026
Model Releases

Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?

DGX agent

arXiv:2605.20448v1 Announce Type: new Abstract: Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-c

model-releasesarxiv-cs-cv
21 May 2026
Research

DrawMotion: Generating 3D Human Motions by Freehand Drawing

DGX agent

arXiv:2605.20955v1 Announce Type: new Abstract: Text-to-motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey t

researcharxiv-cs-cv
21 May 2026
Model Releases

DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions

DGX agent

arXiv:2605.21273v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) commonly introduce natural-language reasoning as an intermediate interface for end-to-end planning,

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Dynamic Video Generation: Shaping Video Generation Across Time and Space

DGX agent

arXiv:2605.21042v1 Announce Type: new Abstract: Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due t

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation

DGX agent

arXiv:2605.21028v1 Announce Type: new Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity w

model-releasesarxiv-cs-cv
21 May 2026
Local Ai

E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference

DGX agent

arXiv:2605.20717v1 Announce Type: cross Abstract: This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edg

local-aiarxiv-cs-cv
21 May 2026
Applications

E2GS: Event Enhanced Gaussian Splatting

DGX agent

arXiv:2406.14978v2 Announce Type: replace Abstract: Event cameras, known for their high dynamic range, absence of motion blur, and low energy usage, have recently found a wide range of applications th

applicationsarxiv-cs-cv
21 May 2026
Research

Early High-Frequency Injection for Geometry-Sensitive OOD Detection

DGX agent

arXiv:2605.20728v1 Announce Type: new Abstract: Post-hoc OOD detectors score logits or features after training, so their success depends on the geometry already encoded in the representation. We revis

researcharxiv-cs-cv
21 May 2026
Research

Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

DGX agent

arXiv:2605.20254v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, bec

researcharxiv-cs-cv
21 May 2026
Research

Efficient training for compact compression models via sequential distillation

DGX agent

arXiv:2601.05639v2 Announce Type: replace Abstract: Deep learning models for image compression often face practical limitations in hardware-constrained applications. Although these models achieve high

researcharxiv-cs-cv
21 May 2026
Research

ELEMENT: Multi-Modal Retinal Vessel Segmentation Based on a Coupled Region Growing and Machine Learning Approach

DGX agent

arXiv:2605.20458v1 Announce Type: new Abstract: Vascular structures in the retina contain important information for the detection and analysis of ocular diseases, including age-related macular degener

researcharxiv-cs-cv
21 May 2026
Local Ai

End-to-End Unmixing with Material Prompts for Hyperspectral Object Tracking

DGX agent

arXiv:2605.20569v1 Announce Type: new Abstract: Hyperspectral imagery encodes rich material properties that can improve tracking robustness under appearance ambiguity, illumination change, and backgro

local-aiarxiv-cs-cv
21 May 2026
Research

End2Reg: Learning Task-Specific Segmentation for Markerless Registration in Spine Surgery

DGX agent

arXiv:2512.13402v2 Announce Type: replace Abstract: Intraoperative navigation in spine surgery demands millimeter-level accuracy. Currently, this is achieved through radiation-intensive intraoperative

researcharxiv-cs-cv
21 May 2026
Research

EPC-3D-Diff: Equivariant Physics Consistent Conditional 3D Latent Diffusion for CBCT to CT Synthesis

DGX agent

arXiv:2605.20470v1 Announce Type: new Abstract: Cone-beam CT (CBCT) is routinely acquired during radiotherapy for patient setup, but its quantitative reliability is degraded by scatter, noise, and rec

researcharxiv-cs-cv
21 May 2026
Model Releases

Exploring Deep Learning and Ultra-Widefield Imaging for Diabetic Retinopathy and Macular Edema

DGX agent

arXiv:2603.08235v2 Announce Type: replace Abstract: Diabetic retinopathy (DR) and diabetic macular edema (DME) are leading causes of preventable blindness among working-age adults. Traditional approac

model-releasesarxiv-cs-cv
21 May 2026
Research

Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning

DGX agent

arXiv:2605.20551v1 Announce Type: new Abstract: Visual Place Recognition (VPR) aims to match a query image to reference images of the same place in a large-scale database. Recent state-of-the-art meth

researcharxiv-cs-cv
21 May 2026
Model Releases

FedCritic: Serverless Federated Critic Learning-based Resource Allocation for Multi-Cell OFDMA in 6G

DGX agent

arXiv:2605.21418v1 Announce Type: cross Abstract: In sixth-generation (6G) ultra-dense networks, aggressive frequency reuse amplifies inter-cell interference (ICI), making multi-cell orthogonal freque

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy

DGX agent

arXiv:2605.20965v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have shown remarkable performance on a wide range of vision-language tasks. Despite this progress, they are still p

model-releasesarxiv-cs-cv
21 May 2026
Applications

Findings of the Counter Turing Test: AI-Generated Image Detection

DGX agent

arXiv:2605.20787v1 Announce Type: new Abstract: The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of s

applicationsarxiv-cs-cv
21 May 2026
Safety

FineVision: Open Data Is All You Need

DGX agent

arXiv:2510.17269v2 Announce Type: replace Abstract: The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduc

safetyarxiv-cs-cv
21 May 2026
Safety

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

DGX agent

arXiv:2605.20910v1 Announce Type: new Abstract: Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free app

safetyarxiv-cs-cv
21 May 2026
Research

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models

DGX agent

arXiv:2605.20950v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a bottleneck of prohibitive computational costs arising from massive visual token sequences during inference. Existin

researcharxiv-cs-cv
21 May 2026
Model Releases

Free-Grained Hierarchical Visual Recognition

DGX agent

arXiv:2510.14737v3 Announce Type: replace Abstract: Hierarchical image recognition seeks to predict class labels along a semantic taxonomy, from broad categories to specific ones, typically under the

model-releasesarxiv-cs-cv
21 May 2026
Applications

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

DGX agent

arXiv:2605.20892v1 Announce Type: new Abstract: Fine-grained fruit classification is a critical yet challenging task in agricultural computer vision, primarily hindered by a severe shortage of high-qu

applicationsarxiv-cs-cv
21 May 2026
Local Ai

FTerViT: Fully Ternary Vision Transformer

DGX agent

arXiv:2605.21171v1 Announce Type: new Abstract: Ternary Vision Transformers offer substantial model compression, however state-of-the-art methods only ternarize the encoder layers, leaving patch embed

local-aiarxiv-cs-cv
21 May 2026
Model Releases

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation

DGX agent

arXiv:2605.20316v1 Announce Type: new Abstract: Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified visi

model-releasesarxiv-cs-cv
21 May 2026
Research

FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction

DGX agent

arXiv:2605.20287v1 Announce Type: cross Abstract: Standard cells form the building blocks of digital circuits, so their delay and power critically influence chip-level performance; yet characterizatio

researcharxiv-cs-cv
21 May 2026
Safety

GAMR: Geometric-Aware Manifold Regularization with Virtual Outlier Synthesis for Learning with Noisy Labels

DGX agent

arXiv:2605.20727v1 Announce Type: new Abstract: Deep neural networks (DNNs) experience significant performance degradation when processing noisy labels, primarily due to overfitting on mislabeled data

safetyarxiv-cs-cv
21 May 2026
Research

Gaze into the Details: Locality-Sensitive Enhancement for OCTA Retinal Vessel Segmentation

DGX agent

arXiv:2605.20651v1 Announce Type: new Abstract: Existing deep learning frameworks for Optical Coherence Tomography Angiography (OCTA) vessel segmentation are largely derived from the U-Net architectur

researcharxiv-cs-cv
21 May 2026
← Previous
1…150151152153154…263
Next →