AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

Model ReleasesDGX agent

arXiv:2602.19946v4 Announce Type: replace Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

ResearchDGX agent

arXiv:2507.13868v2 Announce Type: replace Abstract: Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their interna

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.17375v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have substantially enhanced their ability across multimodal video understanding benchmarks spanning tem

When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations

Model ReleasesDGX agent

arXiv:2603.22368v2 Announce Type: replace Abstract: Visualizations help communicate data insights, but deceptive data representations can distort their interpretation and propagate misinformation. Whi

When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization

Model ReleasesDGX agent

arXiv:2604.16855v1 Announce Type: new Abstract: Camouflaged object detection (COD) segments objects that intentionally blend with the background, so predictions depend on subtle texture and boundary c

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

Local AiDGX agent

arXiv:2604.17422v1 Announce Type: new Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of proces

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals

ResearchDGX agent

arXiv:2604.16745v1 Announce Type: cross Abstract: Training-free token reduction methods for Vision Transformers (ToMe, ToFu, PiToMe, and MCTF) employ different scoring mechanisms, yet they share a clo

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

AgentsDGX agent

arXiv:2604.18484v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex

ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection

ResearchDGX agent

arXiv:2604.17949v1 Announce Type: new Abstract: Deep learning-based industrial anomaly detectors often behave as black boxes, making it hard to justify decisions with physically meaningful defect evid

20 Apr 2026

A Single Image and Multimodality Is All You Need for Novel View Synthesis

Local AiDGX agent

arXiv:2602.17909v2 Announce Type: replace Abstract: Diffusion-based approaches have recently demonstrated strong performance for single-image novel view synthesis by conditioning generative models on

Adapting in the Dark: Efficient and Stable Test-Time Adaptation for Black-Box Models

Local AiDGX agent

arXiv:2604.15609v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) for black-box models accessible only via APIs remains a largely unexplored challenge. Existing approaches such as post-hoc

AdaVFM: Adaptive Vision Foundation Models for Edge Intelligence via LLM-Guided Execution

Local AiDGX agent

arXiv:2604.15622v1 Announce Type: new Abstract: Language-aligned vision foundation models (VFMs) enable versatile visual understanding for always-on contextual AI, but their deployment on edge devices

AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.16067v1 Announce Type: cross Abstract: Adapting pre-trained vision-language models (VLMs) for robotic control requires injecting high-magnitude continuous gradients from a flow-matching act

AeroDeshadow: Physics-Guided Shadow Synthesis and Penumbra-Aware Deshadowing for Aerospace Imagery

ApplicationsDGX agent

arXiv:2604.15903v1 Announce Type: new Abstract: Shadows are prevalent in high-resolution aerospace imagery (ASI). They often cause spectral distortion and information loss, which degrade downstream in

AHS: Adaptive Head Synthesis via Synthetic Data Augmentations

ApplicationsDGX agent

arXiv:2604.15857v1 Announce Type: new Abstract: Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping, where

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow

ResearchDGX agent

arXiv:2604.15809v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grou

An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval

ApplicationsDGX agent

arXiv:2503.22171v2 Announce Type: replace Abstract: Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. Mainstream research paradigm necessitates real-world person images with ma

APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition

Model ReleasesDGX agent

arXiv:2604.15708v1 Announce Type: new Abstract: The advent of deep neural networks has led to remarkable progress in 3D point cloud recognition, but they remain vulnerable to adversarial attacks. Alth

Art3D: Training-Free 3D Generation from Flat-Colored Illustration

Model ReleasesDGX agent

arXiv:2504.10466v2 Announce Type: replace Abstract: Large-scale pre-trained image-to-3D generative models have exhibited remarkable capabilities in diverse shape generations. However, most of them str

AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis

AgentsDGX agent

arXiv:2604.16024v1 Announce Type: cross Abstract: Vision Language Models (VLMs) have been applied to several specific domains and have shown strong problem-solving capabilities. However, astronomical

AutoDrive-R^2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

SafetyDGX agent

arXiv:2509.01944v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models in autonomous driving systems have recently demonstrated transformative potential by integrating multimoda

Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration

ResearchDGX agent

arXiv:2604.15829v1 Announce Type: new Abstract: Text-to-image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to impl

Breakout-picker: Reducing false positives in deep learning-based borehole breakout characterization from acoustic image logs

ResearchDGX agent

arXiv:2604.16011v1 Announce Type: new Abstract: Borehole breakouts are stress-induced spalling on the borehole wall, which are identifiable in acoustic image logs as paired zones with near-symmetry az

CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity Awareness

SafetyDGX agent

arXiv:2602.22159v2 Announce Type: replace Abstract: Arbitrary-Scale SR (ASISR) remains fundamentally limited by cross-scale distribution shift: once the inference scale leaves the training range, nois

Causal Bootstrapped Alignment for Unsupervised Video-Based Visible-Infrared Person Re-Identification

SafetyDGX agent

arXiv:2604.15631v1 Announce Type: new Abstract: VVI-ReID is a critical technique for all-day surveillance, where temporal information provides additional cues beyond static images. However, existing a

ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation

Model ReleasesDGX agent

arXiv:2508.10635v3 Announce Type: replace Abstract: Understanding environmental changes from remote sensing imagery is vital for climate resilience, urban planning, and ecosystem monitoring. Yet, curr

CLOTH-HUGS: Cloth Aware Human Gaussian Splatting

ResearchDGX agent

arXiv:2604.15875v1 Announce Type: new Abstract: We present Cloth-HUGS, a Gaussian Splatting based neural rendering framework for photorealistic clothed human reconstruction that explicitly disentangle

CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting

ResearchDGX agent

arXiv:2604.16240v1 Announce Type: new Abstract: Time-to-Collision (TTC) forecasting is a critical task in collision prevention, requiring precise temporal prediction and comprehending both local and g

Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep Learning

ResearchDGX agent

arXiv:2501.05281v2 Announce Type: replace Abstract: Continuous monitoring of glacier calving fronts is essential for sea level rise projections. This study benchmarks Deep Learning systems for front d

Concept-wise Attention for Fine-grained Concept Bottleneck Models

SafetyDGX agent

arXiv:2604.15748v1 Announce Type: new Abstract: Recently impressive performance has been achieved in Concept Bottleneck Models (CBM) by utilizing the image-text alignment learned by a large pre-traine

Continual Hand-Eye Calibration for Open-world Robotic Manipulation

Local AiDGX agent

arXiv:2604.15814v1 Announce Type: new Abstract: Hand-eye calibration through visual localization is a critical capability for robotic manipulation in open-world environments. However, most deep learni

CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment

HardwareDGX agent

arXiv:2604.15665v1 Announce Type: new Abstract: Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most research

Cross-modal learning for plankton recognition

TutorialsDGX agent

arXiv:2603.16427v2 Announce Type: replace Abstract: This paper considers self-supervised cross-modal coordination as a strategy enabling utilization of multiple modalities and large volumes of unlabel

CTSCAN: Evaluation Leakage in Chest CT Segmentation and a Reproducible Patient-Disjoint Benchmark

Model ReleasesDGX agent

arXiv:2604.15561v1 Announce Type: cross Abstract: Reported chest CT segmentation performance can be strongly inflated when train and test partitions mix slices from the same study. We present CTSCAN,

CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

Model ReleasesDGX agent

arXiv:2604.15555v1 Announce Type: new Abstract: Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing

DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs

ApplicationsDGX agent

arXiv:2604.16201v1 Announce Type: cross Abstract: Consumer LiDARs in mobile devices and robots typically output a single depth value per pixel. Yet internally, they record full time-resolved histogram

DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates

Model ReleasesDGX agent

arXiv:2604.16099v1 Announce Type: new Abstract: Tables condense key transactional and administrative information into compact layouts, but practical extraction requires more than text recognition: sys

Dental Panoramic Radiograph Analysis Using YOLO26 From Tooth Detection to Disease Diagnosis

ResearchDGX agent

arXiv:2604.16231v1 Announce Type: new Abstract: Panoramic radiography is a fundamental diagnostic tool in dentistry, offering a comprehensive view of the entire dentition with minimal radiation exposu

DINOv3 Beats Specialized Detectors: A Simple Foundation Model Baseline for Image Forensics

Local AiDGX agent

arXiv:2604.16083v1 Announce Type: new Abstract: With the rapid advancement of deep generative models, realistic fake images have become increasingly accessible, yet existing localization methods rely

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World

Model ReleasesDGX agent

arXiv:2512.23421v3 Announce Type: replace Abstract: World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the rea

DualTrack: Sensorless 3D Ultrasound needs Local and Global Context

Model ReleasesDGX agent

arXiv:2509.09530v2 Announce Type: replace Abstract: Three-dimensional ultrasound (US) offers many clinical advantages over conventional 2D imaging, yet its widespread adoption is limited by the cost a

DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo

ResearchDGX agent

arXiv:2506.13215v2 Announce Type: replace Abstract: Recently, patch deformation-based methods have demonstrated significant effectiveness in multi-view stereo due to their incorporation of deformable

DyTact: Capturing Dynamic Contacts in Hand-Object Manipulation

SafetyDGX agent

arXiv:2506.03103v2 Announce Type: replace Abstract: Reconstructing dynamic hand-object contacts is essential for realistic manipulation in AI character animation, XR, and robotics, yet it remains chal

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence

ResearchDGX agent

arXiv:2509.14977v2 Announce Type: replace Abstract: Ultrasound imaging has become the preferred imaging modality for early cancer screening due to its advantages of non-ionizing radiation, low cost, a

Efficient Video Diffusion Models: Advancements and Challenges

ApplicationsDGX agent

arXiv:2604.15911v1 Announce Type: new Abstract: Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains co

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

SafetyDGX agent

arXiv:2604.16044v1 Announce Type: new Abstract: Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these mo

Enhancing Hazy Wildlife Imagery: AnimalHaze3k and IncepDehazeGan

ResearchDGX agent

arXiv:2604.16284v1 Announce Type: new Abstract: Atmospheric haze significantly degrades wildlife imagery, impeding computer vision applications critical for conservation, such as animal detection, tra

EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond

ResearchDGX agent

arXiv:2411.18328v2 Announce Type: replace Abstract: Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with tradition

Fed3D: Federated 3D Object Detection

Local AiDGX agent

arXiv:2604.15795v1 Announce Type: new Abstract: 3D object detection models trained in one server plays an important role in autonomous driving, robotics manipulation, and augmented reality scenarios.

FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound

Model ReleasesDGX agent

arXiv:2512.22278v2 Announce Type: replace Abstract: The growing demand for prenatal ultrasound imaging has intensified a global shortage of trained sonographers, creating barriers to essential fetal h

Find, Fix, Reason: Context Repair for Video Reasoning

SafetyDGX agent

arXiv:2604.16243v1 Announce Type: new Abstract: Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

Model ReleasesDGX agent

arXiv:2604.16298v1 Announce Type: new Abstract: UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous mult

Frequency-Aware Flow Matching for High-Quality Image Generation

Model ReleasesDGX agent

arXiv:2604.15521v1 Announce Type: new Abstract: Flow matching models have emerged as a powerful framework for realistic image generation by learning to reverse a corruption process that progressively

From Articles to Canopies: Knowledge-Driven Pseudo-Labelling for Tree Species Classification using LLM Experts

ApplicationsDGX agent

arXiv:2604.16115v1 Announce Type: new Abstract: Hyperspectral tree species classification is challenging due to limited and imbalanced class labels, spectral mixing (overlapping light signatures from

From Competition to Coopetition: Coopetitive Training-Free Image Editing Based on Text Guidance

SafetyDGX agent

arXiv:2604.15948v1 Announce Type: new Abstract: Text-guided image editing, a pivotal task in modern multimedia content creation, has seen remarkable progress with training-free methods that eliminate

From Limited Labels to Open Domains:An Efficient Learning Method for Drone-view Geo-Localization

Local AiDGX agent

arXiv:2503.07520v5 Announce Type: replace Abstract: Traditional supervised drone-view geo-localization (DVGL) methods heavily depend on paired training data and encounter difficulties in learning cros

From Zero to Detail: A Progressive Spectral Decoupling Paradigm for UHD Image Restoration with New Benchmark

Model ReleasesDGX agent

arXiv:2604.15654v1 Announce Type: new Abstract: Ultra-high-definition (UHD) image restoration poses unique challenges due to the high spatial resolution, diverse content, and fine-grained structures p

GaussianFlow SLAM: Monocular Gaussian Splatting SLAM Guided by GaussianFlow

TutorialsDGX agent

arXiv:2604.15612v1 Announce Type: cross Abstract: Gaussian splatting has recently gained traction as a compelling map representation for SLAM systems, enabling dense and photo-realistic scene modeling

GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos

ApplicationsDGX agent

arXiv:2604.16214v1 Announce Type: new Abstract: Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments.

GenHSI: Controllable Generation of Human-Scene Interaction Videos

ResearchDGX agent

arXiv:2506.19840v2 Announce Type: replace Abstract: Large-scale pre-trained video diffusion models have exhibited remarkable capabilities in diverse video generation. However, existing solutions face

← Previous
1…186187188189190…209
Next →