AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
6 May 2026

StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning

Model ReleasesDGX agent

arXiv:2605.03927v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

SafetyDGX agent

arXiv:2605.03849v1 Announce Type: new Abstract: Distillation-based acceleration has become foundational for making autoregressive streaming video diffusion models practical, with distribution matching

Synthetic Data Generation for Long-Tail Medical Image Classification: A Case Study in Skin Lesions

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.03221v1 Announce Type: new Abstract: Long-tailed class distributions are pervasive in multi-class medical datasets and pose significant challenges for deep learning models which typically u

TACO: Trajectory Aligning Cross-view Optimisation

Local AiDGX agent

arXiv:2605.03315v1 Announce Type: new Abstract: Cross-View Geo-localisation (CVGL) matches ground imagery against satellite tiles to give absolute position fixes, an alternative to GNSS where signals

Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing

Model ReleasesDGX agent

arXiv:2605.03909v1 Announce Type: cross Abstract: Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor conf

Test-Time Training with KV Binding Is Secretly Linear Attention

HardwareDGX agent

arXiv:2602.21204v3 Announce Type: replace-cross Abstract: Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a

Text-Conditional JEPA for Learning Semantically Rich Visual Representations

TutorialsDGX agent

arXiv:2605.03245v1 Announce Type: cross Abstract: Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature pre

The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection

Local AiDGX agent

arXiv:2605.03642v1 Announce Type: new Abstract: Open-vocabulary object detection aims to recognize objects from an open set of categories, which leverages vision-language models (VLMs) pre-trained on

Towards an End-to-End System for 3D Tracking of Physical Objects in Virtual Immersive Environments

ResearchDGX agent

arXiv:2605.02901v1 Announce Type: cross Abstract: This work aims to establish an end-to-end system for tracking of physical 3D objects for virtual reality (VR) applications. We focus on training appli

Tracing Like a Clinician: Anatomy-Guided Spatial Priors for Cephalometric Landmark Detection

SafetyDGX agent

arXiv:2605.03358v1 Announce Type: new Abstract: When orthodontists trace cephalometric radiographs, they follow a structured workflow: identify the soft tissue profile, partition the skull into anatom

TsallisPGD: Adaptive Gradient Weighting for Adversarial Attacks on Semantic Segmentation

ResearchDGX agent

arXiv:2605.03405v1 Announce Type: new Abstract: Attacking semantic segmentation models is significantly harder than image classification models because an attacker must flip thousands of pixel predict

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

Model ReleasesDGX agent

arXiv:2605.03950v1 Announce Type: new Abstract: Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual ev

Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers

Local AiDGX agent

arXiv:2605.03614v1 Announce Type: new Abstract: Visual affordances identify regions in an image with potential interactions, offering a novel paradigm for scene understanding. Recognizing affordances

UniCorrn: Unified Correspondence Transformer Across 2D and 3D

ResearchDGX agent

arXiv:2605.04044v1 Announce Type: new Abstract: Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the f

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

ResearchDGX agent

arXiv:2605.03716v1 Announce Type: new Abstract: Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing met

Unsupervised Monocular Road Segmentation for Autonomous Driving via Scene Geometry

Local AiDGX agent

arXiv:2510.16790v2 Announce Type: replace Abstract: This paper presents a fully unsupervised approach for binary road segmentation (road vs. non-road), eliminating the reliance on costly manually labe

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Model ReleasesDGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

AgentsDGX agent

arXiv:2603.28489v2 Announce Type: cross Abstract: The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as pot

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection

Local AiDGX agent

arXiv:2605.03456v1 Announce Type: new Abstract: Open-world object detection aims to localize and recognize objects beyond a fixed closed-set label space. It is commonly divided into two categories, i.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.03351v1 Announce Type: new Abstract: Video vision-language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipe

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition

Model ReleasesDGX agent

arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

Model ReleasesDGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

5 May 2026

3D Gaussian Splatting against Moving Objects for High-Fidelity Street Scene Reconstruction

AgentsDGX agent

arXiv:2503.12001v4 Announce Type: replace Abstract: The accurate reconstruction of dynamic street scenes is critical for applications in autonomous driving, augmented reality, and virtual reality. Tra

A Comprehensive Review of Fish Feeding Behavior Analysis in Aquaculture: Tasks, Techniques, and Applications

ApplicationsDGX agent

arXiv:2502.15311v3 Announce Type: replace Abstract: Fish feeding behavior analysis is a key foundation for intelligent feeding and precision aquaculture management, and plays an important role in impr

A Coupled Fourth Order Telegraph Diffusion Framework Using Grayscale Indicators for Image Despeckling

ResearchDGX agent

arXiv:2605.00881v1 Announce Type: cross Abstract: Speckle noise severely limits the quality of images acquired from coherent imaging systems such as Synthetic Aperture Radar (SAR) and medical ultrasou

A Hybrid Approach for Closing the Sim2real Appearance Gap in Game Engine Synthetic Datasets

ApplicationsDGX agent

arXiv:2605.02291v1 Announce Type: new Abstract: Video game engines have been an important source for generating large volumes of visual synthetic datasets for training and evaluating computer vision a

A Light Weight Multi-Features-View Convolution Neural Network For Plant Disease Identification

Model ReleasesDGX agent

arXiv:2605.00903v1 Announce Type: new Abstract: Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. Howe

A Proof-of-Concept Study of Multitask Learning for Cranial Synthetic CT Generation Across Heterogeneous MRI Field Strengths

ResearchDGX agent

arXiv:2605.00923v1 Announce Type: cross Abstract: Accurate synthesis of computed tomography (CT) images from magnetic resonance imaging (MRI) is clinically valuable for cranial applications such as at

Act in Collusion: Distributed Multi-Target Backdoor Attacks in Federated Learning

ResearchDGX agent

arXiv:2411.03926v3 Announce Type: replace Abstract: Federated learning (FL) is widely used in Internet-of-Things (IoT) systems, but its distributed training process also exposes it to backdoor attacks

Act2See: Emergent Active Visual Perception for Video Reasoning

ResearchDGX agent

arXiv:2605.01657v1 Announce Type: new Abstract: Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic in

Active Reasoning Vision-Language Models via Sequential Experimental Design

ResearchDGX agent

arXiv:2605.01345v1 Announce Type: new Abstract: Visual perception in modern Vision-Language Models (VLMs) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field of view inevita

Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion

ResearchDGX agent

arXiv:2605.02849v1 Announce Type: new Abstract: Diffusion models provide a powerful generative prior for perceptual reconstruction at ultra-low bitrates, but effective video compression requires contr

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

Model ReleasesDGX agent

arXiv:2506.08849v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

Model ReleasesDGX agent

arXiv:2605.01741v1 Announce Type: new Abstract: Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric anal

Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2605.00880v1 Announce Type: new Abstract: Autonomous driving (AD) is evolving towards end-to-end (E2E) frameworks through two primary paradigms: monolithic models exemplified by Vision-Language-

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

SafetyDGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

AgriKD: Cross-Architecture Knowledge Distillation for Efficient Leaf Disease Classification

HardwareDGX agent

arXiv:2605.01355v1 Announce Type: new Abstract: Automated leaf disease classification is critical for early disease detection in resource-constrained field environments. Vision Transformers (ViTs) pro

AlbumFill: Album-Guided Reasoning and Retrieval for Personalized Image Completion

TutorialsDGX agent

arXiv:2605.02892v1 Announce Type: new Abstract: Personalized image completion aims to restore occluded regions in personal photos while preserving identity and appearance. Existing methods either rely

Almost for Free: Crafting Adversarial Examples with Convolutional Image Filters

ResearchDGX agent

arXiv:2605.01098v1 Announce Type: cross Abstract: Adversarial examples in machine learning are typically generated using gradients, obtained either directly through access to the model or approximated

AnchorD: Metric Grounding of Monocular Depth Using Factor Graphs

Model ReleasesDGX agent

arXiv:2605.02667v1 Announce Type: cross Abstract: Dense and accurate depth estimation is essential for robotic manipulation, grasping, and navigation, yet currently available depth sensors are prone t

Anomaly-Preference Image Generation

SafetyDGX agent

arXiv:2605.02439v1 Announce Type: new Abstract: Synthesizing realistic and diverse anomalous samples from limited data is vital for robust model generalization. However, existing methods struggle to r

Application Research of a Deep Learning Model Integrating CycleGAN and YOLO in PCB Infrared Defect Detection

ResearchDGX agent

arXiv:2601.00237v2 Announce Type: replace Abstract: This paper addresses the critical bottleneck of infrared (IR) data scarcity in Printed Circuit Board (PCB) defect detection by proposing a cross-mod

Asymmetric Invertible Threat: Learning Reversible Privacy Defense for Face Recognition

TutorialsDGX agent

arXiv:2605.01217v1 Announce Type: new Abstract: Face Recognition systems are widely deployed in real-world applications, but they also raise privacy concerns due to unauthorized collection and misuse

AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT

Model ReleasesDGX agent

arXiv:2605.01480v1 Announce Type: new Abstract: We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-im

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding

AgentsDGX agent

arXiv:2605.02630v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. Howeve

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection

ResearchDGX agent

arXiv:2605.02567v1 Announce Type: new Abstract: The rapid advancement of generative Artificial Intelligence (AI) has introduced significant challenges for reliable AI-generated image detection. Existi

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

Model ReleasesDGX agent

arXiv:2506.09082v5 Announce Type: replace Abstract: The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as gener

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

AgentsDGX agent

arXiv:2512.10571v4 Announce Type: replace Abstract: Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, exi

BadmintonGRF: A Multimodal Dataset and Benchmark for Markerless Ground Reaction Force Estimation in Badminton

Model ReleasesDGX agent

arXiv:2605.01876v1 Announce Type: new Abstract: Multimodal resources for non-periodic court sports with laboratory-grade sensing remain scarce: few publicly pair instrumented ground reaction force (GR

Behavior-Grounded Lane Representation Learning for Multi-Task Traffic Digital Twins

SafetyDGX agent

arXiv:2605.01901v1 Announce Type: new Abstract: Traffic digital twins are powerful tools for advanced traffic management, and most systems are built on static geometric representations. However, these

Beyond Known Objects: A Novel Framework for Open-Set Object Detection using Negative-Aware Norm

AgentsDGX agent

arXiv:2605.02284v1 Announce Type: new Abstract: Open-Set Object Detection (OSOD) is crucial for autonomous driving, where perception systems must recognize and localize both known and previously unsee

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs

SafetyDGX agent

arXiv:2605.01324v1 Announce Type: new Abstract: Although reinforcement learning (RL) has significantly advanced reasoning capabilities in large multimodal language models (MLLMs), its efficacy remains

Biological Spatial Priors Regularize Foundation Model Representations for Cross-Site MSI Generalization in Colorectal Cancer

Local AiDGX agent

arXiv:2605.02660v1 Announce Type: cross Abstract: Predicting microsatellite instability (MSI) status from routine hematoxylin and eosin (H&E) whole slide images (WSIs) offers a practical alternative t

BoostDream: Efficient Refining for High-Quality Text-to-3D Generation from Multi-View Diffusion

ResearchDGX agent

arXiv:2401.16764v5 Announce Type: replace Abstract: Witnessing the evolution of text-to-image diffusion models, significant strides have been made in text-to-3D generation. Currently, two primary para

Boosting Multimodal Remote Sensing Image Classification with Transformer-based Heterogeneously Salient Graph Representation

Model ReleasesDGX agent

arXiv:2311.10320v3 Announce Type: replace Abstract: Data collected by different modalities can provide a wealth of complementary information, such as hyperspectral image (HSI) to offer rich spectral-s

Breaking the Resolution Barrier: Arbitrary-resolution Deep Image Steganography Framework

ResearchDGX agent

arXiv:2601.15739v2 Announce Type: replace Abstract: Deep image steganography (DIS) has achieved significant results in capacity and invisibility. However, current paradigms enforce the secret image to

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

Model ReleasesDGX agent

arXiv:2605.00873v1 Announce Type: cross Abstract: The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing benchmarks

CADFit: Precise Mesh-to-CAD Program Generation with Hybrid Optimization

ApplicationsDGX agent

arXiv:2605.01171v1 Announce Type: new Abstract: Despite recent progress, recovering parametric CAD construction sequences from geometric input, such as meshes or point clouds, is a key challenge for d

CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models

ApplicationsDGX agent

arXiv:2605.01925v1 Announce Type: new Abstract: We introduce CADFS, a data-centric framework that enables large vision-language models to generate complex CAD design histories. Existing generative CAD

← Previous
1…154155156157158…209
Next →