AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Hardware

Test-Time Training with KV Binding Is Secretly Linear Attention

DGX agent

arXiv:2602.21204v3 Announce Type: replace-cross Abstract: Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a

hardwarearxiv-cs-cv
6 May 2026
Tutorials
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Text-Conditional JEPA for Learning Semantically Rich Visual Representations

DGX agent

arXiv:2605.03245v1 Announce Type: cross Abstract: Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature pre

tutorialsarxiv-cs-cv
6 May 2026
Local Ai

The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection

DGX agent

arXiv:2605.03642v1 Announce Type: new Abstract: Open-vocabulary object detection aims to recognize objects from an open set of categories, which leverages vision-language models (VLMs) pre-trained on

local-aiarxiv-cs-cv
6 May 2026
Research

Towards an End-to-End System for 3D Tracking of Physical Objects in Virtual Immersive Environments

DGX agent

arXiv:2605.02901v1 Announce Type: cross Abstract: This work aims to establish an end-to-end system for tracking of physical 3D objects for virtual reality (VR) applications. We focus on training appli

researcharxiv-cs-cv
6 May 2026
Safety

Tracing Like a Clinician: Anatomy-Guided Spatial Priors for Cephalometric Landmark Detection

DGX agent

arXiv:2605.03358v1 Announce Type: new Abstract: When orthodontists trace cephalometric radiographs, they follow a structured workflow: identify the soft tissue profile, partition the skull into anatom

safetyarxiv-cs-cv
6 May 2026
Research

TsallisPGD: Adaptive Gradient Weighting for Adversarial Attacks on Semantic Segmentation

DGX agent

arXiv:2605.03405v1 Announce Type: new Abstract: Attacking semantic segmentation models is significantly harder than image classification models because an attacker must flip thousands of pixel predict

researcharxiv-cs-cv
6 May 2026
Model Releases

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

DGX agent

arXiv:2605.03950v1 Announce Type: new Abstract: Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual ev

model-releasesarxiv-cs-cv
6 May 2026
Local Ai

Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers

DGX agent

arXiv:2605.03614v1 Announce Type: new Abstract: Visual affordances identify regions in an image with potential interactions, offering a novel paradigm for scene understanding. Recognizing affordances

local-aiarxiv-cs-cv
6 May 2026
Research

UniCorrn: Unified Correspondence Transformer Across 2D and 3D

DGX agent

arXiv:2605.04044v1 Announce Type: new Abstract: Visual correspondence across image-to-image (2D-2D), image-to-point cloud (2D-3D), and point cloud-to-point cloud (3D-3D) geometric matching forms the f

researcharxiv-cs-cv
6 May 2026
Research

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

DGX agent

arXiv:2605.03716v1 Announce Type: new Abstract: Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing met

researcharxiv-cs-cv
6 May 2026
Local Ai

Unsupervised Monocular Road Segmentation for Autonomous Driving via Scene Geometry

DGX agent

arXiv:2510.16790v2 Announce Type: replace Abstract: This paper presents a fully unsupervised approach for binary road segmentation (road vs. non-road), eliminating the reliance on costly manually labe

local-aiarxiv-cs-cv
6 May 2026
Model Releases

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

DGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

model-releasesarxiv-cs-cv
6 May 2026
Agents

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

DGX agent

arXiv:2603.28489v2 Announce Type: cross Abstract: The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as pot

agentsarxiv-cs-cv
6 May 2026
Local Ai

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection

DGX agent

arXiv:2605.03456v1 Announce Type: new Abstract: Open-world object detection aims to localize and recognize objects beyond a fixed closed-set label space. It is commonly divided into two categories, i.

local-aiarxiv-cs-cv
6 May 2026
Model Releases

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models

DGX agent

arXiv:2605.03351v1 Announce Type: new Abstract: Video vision-language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipe

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

DGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition

DGX agent

arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

DGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

model-releasesarxiv-cs-cv
6 May 2026
Agents

3D Gaussian Splatting against Moving Objects for High-Fidelity Street Scene Reconstruction

DGX agent

arXiv:2503.12001v4 Announce Type: replace Abstract: The accurate reconstruction of dynamic street scenes is critical for applications in autonomous driving, augmented reality, and virtual reality. Tra

agentsarxiv-cs-cv
5 May 2026
Applications

A Comprehensive Review of Fish Feeding Behavior Analysis in Aquaculture: Tasks, Techniques, and Applications

DGX agent

arXiv:2502.15311v3 Announce Type: replace Abstract: Fish feeding behavior analysis is a key foundation for intelligent feeding and precision aquaculture management, and plays an important role in impr

applicationsarxiv-cs-cv
5 May 2026
Research

A Coupled Fourth Order Telegraph Diffusion Framework Using Grayscale Indicators for Image Despeckling

DGX agent

arXiv:2605.00881v1 Announce Type: cross Abstract: Speckle noise severely limits the quality of images acquired from coherent imaging systems such as Synthetic Aperture Radar (SAR) and medical ultrasou

researcharxiv-cs-cv
5 May 2026
Applications

A Hybrid Approach for Closing the Sim2real Appearance Gap in Game Engine Synthetic Datasets

DGX agent

arXiv:2605.02291v1 Announce Type: new Abstract: Video game engines have been an important source for generating large volumes of visual synthetic datasets for training and evaluating computer vision a

applicationsarxiv-cs-cv
5 May 2026
Model Releases

A Light Weight Multi-Features-View Convolution Neural Network For Plant Disease Identification

DGX agent

arXiv:2605.00903v1 Announce Type: new Abstract: Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. Howe

model-releasesarxiv-cs-cv
5 May 2026
Research

A Proof-of-Concept Study of Multitask Learning for Cranial Synthetic CT Generation Across Heterogeneous MRI Field Strengths

DGX agent

arXiv:2605.00923v1 Announce Type: cross Abstract: Accurate synthesis of computed tomography (CT) images from magnetic resonance imaging (MRI) is clinically valuable for cranial applications such as at

researcharxiv-cs-cv
5 May 2026
Research

Act in Collusion: Distributed Multi-Target Backdoor Attacks in Federated Learning

DGX agent

arXiv:2411.03926v3 Announce Type: replace Abstract: Federated learning (FL) is widely used in Internet-of-Things (IoT) systems, but its distributed training process also exposes it to backdoor attacks

researcharxiv-cs-cv
5 May 2026
Research

Act2See: Emergent Active Visual Perception for Video Reasoning

DGX agent

arXiv:2605.01657v1 Announce Type: new Abstract: Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic in

researcharxiv-cs-cv
5 May 2026
Research

Active Reasoning Vision-Language Models via Sequential Experimental Design

DGX agent

arXiv:2605.01345v1 Announce Type: new Abstract: Visual perception in modern Vision-Language Models (VLMs) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field of view inevita

researcharxiv-cs-cv
5 May 2026
Research

Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion

DGX agent

arXiv:2605.02849v1 Announce Type: new Abstract: Diffusion models provide a powerful generative prior for perceptual reconstruction at ultra-low bitrates, but effective video compression requires contr

researcharxiv-cs-cv
5 May 2026
Model Releases

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

DGX agent

arXiv:2506.08849v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

DGX agent

arXiv:2605.01741v1 Announce Type: new Abstract: Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric anal

model-releasesarxiv-cs-cv
5 May 2026
Agents

Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving

DGX agent

arXiv:2605.00880v1 Announce Type: new Abstract: Autonomous driving (AD) is evolving towards end-to-end (E2E) frameworks through two primary paradigms: monolithic models exemplified by Vision-Language-

agentsarxiv-cs-cv
5 May 2026
Safety

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

DGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

safetyarxiv-cs-cv
5 May 2026
Hardware

AgriKD: Cross-Architecture Knowledge Distillation for Efficient Leaf Disease Classification

DGX agent

arXiv:2605.01355v1 Announce Type: new Abstract: Automated leaf disease classification is critical for early disease detection in resource-constrained field environments. Vision Transformers (ViTs) pro

hardwarearxiv-cs-cv
5 May 2026
Tutorials

AlbumFill: Album-Guided Reasoning and Retrieval for Personalized Image Completion

DGX agent

arXiv:2605.02892v1 Announce Type: new Abstract: Personalized image completion aims to restore occluded regions in personal photos while preserving identity and appearance. Existing methods either rely

tutorialsarxiv-cs-cv
5 May 2026
Research

Almost for Free: Crafting Adversarial Examples with Convolutional Image Filters

DGX agent

arXiv:2605.01098v1 Announce Type: cross Abstract: Adversarial examples in machine learning are typically generated using gradients, obtained either directly through access to the model or approximated

researcharxiv-cs-cv
5 May 2026
Model Releases

AnchorD: Metric Grounding of Monocular Depth Using Factor Graphs

DGX agent

arXiv:2605.02667v1 Announce Type: cross Abstract: Dense and accurate depth estimation is essential for robotic manipulation, grasping, and navigation, yet currently available depth sensors are prone t

model-releasesarxiv-cs-cv
5 May 2026
Safety

Anomaly-Preference Image Generation

DGX agent

arXiv:2605.02439v1 Announce Type: new Abstract: Synthesizing realistic and diverse anomalous samples from limited data is vital for robust model generalization. However, existing methods struggle to r

safetyarxiv-cs-cv
5 May 2026
Research

Application Research of a Deep Learning Model Integrating CycleGAN and YOLO in PCB Infrared Defect Detection

DGX agent

arXiv:2601.00237v2 Announce Type: replace Abstract: This paper addresses the critical bottleneck of infrared (IR) data scarcity in Printed Circuit Board (PCB) defect detection by proposing a cross-mod

researcharxiv-cs-cv
5 May 2026
Tutorials

Asymmetric Invertible Threat: Learning Reversible Privacy Defense for Face Recognition

DGX agent

arXiv:2605.01217v1 Announce Type: new Abstract: Face Recognition systems are widely deployed in real-world applications, but they also raise privacy concerns due to unauthorized collection and misuse

tutorialsarxiv-cs-cv
5 May 2026
Model Releases

AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT

DGX agent

arXiv:2605.01480v1 Announce Type: new Abstract: We study training-free image editing on Qwen-Image-Edit-2511, a 60-block multi-modal diffusion transformer (MMDiT) that concatenates noise and source-im

model-releasesarxiv-cs-cv
5 May 2026
Agents

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding

DGX agent

arXiv:2605.02630v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. Howeve

agentsarxiv-cs-cv
5 May 2026
Research

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection

DGX agent

arXiv:2605.02567v1 Announce Type: new Abstract: The rapid advancement of generative Artificial Intelligence (AI) has introduced significant challenges for reliable AI-generated image detection. Existi

researcharxiv-cs-cv
5 May 2026
Model Releases

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

DGX agent

arXiv:2506.09082v5 Announce Type: replace Abstract: The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as gener

model-releasesarxiv-cs-cv
5 May 2026
Agents

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

DGX agent

arXiv:2512.10571v4 Announce Type: replace Abstract: Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, exi

agentsarxiv-cs-cv
5 May 2026
Model Releases

BadmintonGRF: A Multimodal Dataset and Benchmark for Markerless Ground Reaction Force Estimation in Badminton

DGX agent

arXiv:2605.01876v1 Announce Type: new Abstract: Multimodal resources for non-periodic court sports with laboratory-grade sensing remain scarce: few publicly pair instrumented ground reaction force (GR

model-releasesarxiv-cs-cv
5 May 2026
Safety

Behavior-Grounded Lane Representation Learning for Multi-Task Traffic Digital Twins

DGX agent

arXiv:2605.01901v1 Announce Type: new Abstract: Traffic digital twins are powerful tools for advanced traffic management, and most systems are built on static geometric representations. However, these

safetyarxiv-cs-cv
5 May 2026
Agents

Beyond Known Objects: A Novel Framework for Open-Set Object Detection using Negative-Aware Norm

DGX agent

arXiv:2605.02284v1 Announce Type: new Abstract: Open-Set Object Detection (OSOD) is crucial for autonomous driving, where perception systems must recognize and localize both known and previously unsee

agentsarxiv-cs-cv
5 May 2026
Safety

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs

DGX agent

arXiv:2605.01324v1 Announce Type: new Abstract: Although reinforcement learning (RL) has significantly advanced reasoning capabilities in large multimodal language models (MLLMs), its efficacy remains

safetyarxiv-cs-cv
5 May 2026
← Previous
1…195196197198199…263
Next →