AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
25 May 2026

Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation

AgentsDGX agent

arXiv:2605.23257v1 Announce Type: cross Abstract: Navigating under non-stationary environment shifts poses a critical challenge for a Vision-and-Language Navigation (VLN) agent deployed in the wild. Y

U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025

ResearchDGX agent

arXiv:2605.23274v1 Announce Type: new Abstract: Retrieving events from large-scale video datasets is challenging due to complex temporal, spatial, and multimodal information. This paper presents U-CES

UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries

SafetyDGX agent

arXiv:2507.23372v2 Announce Type: replace Abstract: Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each othe


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

UniReg: A Universal Model for Controllable CT Image Registration

SafetyDGX agent

arXiv:2503.12868v2 Announce Type: replace Abstract: Learning-based medical image registration has matched the accuracy of conventional methods while offering superior computational efficiency. However

Using Ensemble Diffusion to Estimate Uncertainty for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2506.00560v2 Announce Type: replace-cross Abstract: End-to-end planning systems for autonomous driving are rapidly improving, especially in closed-loop simulation environments like CARLA. Many s

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation

Model ReleasesDGX agent

arXiv:2605.23381v1 Announce Type: new Abstract: Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

Model ReleasesDGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

Model ReleasesDGX agent

arXiv:2605.23518v1 Announce Type: new Abstract: Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in m

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

Model ReleasesDGX agent

arXiv:2605.23141v1 Announce Type: new Abstract: A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipula

Vision Transformers Need Better Token Interaction

SafetyDGX agent

arXiv:2605.23868v1 Announce Type: new Abstract: Vision Transformers (ViTs) can learn strong image-level representations while their patch representations become less effective for dense prediction dur

What Linear Probes Miss: Multi-View Probing for Weight-Space Learning

Model ReleasesDGX agent

arXiv:2605.23410v1 Announce Type: cross Abstract: The explosive growth of open-source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentat

22 May 2026

3D LULC classification using multispectral LiDAR and deep learning: current and prospective schemes

Model ReleasesDGX agent

arXiv:2605.22328v1 Announce Type: new Abstract: Land Use Land Cover (LULC) classification is essential for national 3D mapping, geospatial analysis, and sustainable planning. Multispectral (MS) LiDAR

4D-GSW: Kinematic-Aware Spatio-Temporal Consistent Watermarking for 4D Gaussian Splatting

ResearchDGX agent

arXiv:2605.22342v1 Announce Type: new Abstract: While 4D Gaussian Splatting (4DGS) has revolutionized high-fidelity dynamic reconstruction, safeguarding the intellectual property of these assets remai

4D Radar Semantic Segmentation of People in Field Conditions Using Temporal Multi-View Networks

ApplicationsDGX agent

arXiv:2404.05307v2 Announce Type: replace Abstract: Reliable people detection is crucial for the safe autonomy of mobile robots and heavy vehicles, both on roads and in industrial settings like mining

A Robust Semantic Segmentation Pipeline for the CVPR 2026 8th UG2+ Challenge Track 2

ResearchDGX agent

arXiv:2605.22216v1 Announce Type: new Abstract: This report presents our solution for the WeatherProof Dataset Challenge, namely CVPR 2026 8th UG2+ Challenge Track 2: Semantic Segmentation in Adverse

A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient

Model ReleasesDGX agent

arXiv:2605.21500v1 Announce Type: cross Abstract: Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offeri

Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?

ResearchDGX agent

arXiv:2605.21642v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly augmented with continuous or latent non-textual tokens intended to support 'visual thinking.' Despite imp

Accelerating Vision Foundation Models with Drop-in Depthwise Convolution

ResearchDGX agent

arXiv:2605.22132v1 Announce Type: new Abstract: Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones

AesFormer: Transform Everyday Photos into Beautiful Memories

Model ReleasesDGX agent

arXiv:2605.22126v1 Announce Type: new Abstract: In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that ex

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

Model ReleasesDGX agent

arXiv:2605.22366v1 Announce Type: new Abstract: Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However,

AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

Model ReleasesDGX agent

arXiv:2605.22034v1 Announce Type: new Abstract: Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, en

AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

Model ReleasesDGX agent

arXiv:2512.20538v2 Announce Type: replace Abstract: Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter,

An Evidence Hierarchy for Bayesian Object Classification via OSINT-Aided Heterogeneous Sensor Fusion

ResearchDGX agent

arXiv:2605.22259v1 Announce Type: cross Abstract: Heterogeneous sensor fusion is vital for detecting, localizing, and classifying CBRNE threats. However, individual sensors are often only capable of d

An Open Multi-Center Whole-Body FDG PET/CT Foundation Model for Tumor Segmentation

ResearchDGX agent

arXiv:2605.21835v1 Announce Type: cross Abstract: The synergistic interpretation of anatomical information from computed tomography (CT) and metabolic information from positron emission tomography (PE

AtomicMotion: Learning Human Motion From Different Human Parts

Local AiDGX agent

arXiv:2605.22631v1 Announce Type: new Abstract: Accurately reconstructing full-body poses from sparse head and hand trajectories is a foundational challenge for immersive AR/VR telepresence. Current m

Attacking the Spike: On the Transferability and Security of Spiking Neural Networks to Adversarial Examples

ResearchDGX agent

arXiv:2209.03358v5 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs) have attracted much attention for their high energy efficiency and recent advances in classification performanc

AVI-HT: Adaptive Vision-IMU Fusion for 3D Hand Tracking

ApplicationsDGX agent

arXiv:2605.21714v1 Announce Type: new Abstract: We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU si

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.22816v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of

Balancing Uncertainty and Diversity of Samples: Leveraging Diversity of Least, High Confidence Samples for Effective Active Learning

TutorialsDGX agent

arXiv:2605.22169v1 Announce Type: new Abstract: Deep learning models, including Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have achieved state-of-the-art performance on vario

Bernini: Latent Semantic Planning for Video Diffusion

ResearchDGX agent

arXiv:2605.22344v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimo

Beyond Chamfer Distance: Granular Order-aware Evaluation Metric For Online Mapping

Model ReleasesDGX agent

arXiv:2605.22578v1 Announce Type: new Abstract: Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (S

BodyReLux: Temporally Consistent Full-Body Video Relighting

ApplicationsDGX agent

arXiv:2605.21766v1 Announce Type: new Abstract: Being able to relight human performance is a fundamental task for post production and content creation. We present BodyReLux, a subject-specific video d

Bounding-Box Trajectories Matter for Video Anomaly Detection

SafetyDGX agent

arXiv:2605.21957v1 Announce Type: new Abstract: Video anomaly detection is critical for public safety and security, yet remains highly challenging despite extensive research due to large variations in

Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded Generations

ResearchDGX agent

arXiv:2605.22050v1 Announce Type: new Abstract: While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks.

Cambrian-P: Pose-Grounded Video Understanding

ResearchDGX agent

arXiv:2605.22819v1 Announce Type: new Abstract: Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video fram

Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction

ApplicationsDGX agent

arXiv:2602.06676v4 Announce Type: replace Abstract: Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real-world forensic scenarios. Compare

Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement

SafetyDGX agent

arXiv:2605.22547v1 Announce Type: new Abstract: Deep learning has brought significant progress to medical image classification, yet most existing methods still rely on isolated visual evidence and can

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

ResearchDGX agent

arXiv:2602.02214v3 Announce Type: replace Abstract: To achieve real-time interactive video generation, current methods distill pretrained bidirectional video diffusion models into few-step autoregress

Cell Phantom Video Generation in Elliptical Fourier Descriptor Domain

ResearchDGX agent

arXiv:2605.22563v1 Announce Type: new Abstract: Training Deep Neural Networks for tracking individual cells in biomedical videos requires a large amount of annotated data. The annotation of videos for

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

SafetyDGX agent

arXiv:2505.16416v3 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in large language models, but when applied to vision-language models (VLMs) it couples text and i

COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition

Model ReleasesDGX agent

arXiv:2605.22068v1 Announce Type: new Abstract: We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained gra

Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models

TutorialsDGX agent

arXiv:2605.22679v1 Announce Type: new Abstract: Vision-language models learn powerful multimodal embeddings, yet their internal semantics remain opaque. While sparse autoencoders (SAEs) can extract in

ConvNeXt-FD: A Fractal-Based Deep Model for Robust Biomedical Image Segmentation

ResearchDGX agent

arXiv:2605.22002v1 Announce Type: new Abstract: Biomedical image segmentation is a critical task in medical diagnosis and treatment planning, enabling precise delineation of anatomical structures and

Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions

TutorialsDGX agent

arXiv:2605.22697v1 Announce Type: new Abstract: Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real-world scenarios, where action cate

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.21854v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and co

CryoNet: A Deep Learning Framework for Multi-Modal Debris-Covered Glacier Mapping. A Case Study of the Poiqu Basin, Central Himalaya

ApplicationsDGX agent

arXiv:2605.21527v1 Announce Type: cross Abstract: Glaciers play a critical role as freshwater reserves and indicators of climate change, yet their automatic delineation, especially for debris-covered

D3Seg: Dependency-Aware Diffusion for Brain Tumor Segmentation with Missing Modalities

ResearchDGX agent

arXiv:2605.22249v1 Announce Type: new Abstract: Accurate brain tumor segmentation using multiparametric MRI is critical for effective treatment planning. However, in clinical settings, complete acquis

Decoupling Ego-Motion from Target Dynamics via Dual-Interval Motion Cues for UAV Detection

ResearchDGX agent

arXiv:2605.22605v1 Announce Type: cross Abstract: Object detection from Unmanned Aerial Vehicles (UAVs) is challenged by severe ego-motion, camera jitter, and large scale variations. While modern dete

DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders

ResearchDGX agent

arXiv:2605.22777v1 Announce Type: new Abstract: Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations th

Demystifying Transition Matching: When and Why It Can Beat Flow Matching

ApplicationsDGX agent

arXiv:2510.17991v3 Announce Type: replace-cross Abstract: Flow Matching (FM) underpins many state-of-the-art generative models, yet recent results indicate that Transition Matching (TM) can achieve hi

Depth Augmented and FE Free 3D/2D Liver Registration for Laparoscopic Liver AR

SafetyDGX agent

arXiv:2602.17517v2 Announce Type: replace Abstract: Augmented reality (AR) guidance in laparoscopic liver surgery requires accurate registration of preoperative 3D models to intraoperative 2D video, b

Detection of Virus and Small Cell Patches in Foci Images Using Switchable Convolution and Feature Pyramid Networks

ResearchDGX agent

arXiv:2605.22290v1 Announce Type: new Abstract: Accurate detection and counting of virus patches in focus-forming unit (FFU) images, also known as foci images, are important for quantifying viral infe

Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction

AgentsDGX agent

arXiv:2605.22420v1 Announce Type: new Abstract: Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural re

Direct content-based retrieval from music scores images

ResearchDGX agent

arXiv:2605.22255v1 Announce Type: new Abstract: The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metad

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

SafetyDGX agent

arXiv:2605.22061v1 Announce Type: new Abstract: Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core ch

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

Model ReleasesDGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

Model ReleasesDGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

HardwareDGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

Model ReleasesDGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

← Previous
1…116117118119120…211
Next →