AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Where to Refine, When to Stop: Rethinking Redundancy via Latent Discrepancy for Efficient Visual Autoregressive Generation

DGX agent

arXiv:2606.00310v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models deliver high-quality image generation but suffer from significant inference latency at high resolutions. Recent accel

researcharxiv-cs-cv
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata

DGX agent

arXiv:2602.12819v2 Announce Type: replace-cross Abstract: In this paper, we present WISE, an open-source audiovisual search engine which integrates a range of multimodal retrieval capabilities into a

applicationsarxiv-cs-cv
2 Jun 2026
Model Releases

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

DGX agent

arXiv:2603.06331v2 Announce Type: replace Abstract: Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactiv

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

DGX agent

arXiv:2512.10958v2 Announce Type: replace Abstract: Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fa

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

DGX agent

arXiv:2605.24892v2 Announce Type: replace Abstract: Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and gen

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

DGX agent

arXiv:2606.02482v1 Announce Type: new Abstract: While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and mul

model-releasesarxiv-cs-cv
2 Jun 2026
Agents

XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation

DGX agent

arXiv:2510.13565v2 Announce Type: replace Abstract: Depth estimation remains central to autonomous driving, and radar-camera fusion offers robustness in adverse conditions by providing complementary g

agentsarxiv-cs-cv
2 Jun 2026
Research

Zero-Shot Multi-Animal Tracking in the Wild

DGX agent

arXiv:2511.02591v2 Announce Type: replace Abstract: Multi-animal tracking is crucial for understanding animal ecology and behavior, yet remains challenging due to variations in habitat, motion pattern

researcharxiv-cs-cv
2 Jun 2026
Model Releases

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

DGX agent

arXiv:2605.30469v1 Announce Type: cross Abstract: 3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural

model-releasesarxiv-cs-cv
1 Jun 2026
Research

A Context-Aware Middleware for Medical Image Based Reports: An approach based on image feature extraction and association rules

DGX agent

arXiv:2605.30699v1 Announce Type: cross Abstract: This work proposes a context-aware middleware for medical workflow organization and efficiency improvement. In hospitals, laboratories and teleradiolo

researcharxiv-cs-cv
1 Jun 2026
Model Releases

A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss

DGX agent

arXiv:2509.10114v2 Announce Type: replace Abstract: Face image quality assessment (FIQA) plays a critical role in face recognition and verification systems, especially in uncontrolled, real-world envi

model-releasesarxiv-cs-cv
1 Jun 2026
Research

A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications

DGX agent

arXiv:2601.22202v2 Announce Type: replace-cross Abstract: Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw d

researcharxiv-cs-cv
1 Jun 2026
Research

A Unifying View of Variational Generative Wasserstein Flows

DGX agent

arXiv:2605.31369v1 Announce Type: cross Abstract: Many modern generative models can be viewed as minimizing divergences between probability distributions, yet they rely on different algorithmic and ge

researcharxiv-cs-cv
1 Jun 2026
Applications

AdvScene: Rethinking Adversarial Patch Evaluation Through Scene Robustness

DGX agent

arXiv:2605.30578v1 Announce Type: cross Abstract: Adversarial patches are physical patterns attached to real objects to mislead AI vision systems. Their real-world risk is not determined by a single s

applicationsarxiv-cs-cv
1 Jun 2026
Safety

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

DGX agent

arXiv:2605.30742v1 Announce Type: new Abstract: This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topi

safetyarxiv-cs-cv
1 Jun 2026
Applications

Astra: a generalizable report generation foundation model for 3D computed tomography

DGX agent

arXiv:2605.31437v1 Announce Type: new Abstract: CT interpretation requires radiologists to review hundreds of volumetric slices per examination, making reporting time-consuming and highly expertise-de

applicationsarxiv-cs-cv
1 Jun 2026
Research

Authentication of Copy Detection Patterns via Cross-Camera Dual-Synthetic Referencing

DGX agent

arXiv:2605.31292v1 Announce Type: new Abstract: Copy Detection Patterns (CDPs) are structures printed on physical objects to enable cost-effective authentication. Verification is achieved by comparing

researcharxiv-cs-cv
1 Jun 2026
Model Releases

Automated Prediction of Postoperative Pancreatic Fistula Using Preoperative Computed Tomography

DGX agent

arXiv:2605.31539v1 Announce Type: new Abstract: Postoperative pancreatic fistula (POPF) is a serious complication after pancreatic resection, increasing morbidity, hospital stay, and healthcare costs.

model-releasesarxiv-cs-cv
1 Jun 2026
Research

BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation

DGX agent

arXiv:2511.19394v2 Announce Type: replace Abstract: Segmenting small lesions in medical images remains notoriously difficult. Most prior work tackles this challenge by either designing better architec

researcharxiv-cs-cv
1 Jun 2026
Research

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

DGX agent

arXiv:2605.31246v1 Announce Type: cross Abstract: Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing a

researcharxiv-cs-cv
1 Jun 2026
Applications

Benchmarking Single-Step Inpainting Methods for Multi-Object 3D Gaussian Splatting Scenes

DGX agent

arXiv:2605.30987v1 Announce Type: new Abstract: The tasks of object removal and inpainting 3D Gaussian Splatting (3DGS) scenes face challenges such as 3D consistency across camera views. In comparing

applicationsarxiv-cs-cv
1 Jun 2026
Local Ai

Beyond Accuracy: Evaluating Efficiency, Robustness and Explainability in Deep Learning for Malaria Diagnosis

DGX agent

arXiv:2605.30734v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in sub-Saharan Africa, where scarce diagnostic infrastructure makes timely, accurate diagnosis particular

local-aiarxiv-cs-cv
1 Jun 2026
Safety

BIAS-ID: A Framework for Analyzing Transformation Biases in AI-Generated Image Detectors

DGX agent

arXiv:2605.31153v1 Announce Type: new Abstract: Given the surge of harmful AI-generated imagery online, reliably distinguishing authentic images from generated ones has become an urgent research topic

safetyarxiv-cs-cv
1 Jun 2026
Safety

BiSegMamba: Efficient Bidirectional Tri-Oriented Mamba for 3D Medical Image Segmentation

DGX agent

arXiv:2605.30972v1 Announce Type: new Abstract: Accurate 3D medical image segmentation requires both long-range volumetric context and fine boundary preservation. CNN-based methods have limited global

safetyarxiv-cs-cv
1 Jun 2026
Safety

Calibrated Uncertainty for Trustworthy Clinical Gait Analysis Using Probabilistic Multiview Markerless Motion Capture

DGX agent

arXiv:2601.22412v2 Announce Type: replace Abstract: Video-based human movement analysis holds potential for movement assessment in clinical practice and research. However, the clinical implementation

safetyarxiv-cs-cv
1 Jun 2026
Research

CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

DGX agent

arXiv:2605.30774v1 Announce Type: new Abstract: Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly i

researcharxiv-cs-cv
1 Jun 2026
Agents

Can BEV Perception Gracefully Degrade under Sensor Failures?

DGX agent

arXiv:2605.30983v1 Announce Type: new Abstract: Despite the remarkable success of multi-modal bird's-eye view (BEV) perception in autonomous driving, current systems exhibit a critical vulnerability:

agentsarxiv-cs-cv
1 Jun 2026
Research

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation

DGX agent

arXiv:2510.22067v3 Announce Type: replace Abstract: Vision language models (VLMs) often generate hallucination, i.e., content that cannot be substantiated by either textual or visual inputs. Prior wor

researcharxiv-cs-cv
1 Jun 2026
Research

Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing

DGX agent

arXiv:2605.30467v1 Announce Type: new Abstract: This study introduces a novel Arctic-focused remote sensing foundation model (RSFM) by combining diversity-aware regional-scale image curation with mask

researcharxiv-cs-cv
1 Jun 2026
Model Releases

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

DGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Count Anything

DGX agent

arXiv:2605.30846v1 Announce Type: new Abstract: Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing c

model-releasesarxiv-cs-cv
1 Jun 2026
Applications

Cross-Modal Clinical Knowledge Integration for Mammography Report Generation

DGX agent

arXiv:2605.31093v1 Announce Type: new Abstract: Breast cancer is a major global health concern, and mammography screening plays a central role in early detection. The large volume of screening examina

applicationsarxiv-cs-cv
1 Jun 2026
Research

D-SECURE: Dual-Source Evidence Combination for Unified Reasoning in Misinformation Detection

DGX agent

arXiv:2602.14441v2 Announce Type: replace Abstract: Multimodal misinformation increasingly mixes realistic im-age edits with fluent but misleading text, producing persuasive posts that are difficult t

researcharxiv-cs-cv
1 Jun 2026
Research

DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring

DGX agent

arXiv:2509.18898v2 Announce Type: replace Abstract: In this paper, we propose the first Structure-from-Motion (SfM)-free deblurring 3D Gaussian Splatting method via event camera, dubbed DeblurSplat. W

researcharxiv-cs-cv
1 Jun 2026
Research

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

DGX agent

arXiv:2605.31336v1 Announce Type: new Abstract: Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal

researcharxiv-cs-cv
1 Jun 2026
Safety

Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery

DGX agent

arXiv:2509.20941v2 Announce Type: replace Abstract: As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations nece

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

DGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

DGX agent

arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation

DGX agent

arXiv:2605.30444v1 Announce Type: new Abstract: Recent advances in 4D Human-Object Interaction (HOI) generation have enabled increasingly realistic motion synthesis, particularly for single-object man

model-releasesarxiv-cs-cv
1 Jun 2026
Research

DisPlace: Discriminative Place Projections for Multi-Reference Visual Place Recognition

DGX agent

arXiv:2605.30769v1 Announce Type: new Abstract: A key challenge in Visual Place Recognition (VPR) is matching query images against reference maps captured under diverse environmental conditions and vi

researcharxiv-cs-cv
1 Jun 2026
Safety

DiTTo: Scalable Order-aware All-in-One Image Restoration Agent

DGX agent

arXiv:2605.30915v1 Announce Type: new Abstract: Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration q

safetyarxiv-cs-cv
1 Jun 2026
Research

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

DGX agent

arXiv:2605.30713v1 Announce Type: cross Abstract: Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their applicati

researcharxiv-cs-cv
1 Jun 2026
Model Releases

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

DGX agent

arXiv:2605.31271v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise

model-releasesarxiv-cs-cv
1 Jun 2026
Hardware

DSD-GS: Dynamic-Static Decomposition of Gaussian Splatting for Efficient and High-Fidelity Dynamic Scene Reconstruction

DGX agent

arXiv:2605.30863v1 Announce Type: new Abstract: Dynamic scene reconstruction and novel view synthesis are fundamental to next-generation visual intelligence applications such as virtual reality, robot

hardwarearxiv-cs-cv
1 Jun 2026
Model Releases

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

DGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

model-releasesarxiv-cs-cv
1 Jun 2026
Research

EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment

DGX agent

arXiv:2602.01173v3 Announce Type: replace Abstract: Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowerin

researcharxiv-cs-cv
1 Jun 2026
Model Releases

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision

DGX agent

arXiv:2605.31557v1 Announce Type: new Abstract: Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchma

model-releasesarxiv-cs-cv
1 Jun 2026
Hardware

Elastic ViTs from Pretrained Models without Retraining

DGX agent

arXiv:2510.17700v2 Announce Type: replace Abstract: Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deploym

hardwarearxiv-cs-cv
1 Jun 2026
← Previous
1…131132133134135…263
Next →