AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Local Ai

Towards Interpretable Foundation Models for Retinal Fundus Images

DGX agent

arXiv:2603.18846v2 Announce Type: replace Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL

local-aiarxiv-cs-cv
15 Apr 2026
Research
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

DGX agent

arXiv:2604.12309v1 Announce Type: new Abstract: We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generati

researcharxiv-cs-cv
15 Apr 2026
Tutorials

Ultra-low-light computer vision using trained photon correlations

DGX agent

arXiv:2604.11993v1 Announce Type: new Abstract: Illumination using correlated photon sources has been established as an approach to allowing high-fidelity images to be reconstructed from noisy camera

tutorialsarxiv-cs-cv
15 Apr 2026
Safety

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

DGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

safetyarxiv-cs-cv
15 Apr 2026
Research

UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators

DGX agent

arXiv:2604.11843v1 Announce Type: new Abstract: Invisible watermarking for autoregressive (AR) image generation has recently gained attention as a means of protecting image ownership and tracing AI-ge

researcharxiv-cs-cv
15 Apr 2026
Model Releases

Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization

DGX agent

arXiv:2604.12346v1 Announce Type: new Abstract: Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notorious

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos

DGX agent

arXiv:2604.11913v1 Announce Type: new Abstract: Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely

model-releasesarxiv-cs-cv
15 Apr 2026
Research

Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling

DGX agent

arXiv:2505.17384v2 Announce Type: replace-cross Abstract: Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a

researcharxiv-cs-cv
15 Apr 2026
Local Ai

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

DGX agent

arXiv:2604.12887v1 Announce Type: new Abstract: Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what

local-aiarxiv-cs-cv
15 Apr 2026
Research

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

DGX agent

arXiv:2604.12159v1 Announce Type: new Abstract: The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensi

researcharxiv-cs-cv
15 Apr 2026
Local Ai

ViLL-E: Video LLM Embeddings for Retrieval

DGX agent

arXiv:2604.12148v1 Announce Type: new Abstract: Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Caption

local-aiarxiv-cs-cv
15 Apr 2026
Research

Vision Transformers Need More Than Registers

DGX agent

arXiv:2602.22394v2 Announce Type: replace Abstract: Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, art

researcharxiv-cs-cv
15 Apr 2026
Research

Visual Diffusion Models are Geometric Solvers

DGX agent

arXiv:2510.21697v2 Announce Type: replace Abstract: In this paper we show that visual diffusion models can serve as effective geometric solvers: they can directly reason about geometric problems by wo

researcharxiv-cs-cv
15 Apr 2026
Local Ai

VPTracker: Global Vision-Language Tracking via Visual Prompt

DGX agent

arXiv:2512.22799v2 Announce Type: replace Abstract: Vision-Language Tracking aims to continuously localize objects described by a visual template and a language description. Existing methods, however,

local-aiarxiv-cs-cv
15 Apr 2026
Safety

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

DGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

safetyarxiv-cs-cv
15 Apr 2026
Research

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding

DGX agent

arXiv:2604.12358v1 Announce Type: new Abstract: Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that

researcharxiv-cs-cv
15 Apr 2026
Safety

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

DGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

safetyarxiv-cs-cv
14 Apr 2026
Research

3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis

DGX agent

arXiv:2604.11211v1 Announce Type: new Abstract: Real-time free-viewpoint rendering requires balancing multi-camera redundancy with the latency constraints of interactive applications. We address this

researcharxiv-cs-cv
14 Apr 2026
Model Releases

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

DGX agent

arXiv:2604.10456v1 Announce Type: new Abstract: The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems.

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

A Comparative Study of Modern Object Detectors for Robust Apple Detection in Orchard Imagery

DGX agent

arXiv:2604.09996v1 Announce Type: new Abstract: Accurate apple detection in orchard images is important for yield prediction, fruit counting, robotic harvesting, and crop monitoring. However, changing

model-releasesarxiv-cs-cv
14 Apr 2026
Research

A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches

DGX agent

arXiv:2604.10246v1 Announce Type: new Abstract: Photogrammetric 3D reconstruction has long relied on traditional Structure-from-Motion (SfM) and Multi-View Stereo (MVS) methods, which provide high acc

researcharxiv-cs-cv
14 Apr 2026
Tutorials

A Data-driven Loss Weighting Scheme across Heterogeneous Tasks for Image Denoising

DGX agent

arXiv:2301.06081v4 Announce Type: replace-cross Abstract: In a variational denoising model, weight in the data fidelity term plays the role of enhancing the noise-removal capability. It is profoundly

tutorialsarxiv-cs-cv
14 Apr 2026
Applications

A Deep Equilibrium Network for Hyperspectral Unmixing

DGX agent

arXiv:2604.11279v1 Announce Type: new Abstract: Hyperspectral unmixing (HU) is crucial for analyzing hyperspectral imagery, yet achieving accurate unmixing remains challenging. While traditional metho

applicationsarxiv-cs-cv
14 Apr 2026
Research

A Faster Path to Continual Learning

DGX agent

arXiv:2604.11064v1 Announce Type: cross Abstract: Continual Learning (CL) aims to train neural networks on a dynamic stream of tasks without forgetting previously learned knowledge. Among optimization

researcharxiv-cs-cv
14 Apr 2026
Local Ai

A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video

DGX agent

arXiv:2604.09685v1 Announce Type: new Abstract: We describe a zero-shot pipeline developed for the ACCIDENT @ CVPR 2026 challenge. The challenge requires predicting when, where, and what type of traff

local-aiarxiv-cs-cv
14 Apr 2026
Research

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

DGX agent

arXiv:2508.09977v4 Announce Type: replace Abstract: In the context of novel view synthesis, 3D Gaussian Splatting (3DGS) has recently emerged as an efficient and competitive counterpart to Neural Radi

researcharxiv-cs-cv
14 Apr 2026
Agents

A Survey on Deep Learning Techniques for Action Anticipation

DGX agent

arXiv:2309.17257v2 Announce Type: replace Abstract: The ability to anticipate possible future human actions is essential for a wide range of applications, including autonomous driving and human-robot

agentsarxiv-cs-cv
14 Apr 2026
Local Ai

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition

DGX agent

arXiv:2603.12221v2 Announce Type: replace Abstract: This paper addresses the expression (EXPR) recognition challenge in the 10th Affective Behavior Analysis in-the-Wild (ABAW) workshop and competition

local-aiarxiv-cs-cv
14 Apr 2026
Research

A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction

DGX agent

arXiv:2604.10210v1 Announce Type: new Abstract: Learning multi-scale representations is the common strategy to tackle object scale variation in dense prediction tasks. Although existing feature pyrami

researcharxiv-cs-cv
14 Apr 2026
Local Ai

ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents

DGX agent

arXiv:2604.10096v1 Announce Type: new Abstract: Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environmen

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

AC-MIL: Weakly Supervised Atrial LGE-MRI Quality Assessment via Adversarial Concept Disentanglement

DGX agent

arXiv:2604.10303v1 Announce Type: new Abstract: High-quality Late Gadolinium Enhancement (LGE) MRI can be helpful for atrial fibrillation management, yet scan quality is frequently compromised by pati

model-releasesarxiv-cs-cv
14 Apr 2026
Tutorials

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models

DGX agent

arXiv:2511.18082v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models have shown impressive flexibility and generalization, yet their deployment in robotic manipulation remain

tutorialsarxiv-cs-cv
14 Apr 2026
Safety

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

DGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

safetyarxiv-cs-cv
14 Apr 2026
Research

Adversarial Video Promotion Against Text-to-Video Retrieval

DGX agent

arXiv:2508.06964v3 Announce Type: replace Abstract: Thanks to the development of cross-modal models, text-to-video retrieval (T2VR) is advancing rapidly, but its robustness remains largely unexamined.

researcharxiv-cs-cv
14 Apr 2026
Research

Affostruction: 3D Affordance Grounding with Generative Reconstruction

DGX agent

arXiv:2601.09211v2 Announce Type: replace Abstract: This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a te

researcharxiv-cs-cv
14 Apr 2026
Safety

Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning

DGX agent

arXiv:2604.10383v1 Announce Type: new Abstract: Existing multi-agent video generation systems use LLM agents to orchestrate neural video generators, producing visually impressive but semantically unre

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

DGX agent

arXiv:2604.10454v1 Announce Type: new Abstract: Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

DGX agent

arXiv:2604.11730v1 Announce Type: new Abstract: Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits

researcharxiv-cs-cv
14 Apr 2026
Model Releases

AmodalSVG: Amodal Image Vectorization via Semantic Layer Peeling

DGX agent

arXiv:2604.10940v1 Announce Type: new Abstract: We introduce AmodalSVG, a new framework for amodal image vectorization that produces semantically organized and geometrically complete SVG representatio

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Analytical Modeling and Correction of Distance Error in Homography-Based Ground-Plane Mapping

DGX agent

arXiv:2604.10805v1 Announce Type: new Abstract: Accurate distance estimation from monocular cameras is essential for intelligent monitoring systems. In many deployments, image coordinates are mapped t

researcharxiv-cs-cv
14 Apr 2026
Research

Anatomy-Informed Deep Learning for Abdominal Aortic Aneurysm Segmentation

DGX agent

arXiv:2604.10312v1 Announce Type: new Abstract: In CT angiography, the accurate segmentation of abdominal aortic aneurysms (AAAs) is difficult due to large anatomical variability, low-contrast vessel

researcharxiv-cs-cv
14 Apr 2026
Research

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

DGX agent

arXiv:2604.11331v1 Announce Type: new Abstract: 3D scene generation has long been dominated by 2D multi-view or video diffusion models. This is due not only to the lack of scene-level 3D latent repres

researcharxiv-cs-cv
14 Apr 2026
Research

Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?

DGX agent

arXiv:2604.10217v1 Announce Type: new Abstract: Cross-modal optical-SAR (Synthetic Aperture Radar) registration is a bottleneck for disaster-response via remote sensing, yet modern image matchers are

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Are We Recognizing the Jaguar or Its Background? A Diagnostic Framework for Jaguar Re-Identification

DGX agent

arXiv:2604.09690v1 Announce Type: new Abstract: Jaguar re-identification (re-ID) from citizen-science imagery can look strong on standard retrieval metrics while still relying on the wrong evidence, s

model-releasesarxiv-cs-cv
14 Apr 2026
Agents

ArtiCAD: Articulated CAD Assembly Design via Multi-Agent Code Generation

DGX agent

arXiv:2604.10992v1 Announce Type: new Abstract: Parametric Computer-Aided Design (CAD) of articulated assemblies is essential for product development, yet generating these multi-part, movable models f

agentsarxiv-cs-cv
14 Apr 2026
Local Ai

At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections

DGX agent

arXiv:2604.10766v1 Announce Type: new Abstract: Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM cons

local-aiarxiv-cs-cv
14 Apr 2026
Research

Attention-Guided Dual-Stream Learning for Group Engagement Recognition: Fusing Transformer-Encoded Motion Dynamics with Scene Context via Adaptive Gating

DGX agent

arXiv:2604.10078v1 Announce Type: new Abstract: Student engagement is crucial for improving learning outcomes in group activities. Highly engaged students perform better both individually and contribu

researcharxiv-cs-cv
14 Apr 2026
Applications

Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation

DGX agent

arXiv:2511.15875v2 Announce Type: replace Abstract: The automated analysis of historical documents, particularly maps, has drastically benefited from advances in deep learning and its success across v

applicationsarxiv-cs-cv
14 Apr 2026
← Previous
1…242243244245246…259
Next →