AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

Task Alignment: A simple and effective proxy for model merging in computer vision

DGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

safetyarxiv-cs-cv
15 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

DGX agent

arXiv:2512.03963v3 Announce Type: replace Abstract: Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks

local-aiarxiv-cs-cv
15 Apr 2026
Local Ai

Time-reversed Flow Matching with Worst Transport in High-dimensional Latent Space for Image Anomaly Detection

DGX agent

arXiv:2508.05461v3 Announce Type: replace Abstract: Likelihood-based deep generative models have been widely investigated for Image Anomaly Detection (IAD), particularly Normalizing Flows, yet their s

local-aiarxiv-cs-cv
15 Apr 2026
Model Releases

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

DGX agent

arXiv:2604.12012v1 Announce Type: new Abstract: Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classificat

model-releasesarxiv-cs-cv
15 Apr 2026
Local Ai

Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation

DGX agent

arXiv:2512.05812v5 Announce Type: replace-cross Abstract: Scalable multi-agent driving simulation requires behavior models that are both realistic and computationally efficient. We address this by opt

local-aiarxiv-cs-cv
15 Apr 2026
Local Ai

Towards Interpretable Foundation Models for Retinal Fundus Images

DGX agent

arXiv:2603.18846v2 Announce Type: replace Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL

local-aiarxiv-cs-cv
15 Apr 2026
Research

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

DGX agent

arXiv:2604.12309v1 Announce Type: new Abstract: We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generati

researcharxiv-cs-cv
15 Apr 2026
Tutorials

Ultra-low-light computer vision using trained photon correlations

DGX agent

arXiv:2604.11993v1 Announce Type: new Abstract: Illumination using correlated photon sources has been established as an approach to allowing high-fidelity images to be reconstructed from noisy camera

tutorialsarxiv-cs-cv
15 Apr 2026
Safety

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

DGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

safetyarxiv-cs-cv
15 Apr 2026
Research

UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators

DGX agent

arXiv:2604.11843v1 Announce Type: new Abstract: Invisible watermarking for autoregressive (AR) image generation has recently gained attention as a means of protecting image ownership and tracing AI-ge

researcharxiv-cs-cv
15 Apr 2026
Model Releases

Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization

DGX agent

arXiv:2604.12346v1 Announce Type: new Abstract: Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notorious

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos

DGX agent

arXiv:2604.11913v1 Announce Type: new Abstract: Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely

model-releasesarxiv-cs-cv
15 Apr 2026
Research

Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling

DGX agent

arXiv:2505.17384v2 Announce Type: replace-cross Abstract: Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a

researcharxiv-cs-cv
15 Apr 2026
Local Ai

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

DGX agent

arXiv:2604.12887v1 Announce Type: new Abstract: Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what

local-aiarxiv-cs-cv
15 Apr 2026
Research

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

DGX agent

arXiv:2604.12159v1 Announce Type: new Abstract: The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensi

researcharxiv-cs-cv
15 Apr 2026
Local Ai

ViLL-E: Video LLM Embeddings for Retrieval

DGX agent

arXiv:2604.12148v1 Announce Type: new Abstract: Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Caption

local-aiarxiv-cs-cv
15 Apr 2026
Research

Vision Transformers Need More Than Registers

DGX agent

arXiv:2602.22394v2 Announce Type: replace Abstract: Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, art

researcharxiv-cs-cv
15 Apr 2026
Research

Visual Diffusion Models are Geometric Solvers

DGX agent

arXiv:2510.21697v2 Announce Type: replace Abstract: In this paper we show that visual diffusion models can serve as effective geometric solvers: they can directly reason about geometric problems by wo

researcharxiv-cs-cv
15 Apr 2026
Local Ai

VPTracker: Global Vision-Language Tracking via Visual Prompt

DGX agent

arXiv:2512.22799v2 Announce Type: replace Abstract: Vision-Language Tracking aims to continuously localize objects described by a visual template and a language description. Existing methods, however,

local-aiarxiv-cs-cv
15 Apr 2026
Safety

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

DGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

safetyarxiv-cs-cv
15 Apr 2026
Research

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding

DGX agent

arXiv:2604.12358v1 Announce Type: new Abstract: Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that

researcharxiv-cs-cv
15 Apr 2026
Safety

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

DGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

safetyarxiv-cs-cv
14 Apr 2026
Research

3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis

DGX agent

arXiv:2604.11211v1 Announce Type: new Abstract: Real-time free-viewpoint rendering requires balancing multi-camera redundancy with the latency constraints of interactive applications. We address this

researcharxiv-cs-cv
14 Apr 2026
Model Releases

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

DGX agent

arXiv:2604.10456v1 Announce Type: new Abstract: The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems.

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

A Comparative Study of Modern Object Detectors for Robust Apple Detection in Orchard Imagery

DGX agent

arXiv:2604.09996v1 Announce Type: new Abstract: Accurate apple detection in orchard images is important for yield prediction, fruit counting, robotic harvesting, and crop monitoring. However, changing

model-releasesarxiv-cs-cv
14 Apr 2026
Research

A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches

DGX agent

arXiv:2604.10246v1 Announce Type: new Abstract: Photogrammetric 3D reconstruction has long relied on traditional Structure-from-Motion (SfM) and Multi-View Stereo (MVS) methods, which provide high acc

researcharxiv-cs-cv
14 Apr 2026
Tutorials

A Data-driven Loss Weighting Scheme across Heterogeneous Tasks for Image Denoising

DGX agent

arXiv:2301.06081v4 Announce Type: replace-cross Abstract: In a variational denoising model, weight in the data fidelity term plays the role of enhancing the noise-removal capability. It is profoundly

tutorialsarxiv-cs-cv
14 Apr 2026
Applications

A Deep Equilibrium Network for Hyperspectral Unmixing

DGX agent

arXiv:2604.11279v1 Announce Type: new Abstract: Hyperspectral unmixing (HU) is crucial for analyzing hyperspectral imagery, yet achieving accurate unmixing remains challenging. While traditional metho

applicationsarxiv-cs-cv
14 Apr 2026
Research

A Faster Path to Continual Learning

DGX agent

arXiv:2604.11064v1 Announce Type: cross Abstract: Continual Learning (CL) aims to train neural networks on a dynamic stream of tasks without forgetting previously learned knowledge. Among optimization

researcharxiv-cs-cv
14 Apr 2026
Local Ai

A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video

DGX agent

arXiv:2604.09685v1 Announce Type: new Abstract: We describe a zero-shot pipeline developed for the ACCIDENT @ CVPR 2026 challenge. The challenge requires predicting when, where, and what type of traff

local-aiarxiv-cs-cv
14 Apr 2026
Research

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

DGX agent

arXiv:2508.09977v4 Announce Type: replace Abstract: In the context of novel view synthesis, 3D Gaussian Splatting (3DGS) has recently emerged as an efficient and competitive counterpart to Neural Radi

researcharxiv-cs-cv
14 Apr 2026
Agents

A Survey on Deep Learning Techniques for Action Anticipation

DGX agent

arXiv:2309.17257v2 Announce Type: replace Abstract: The ability to anticipate possible future human actions is essential for a wide range of applications, including autonomous driving and human-robot

agentsarxiv-cs-cv
14 Apr 2026
Local Ai

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition

DGX agent

arXiv:2603.12221v2 Announce Type: replace Abstract: This paper addresses the expression (EXPR) recognition challenge in the 10th Affective Behavior Analysis in-the-Wild (ABAW) workshop and competition

local-aiarxiv-cs-cv
14 Apr 2026
Research

A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction

DGX agent

arXiv:2604.10210v1 Announce Type: new Abstract: Learning multi-scale representations is the common strategy to tackle object scale variation in dense prediction tasks. Although existing feature pyrami

researcharxiv-cs-cv
14 Apr 2026
Local Ai

ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents

DGX agent

arXiv:2604.10096v1 Announce Type: new Abstract: Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environmen

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

AC-MIL: Weakly Supervised Atrial LGE-MRI Quality Assessment via Adversarial Concept Disentanglement

DGX agent

arXiv:2604.10303v1 Announce Type: new Abstract: High-quality Late Gadolinium Enhancement (LGE) MRI can be helpful for atrial fibrillation management, yet scan quality is frequently compromised by pati

model-releasesarxiv-cs-cv
14 Apr 2026
Tutorials

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models

DGX agent

arXiv:2511.18082v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models have shown impressive flexibility and generalization, yet their deployment in robotic manipulation remain

tutorialsarxiv-cs-cv
14 Apr 2026
Safety

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

DGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

safetyarxiv-cs-cv
14 Apr 2026
Research

Adversarial Video Promotion Against Text-to-Video Retrieval

DGX agent

arXiv:2508.06964v3 Announce Type: replace Abstract: Thanks to the development of cross-modal models, text-to-video retrieval (T2VR) is advancing rapidly, but its robustness remains largely unexamined.

researcharxiv-cs-cv
14 Apr 2026
Research

Affostruction: 3D Affordance Grounding with Generative Reconstruction

DGX agent

arXiv:2601.09211v2 Announce Type: replace Abstract: This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a te

researcharxiv-cs-cv
14 Apr 2026
Safety

Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning

DGX agent

arXiv:2604.10383v1 Announce Type: new Abstract: Existing multi-agent video generation systems use LLM agents to orchestrate neural video generators, producing visually impressive but semantically unre

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

DGX agent

arXiv:2604.10454v1 Announce Type: new Abstract: Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

DGX agent

arXiv:2604.11730v1 Announce Type: new Abstract: Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits

researcharxiv-cs-cv
14 Apr 2026
Model Releases

AmodalSVG: Amodal Image Vectorization via Semantic Layer Peeling

DGX agent

arXiv:2604.10940v1 Announce Type: new Abstract: We introduce AmodalSVG, a new framework for amodal image vectorization that produces semantically organized and geometrically complete SVG representatio

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Analytical Modeling and Correction of Distance Error in Homography-Based Ground-Plane Mapping

DGX agent

arXiv:2604.10805v1 Announce Type: new Abstract: Accurate distance estimation from monocular cameras is essential for intelligent monitoring systems. In many deployments, image coordinates are mapped t

researcharxiv-cs-cv
14 Apr 2026
Research

Anatomy-Informed Deep Learning for Abdominal Aortic Aneurysm Segmentation

DGX agent

arXiv:2604.10312v1 Announce Type: new Abstract: In CT angiography, the accurate segmentation of abdominal aortic aneurysms (AAAs) is difficult due to large anatomical variability, low-contrast vessel

researcharxiv-cs-cv
14 Apr 2026
Research

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

DGX agent

arXiv:2604.11331v1 Announce Type: new Abstract: 3D scene generation has long been dominated by 2D multi-view or video diffusion models. This is due not only to the lack of scene-level 3D latent repres

researcharxiv-cs-cv
14 Apr 2026
Research

Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?

DGX agent

arXiv:2604.10217v1 Announce Type: new Abstract: Cross-modal optical-SAR (Synthetic Aperture Radar) registration is a bottleneck for disaster-response via remote sensing, yet modern image matchers are

researcharxiv-cs-cv
14 Apr 2026
← Previous
1…244245246247248…261
Next →