AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Safety

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

DGX agent

arXiv:2604.09045v1 Announce Type: new Abstract: Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D s

safetyarxiv-cs-cv
13 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

SCoRe: Clean Image Generation from Diffusion Models Trained on Noisy Images

DGX agent

arXiv:2604.09436v1 Announce Type: new Abstract: Diffusion models trained on noisy datasets often reproduce high-frequency training artifacts, significantly degrading generation quality. To address thi

safetyarxiv-cs-cv
13 Apr 2026
Applications

SelfHVD: Self-Supervised Handheld Video Deblurring

DGX agent

arXiv:2508.08605v2 Announce Type: replace Abstract: Shooting video with handheld shooting devices often results in blurry frames due to shaking hands and other instability factors. Although previous v

applicationsarxiv-cs-cv
13 Apr 2026
Tutorials

ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding

DGX agent

arXiv:2512.03370v3 Announce Type: replace Abstract: We introduce ShelfGaussian, an open-vocabulary multi-modal Gaussian-based 3D scene understanding framework supervised by off-the-shelf vision founda

tutorialsarxiv-cs-cv
13 Apr 2026
Safety

SHIFT: Steering Hidden Intermediates in Flow Transformers

DGX agent

arXiv:2604.09213v1 Announce Type: new Abstract: Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong pro

safetyarxiv-cs-cv
13 Apr 2026
Research

SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation

DGX agent

arXiv:2604.08760v1 Announce Type: new Abstract: Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differe

researcharxiv-cs-cv
13 Apr 2026
Model Releases

SimScale: Learning to Drive via Real-World Simulation at Scale

DGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

model-releasesarxiv-cs-cv
13 Apr 2026
Research

State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition

DGX agent

arXiv:2604.08761v1 Announce Type: new Abstract: Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. E

researcharxiv-cs-cv
13 Apr 2026
Research

Streaming Video Instruction Tuning

DGX agent

arXiv:2512.21334v2 Announce Type: replace Abstract: We present Streamo, a real-time streaming video LLM that serves as a general-purpose interactive assistant. Unlike existing online video models that

researcharxiv-cs-cv
13 Apr 2026
Model Releases

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

DGX agent

arXiv:2604.09000v1 Announce Type: new Abstract: Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memo

model-releasesarxiv-cs-cv
13 Apr 2026
Research

Strips as Tokens: Artist Mesh Generation with Native UV Segmentation

DGX agent

arXiv:2604.09132v1 Announce Type: new Abstract: Recent advancements in autoregressive transformers have demonstrated remarkable potential for generating artist-quality meshes. However, the token order

researcharxiv-cs-cv
13 Apr 2026
Research

Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar Reconstruction

DGX agent

arXiv:2604.09324v1 Announce Type: new Abstract: Reconstructing photorealistic and topology-aware human avatars from monocular videos remains a significant challenge in the fields of computer vision an

researcharxiv-cs-cv
13 Apr 2026
Applications

SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data

DGX agent

arXiv:2604.09411v1 Announce Type: new Abstract: Reliable 3D dynamic perception requires models that can anticipate motion beyond predefined categories, yet progress is hindered by the scarcity of dens

applicationsarxiv-cs-cv
13 Apr 2026
Local Ai

TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction

DGX agent

arXiv:2604.08921v1 Announce Type: new Abstract: Accurate 3D human keypoints localization is a critical technology enabling robots to achieve natural and safe physical interaction with users. Conventio

local-aiarxiv-cs-cv
13 Apr 2026
Local Ai

Tango: Taming Visual Signals for Efficient Video Large Language Models

DGX agent

arXiv:2604.09547v1 Announce Type: new Abstract: Token pruning has emerged as a mainstream approach for developing efficient Video Large Language Models (Video LLMs). This work revisits and advances th

local-aiarxiv-cs-cv
13 Apr 2026
Model Releases

Text-Conditioned Multi-Expert Regression Framework for Fully Automated Multi-Abutment Design

DGX agent

arXiv:2604.09047v1 Announce Type: new Abstract: Dental implant abutments serve as the geometric and biomechanical interface between the implant fixture and the prosthetic crown, yet their design relie

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation

DGX agent

arXiv:2604.09368v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators per

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

TinyNeRV: Compact Neural Video Representations via Capacity Scaling, Distillation, and Low-Precision Inference

DGX agent

arXiv:2604.09220v1 Announce Type: new Abstract: Implicit neural video representations encode entire video sequences within the parameters of a neural network and enable constant time frame reconstruct

model-releasesarxiv-cs-cv
13 Apr 2026
Research

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

DGX agent

arXiv:2603.01400v2 Announce Type: replace Abstract: Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pru

researcharxiv-cs-cv
13 Apr 2026
Safety

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

DGX agent

arXiv:2604.09057v1 Announce Type: new Abstract: Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible moti

safetyarxiv-cs-cv
13 Apr 2026
Local Ai

TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches

DGX agent

arXiv:2604.08945v1 Announce Type: new Abstract: Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is

local-aiarxiv-cs-cv
13 Apr 2026
Model Releases

Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments

DGX agent

arXiv:2604.09038v1 Announce Type: cross Abstract: Robust geo-localization in changing environmental conditions is critical for long-term aerial autonomy. While visual place recognition (VPR) models pe

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

DGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

safetyarxiv-cs-cv
13 Apr 2026
Tutorials

Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models

DGX agent

arXiv:2604.09227v1 Announce Type: cross Abstract: Image generative models have become indispensable tools to yield exquisite high-resolution (HR) images for everyone, ranging from general users to pro

tutorialsarxiv-cs-cv
13 Apr 2026
Model Releases

TurPy: a physics-based and differentiable optical turbulence simulator for algorithmic development and system optimization

DGX agent

arXiv:2604.07248v2 Announce Type: replace-cross Abstract: Developing optical systems for free-space applications requires simulation tools that accurately capture turbulence-induced wavefront distorti

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models

DGX agent

arXiv:2604.02241v2 Announce Type: replace Abstract: Embodied visual tracking is crucial for Unmanned Aerial Vehicles (UAVs) executing complex real-world tasks. In dynamic urban scenarios with complex

model-releasesarxiv-cs-cv
13 Apr 2026
Research

UHD Low-Light Image Enhancement via Real-Time Enhancement Methods with Clifford Information Fusion

DGX agent

arXiv:2604.09321v1 Announce Type: cross Abstract: Considering efficiency, ultra-high-definition (UHD) low-light image restoration is extremely challenging. Existing methods based on Transformer archit

researcharxiv-cs-cv
13 Apr 2026
Model Releases

Unified Multimodal Uncertain Inference

DGX agent

arXiv:2604.08701v1 Announce Type: new Abstract: We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calib

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation

DGX agent

arXiv:2604.09169v1 Announce Type: new Abstract: Semi-supervised semantic segmentation in computational pathology remains challenging due to scarce pixel-level annotations and unreliable pseudo-label s

safetyarxiv-cs-cv
13 Apr 2026
Safety

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

DGX agent

arXiv:2604.09330v1 Announce Type: cross Abstract: Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-wo

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

VAGNet: Vision-based accident anticipation with global features

DGX agent

arXiv:2604.09305v1 Announce Type: new Abstract: Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

ViSAGE @ NTIRE 2026 Challenge on Video Saliency Prediction

DGX agent

arXiv:2604.08613v1 Announce Type: new Abstract: In this report, we present our champion solution for the NTIRE 2026 Challenge on Video Saliency Prediction held in conjunction with CVPR 2026. To exploi

model-releasesarxiv-cs-cv
13 Apr 2026
Applications

VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization

DGX agent

arXiv:2508.13792v2 Announce Type: replace Abstract: The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible intera

applicationsarxiv-cs-cv
13 Apr 2026
Tutorials

What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction

DGX agent

arXiv:2604.08716v1 Announce Type: new Abstract: Virtual Try-On (VTON) has seen rapid advancements, providing a strong foundation for generative fashion tasks. However, the inverse problem, Virtual Try

tutorialsarxiv-cs-cv
13 Apr 2026
Safety

When & How to Write for Personalized Demand-aware Query Rewriting in Video Search

DGX agent

arXiv:2602.17667v2 Announce Type: replace-cross Abstract: In video search systems, user historical behaviors provide rich context for identifying search intent and resolving ambiguity. However, tradit

safetyarxiv-cs-cv
13 Apr 2026
Applications

WildDet3D: Scaling Promptable 3D Detection in the Wild

DGX agent

arXiv:2604.08626v1 Announce Type: new Abstract: Understanding objects in 3D from a single image is a cornerstone of spatial intelligence. A key step toward this goal is monocular 3D object detection--

applicationsarxiv-cs-cv
13 Apr 2026
Research

Zero-Shot Generative De-identification: Inversion-Free Flow for Privacy-Preserving Skin Image Analysis

DGX agent

arXiv:2602.00821v2 Announce Type: replace Abstract: The secure analysis of dermatological images in clinical environments is fundamentally restricted by the critical trade-off between patient privacy

researcharxiv-cs-cv
13 Apr 2026
Model Releases

3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience

DGX agent

arXiv:2604.08042v1 Announce Type: new Abstract: Sketching in 3D space enables expressive reasoning about shape, structure, and spatial relationships, yet generating 3D sketches through natural languag

model-releasesarxiv-cs-cv
10 Apr 2026
Research

A Geometric Algorithm for Blood Vessel Reconstruction from Skeletal Representation

DGX agent

arXiv:2402.12797v4 Announce Type: replace Abstract: We introduce a novel approach for the reconstruction of tubular shapes from skeletal representations. Our method processes all skeletal points as a

researcharxiv-cs-cv
10 Apr 2026
Safety

A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring

DGX agent

arXiv:2604.07395v1 Announce Type: cross Abstract: Robotic manipulation systems that follow language instructions often execute grasp primitives in a largely single-shot manner: a model proposes an act

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification

DGX agent

arXiv:2510.04628v2 Announce Type: replace Abstract: Deep learning-based methods have achieved significant success in remote sensing Earth observation data analysis. Numerous feature fusion techniques

model-releasesarxiv-cs-cv
10 Apr 2026
Research

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

DGX agent

arXiv:2604.08050v1 Announce Type: new Abstract: In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging

researcharxiv-cs-cv
10 Apr 2026
Agents

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

DGX agent

arXiv:2604.08545v1 Announce Type: new Abstract: The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a pro

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

DGX agent

arXiv:2511.20162v2 Announce Type: replace Abstract: Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Adapting Foundation Models for Annotation-Efficient Adnexal Mass Segmentation in Cine Images

DGX agent

arXiv:2604.08045v1 Announce Type: new Abstract: Adnexal mass evaluation via ultrasound is a challenging clinical task, often hindered by subjective interpretation and significant inter-observer variab

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Adaptive Depth-converted-Scale Convolution for Self-supervised Monocular Depth Estimation

DGX agent

arXiv:2604.07665v1 Announce Type: new Abstract: Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object

model-releasesarxiv-cs-cv
10 Apr 2026
Research

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

DGX agent

arXiv:2604.08077v1 Announce Type: new Abstract: Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fi

researcharxiv-cs-cv
10 Apr 2026
Research

Adversarial Evasion Attacks on Computer Vision using SHAP Values

DGX agent

arXiv:2601.10587v2 Announce Type: replace Abstract: The paper introduces a white-box attack on computer vision models using SHAP values. It demonstrates how adversarial evasion attacks can compromise

researcharxiv-cs-cv
10 Apr 2026
← Previous
1…252253254255256…259
Next →