AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Applications

FASTER: Rethinking Real-Time Flow VLAs

DGX agent

arXiv:2603.19199v2 Announce Type: replace-cross Abstract: Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference method

applicationsarxiv-cs-cv
30 Apr 2026
Research

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.26488v1 Announce Type: new Abstract: One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack r

researcharxiv-cs-cv
30 Apr 2026
Research

Federated Medical Image Classification under Class and Domain Imbalance exploiting Synthetic Sample Generation

DGX agent

arXiv:2604.26324v1 Announce Type: new Abstract: Exploiting deep learning in medical imaging faces critical challenges, including strict privacy constraints, heterogeneous imaging devices with varying

researcharxiv-cs-cv
30 Apr 2026
Safety

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding

DGX agent

arXiv:2504.09925v3 Announce Type: replace Abstract: We introduce FLARE, a family of vision language models (VLMs) with a fully vision-language alignment and integration paradigm. Unlike existing appro

safetyarxiv-cs-cv
30 Apr 2026
Research

Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery

DGX agent

arXiv:2602.13780v2 Announce Type: replace Abstract: Remote sensing (RS) change detection is essential for interpreting surface dynamics. Semantic change detection (SCD) further enables pixel-level und

researcharxiv-cs-cv
30 Apr 2026
Research

FunFace: Feature Utility and Norm Estimation for Face Recognition

DGX agent

arXiv:2604.26598v1 Announce Type: new Abstract: Face Recognition (FR) is used in a variety of application domains, from entertainment and banking to security and surveillance. Such applications rely o

researcharxiv-cs-cv
30 Apr 2026
Applications

GaitKD: A Universal Decoupled Distillation Framework for Efficient Gait Recognition

DGX agent

arXiv:2604.26255v1 Announce Type: new Abstract: Gait recognition is an attractive biometric modality for long-range and contact-free identification, but high-performing gait models often rely on deep

applicationsarxiv-cs-cv
30 Apr 2026
Research

GateMOT: Q-Gated Attention for Dense Object Tracking

DGX agent

arXiv:2604.26353v1 Announce Type: new Abstract: While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tr

researcharxiv-cs-cv
30 Apr 2026
Tutorials

Generalized Disguise Makeup Presentation Attack Detection Using an Attention-Guided Patch-Based Framework

DGX agent

arXiv:2604.26025v1 Announce Type: new Abstract: Despite significant advances in facial recognition systems, they remain vulnerable to face presentation attacks. Among them, disguise makeup attacks are

tutorialsarxiv-cs-cv
30 Apr 2026
Model Releases

GIFGuard: Proactive Forensics against Deepfakes in Facial GIFs via Spatiotemporal Watermarking

DGX agent

arXiv:2604.26519v1 Announce Type: new Abstract: The rapid evolution of deepfake technology poses an unprecedented threat to the authenticity of Graphics Interchange Format (GIF) imagery, which serves

model-releasesarxiv-cs-cv
30 Apr 2026
Agents

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

DGX agent

arXiv:2604.26752v1 Announce Type: new Abstract: We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environmen

agentsarxiv-cs-cv
30 Apr 2026
Safety

GNC-Pose: Geometry-Aware GNC-PnP for Accurate 6D Pose Estimation

DGX agent

arXiv:2512.06565v2 Announce Type: replace Abstract: We present GNC-Pose, a fully learning-free monocular 6D object pose estimation pipeline for textured objects that combines rendering-based initializ

safetyarxiv-cs-cv
30 Apr 2026
Model Releases

Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark

DGX agent

arXiv:2604.26893v1 Announce Type: new Abstract: Fine-grained RGBT image semantic segmentation is crucial for all-weather unmanned aerial vehicle (UAV) scene understanding. However, UAV RGBT semantic s

model-releasesarxiv-cs-cv
30 Apr 2026
Research

Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations

DGX agent

arXiv:2604.26678v1 Announce Type: new Abstract: Optical vibration sensing enables recovering the scene sound directly from the surface vibration of nearby objects, turning everyday objects into ``visu

researcharxiv-cs-cv
30 Apr 2026
Applications

High-Dimensional Noise to Low-Dimensional Manifolds: A Manifold-Space Diffusion Framework for Degraded Hyperspectral Image Classification

DGX agent

arXiv:2604.26279v1 Announce Type: new Abstract: Recently, Hyperspectral Image (HSI) classification has attracted increasing attention in remote sensing. However, HSI data are inherently high-dimension

applicationsarxiv-cs-cv
30 Apr 2026
Tutorials

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

DGX agent

arXiv:2604.26227v1 Announce Type: new Abstract: In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed netw

tutorialsarxiv-cs-cv
30 Apr 2026
Model Releases

HumanOmni-Speaker: Identifying Who said What and When

DGX agent

arXiv:2603.21664v2 Announce Type: replace Abstract: While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human intera

model-releasesarxiv-cs-cv
30 Apr 2026
Research

KAYRA: A Microservice Architecture for AI-Assisted Karyotyping with Cloud and On-Premise Deployment

DGX agent

arXiv:2604.26869v1 Announce Type: cross Abstract: We present KAYRA, an end-to-end karyotyping system that operates inside the operational constraints of a clinical cytogenetic laboratory. KAYRA is arc

researcharxiv-cs-cv
30 Apr 2026
Research

Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation

DGX agent

arXiv:2604.26454v1 Announce Type: new Abstract: Monocular depth estimation (MDE) is a fundamental yet inherently ill-posed task. Recent vision foundation models (VFMs), particularly DINO-based transfo

researcharxiv-cs-cv
30 Apr 2026
Tutorials

Learning Sparse BRDF Measurement Samples from Image

DGX agent

arXiv:2604.26740v1 Announce Type: new Abstract: Accurate BRDF acquisition is important for realistic rendering, but dense gonioreflectometer measurements are slow and expensive. We study how to select

tutorialsarxiv-cs-cv
30 Apr 2026
Safety

Learning Vision-Based Omnidirectional Navigation: A Teacher-Student Approach Using Monocular Depth Estimation

DGX agent

arXiv:2603.01999v2 Announce Type: replace-cross Abstract: Reliable obstacle avoidance in industrial settings demands 3D scene understanding, but widely used 2D LiDAR sensors perceive only a single hor

safetyarxiv-cs-cv
30 Apr 2026
Hardware

MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching

DGX agent

arXiv:2604.26799v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) achieves high-quality novel view synthesis with real-time rendering, but its storage cost remains prohibitive for practical

hardwarearxiv-cs-cv
30 Apr 2026
Model Releases

MixerCA: An Efficient and Accurate Model for High-Performance Hyperspectral Image Classification

DGX agent

arXiv:2604.26138v1 Announce Type: new Abstract: Over the past decade, hyperspectral image (HSI) classification has drawn considerable interest due to HSIs' ability to effectively distinguish terrestri

model-releasesarxiv-cs-cv
30 Apr 2026
Research

Motion-Driven Multi-Object Tracking of Model Organisms in Space Science Experiments

DGX agent

arXiv:2604.26321v1 Announce Type: new Abstract: Automated animal behavior analysis relies on long-term, interpretable individual trajectories; however, multi-animal tracking in space science experimen

researcharxiv-cs-cv
30 Apr 2026
Research

MTCurv: Deep learning for direct microtubule curvature mapping in noisy fluorescence microscopy images

DGX agent

arXiv:2604.26517v1 Announce Type: new Abstract: Accurate quantification of the geometry of curvilinear biological structures is essential for understanding cellular mechanics and disease-related morph

researcharxiv-cs-cv
30 Apr 2026
Local Ai

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding

DGX agent

arXiv:2604.26261v1 Announce Type: new Abstract: Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by th

local-aiarxiv-cs-cv
30 Apr 2026
Model Releases

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

DGX agent

arXiv:2604.04135v2 Announce Type: replace Abstract: This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

DGX agent

arXiv:2601.02731v3 Announce Type: replace-cross Abstract: Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers signifi

model-releasesarxiv-cs-cv
30 Apr 2026
Research

OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction

DGX agent

arXiv:2604.26252v1 Announce Type: new Abstract: Predicting social media popularity requires understanding both the intrinsic appeal of content and the external context that determines how it is expose

researcharxiv-cs-cv
30 Apr 2026
Local Ai

OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer

DGX agent

arXiv:2603.05959v3 Announce Type: replace Abstract: Reconstructing 3D geometry from streaming video requires continuous inference under bounded resources. Recent geometric foundation models achieve im

local-aiarxiv-cs-cv
30 Apr 2026
Model Releases

Perception Test 2025: Challenge Summary and a Unified VQA Extension

DGX agent

arXiv:2601.06287v2 Announce Type: replace Abstract: The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2

model-releasesarxiv-cs-cv
30 Apr 2026
Local Ai

Point Cloud Registration via Probabilistic Self-Update Local Correspondence and Line Vector Sets

DGX agent

arXiv:2604.26318v1 Announce Type: new Abstract: Point cloud registration (PCR) is a fundamental task for integrating 3D observations in remote sensing applications. This paper proposes a fast and effe

local-aiarxiv-cs-cv
30 Apr 2026
Research

Privacy-Preserving Clothing Classification using Vision Transformer for Thermal Comfort Estimation

DGX agent

arXiv:2604.26184v1 Announce Type: new Abstract: A privacy-preserving clothing classification scheme is presented to enable secure occupant-centric control (OCC) systems. Although the utilization of ca

researcharxiv-cs-cv
30 Apr 2026
Research

ProcFunc: Function-Oriented Abstractions for Procedural 3D Generation in Python

DGX agent

arXiv:2604.26943v1 Announce Type: new Abstract: We introduce ProcFunc, a library for Blender-based procedural 3D generation in Python. ProcFunc provides a library of easy-to-use Python functions, whic

researcharxiv-cs-cv
30 Apr 2026
Safety

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

DGX agent

arXiv:2510.08547v2 Announce Type: replace-cross Abstract: Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to wor

safetyarxiv-cs-cv
30 Apr 2026
Model Releases

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

DGX agent

arXiv:2604.26067v1 Announce Type: new Abstract: We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary gro

model-releasesarxiv-cs-cv
30 Apr 2026
Research

Real-time Global Illumination for Dynamic 3D Gaussian Scenes

DGX agent

arXiv:2503.17897v2 Announce Type: replace-cross Abstract: We present a real-time global illumination approach along with a pipeline for dynamic 3D Gaussian models and meshes. Building on a formulated

researcharxiv-cs-cv
30 Apr 2026
Research

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

DGX agent

arXiv:2604.26031v1 Announce Type: new Abstract: This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challen

researcharxiv-cs-cv
30 Apr 2026
Research

Revisiting Human-in-the-Loop Object Retrieval with Pre-Trained Vision Transformers

DGX agent

arXiv:2604.00809v2 Announce Type: replace Abstract: Building on existing approaches, we revisit Human-in-the-Loop Object Retrieval, a task that consists of iteratively retrieving images containing obj

researcharxiv-cs-cv
30 Apr 2026
Safety

Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training

DGX agent

arXiv:2604.26496v1 Announce Type: new Abstract: Adversarial Training (AT) is one of the most effective methods for developing robust deep neural networks (DNNs). However, AT faces a trade-off problem

safetyarxiv-cs-cv
30 Apr 2026
Research

Sample Selection Using Multi-Task Autoencoders in Federated Learning with Non-IID Data

DGX agent

arXiv:2604.26116v1 Announce Type: new Abstract: Federated learning is a machine learning paradigm in which multiple devices collaboratively train a model under the supervision of a central server whil

researcharxiv-cs-cv
30 Apr 2026
Local Ai

SAND: Spatially Adaptive Network Depth for Fast Sampling of Neural Implicit Surfaces

DGX agent

arXiv:2604.25936v1 Announce Type: cross Abstract: Implicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of networ

local-aiarxiv-cs-cv
30 Apr 2026
Model Releases

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

DGX agent

arXiv:2604.26883v1 Announce Type: new Abstract: Synthesizing a target concept from a single reference image is challenging in diffusion-based personalized text-to-image generation, particularly for st

model-releasesarxiv-cs-cv
30 Apr 2026
Research

Seamless Indoor-Outdoor Mapping for INGENIOUS First Responders

DGX agent

arXiv:2604.26368v1 Announce Type: new Abstract: In several applications it is desired to have 3D models not only from the outdoor spaces but also from inside the building. In the context of First Resp

researcharxiv-cs-cv
30 Apr 2026
Research

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

DGX agent

arXiv:2604.26262v1 Announce Type: new Abstract: Modern scene reconstruction methods, such as 3D Gaussian Splatting, enable photo-realistic novel view synthesis at real-time speeds. However, their adop

researcharxiv-cs-cv
30 Apr 2026
Research

SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design

DGX agent

arXiv:2511.13285v2 Announce Type: replace Abstract: Artistic design, particularly poster design, often demands rapid yet precise modification of textual content while preserving visual harmony and typ

researcharxiv-cs-cv
30 Apr 2026
Research

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

DGX agent

arXiv:2604.26620v1 Announce Type: new Abstract: Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in th

researcharxiv-cs-cv
30 Apr 2026
Safety

Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

DGX agent

arXiv:2604.26409v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have demonstrated significant success in interpreting Large Language Models (LLMs) by decomposing dense representations into

safetyarxiv-cs-cv
30 Apr 2026
← Previous
1…206207208209210…261
Next →