AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Apr 2026

MTCurv: Deep learning for direct microtubule curvature mapping in noisy fluorescence microscopy images

ResearchDGX agent

arXiv:2604.26517v1 Announce Type: new Abstract: Accurate quantification of the geometry of curvilinear biological structures is essential for understanding cellular mechanics and disease-related morph

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding

Local AiDGX agent

arXiv:2604.26261v1 Announce Type: new Abstract: Zero-shot 3D Visual Grounding (3DVG) is a critical capability for open-world embodied AI. However, existing methods are fundamentally bottlenecked by th

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.04135v2 Announce Type: replace Abstract: This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

Model ReleasesDGX agent

arXiv:2601.02731v3 Announce Type: replace-cross Abstract: Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers signifi

OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction

ResearchDGX agent

arXiv:2604.26252v1 Announce Type: new Abstract: Predicting social media popularity requires understanding both the intrinsic appeal of content and the external context that determines how it is expose

OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer

Local AiDGX agent

arXiv:2603.05959v3 Announce Type: replace Abstract: Reconstructing 3D geometry from streaming video requires continuous inference under bounded resources. Recent geometric foundation models achieve im

Perception Test 2025: Challenge Summary and a Unified VQA Extension

Model ReleasesDGX agent

arXiv:2601.06287v2 Announce Type: replace Abstract: The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2

Point Cloud Registration via Probabilistic Self-Update Local Correspondence and Line Vector Sets

Local AiDGX agent

arXiv:2604.26318v1 Announce Type: new Abstract: Point cloud registration (PCR) is a fundamental task for integrating 3D observations in remote sensing applications. This paper proposes a fast and effe

Privacy-Preserving Clothing Classification using Vision Transformer for Thermal Comfort Estimation

ResearchDGX agent

arXiv:2604.26184v1 Announce Type: new Abstract: A privacy-preserving clothing classification scheme is presented to enable secure occupant-centric control (OCC) systems. Although the utilization of ca

ProcFunc: Function-Oriented Abstractions for Procedural 3D Generation in Python

ResearchDGX agent

arXiv:2604.26943v1 Announce Type: new Abstract: We introduce ProcFunc, a library for Blender-based procedural 3D generation in Python. ProcFunc provides a library of easy-to-use Python functions, whic

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

SafetyDGX agent

arXiv:2510.08547v2 Announce Type: replace-cross Abstract: Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to wor

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.26067v1 Announce Type: new Abstract: We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary gro

Real-time Global Illumination for Dynamic 3D Gaussian Scenes

ResearchDGX agent

arXiv:2503.17897v2 Announce Type: replace-cross Abstract: We present a real-time global illumination approach along with a pipeline for dynamic 3D Gaussian models and meshes. Building on a formulated

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

ResearchDGX agent

arXiv:2604.26031v1 Announce Type: new Abstract: This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challen

Revisiting Human-in-the-Loop Object Retrieval with Pre-Trained Vision Transformers

ResearchDGX agent

arXiv:2604.00809v2 Announce Type: replace Abstract: Building on existing approaches, we revisit Human-in-the-Loop Object Retrieval, a task that consists of iteratively retrieving images containing obj

Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training

SafetyDGX agent

arXiv:2604.26496v1 Announce Type: new Abstract: Adversarial Training (AT) is one of the most effective methods for developing robust deep neural networks (DNNs). However, AT faces a trade-off problem

Sample Selection Using Multi-Task Autoencoders in Federated Learning with Non-IID Data

ResearchDGX agent

arXiv:2604.26116v1 Announce Type: new Abstract: Federated learning is a machine learning paradigm in which multiple devices collaboratively train a model under the supervision of a central server whil

SAND: Spatially Adaptive Network Depth for Fast Sampling of Neural Implicit Surfaces

Local AiDGX agent

arXiv:2604.25936v1 Announce Type: cross Abstract: Implicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of networ

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

Model ReleasesDGX agent

arXiv:2604.26883v1 Announce Type: new Abstract: Synthesizing a target concept from a single reference image is challenging in diffusion-based personalized text-to-image generation, particularly for st

Seamless Indoor-Outdoor Mapping for INGENIOUS First Responders

ResearchDGX agent

arXiv:2604.26368v1 Announce Type: new Abstract: In several applications it is desired to have 3D models not only from the outdoor spaces but also from inside the building. In the context of First Resp

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

ResearchDGX agent

arXiv:2604.26262v1 Announce Type: new Abstract: Modern scene reconstruction methods, such as 3D Gaussian Splatting, enable photo-realistic novel view synthesis at real-time speeds. However, their adop

SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design

ResearchDGX agent

arXiv:2511.13285v2 Announce Type: replace Abstract: Artistic design, particularly poster design, often demands rapid yet precise modification of textual content while preserving visual harmony and typ

SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

ResearchDGX agent

arXiv:2604.26620v1 Announce Type: new Abstract: Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in th

Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

SafetyDGX agent

arXiv:2604.26409v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have demonstrated significant success in interpreting Large Language Models (LLMs) by decomposing dense representations into

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness

ResearchDGX agent

arXiv:2604.26341v1 Announce Type: new Abstract: Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading

Model ReleasesDGX agent

arXiv:2604.26614v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measuremen

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

AgentsDGX agent

arXiv:2508.01875v4 Announce Type: replace Abstract: Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offl

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2604.26772v1 Announce Type: new Abstract: Recent methods demonstrate that large-scale pretrained models, such as CLIP vision transformers, effectively detect AI-generated images (AIGIs) from uns

Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention

Local AiDGX agent

arXiv:2511.20032v3 Announce Type: replace Abstract: Visual attention serves as the primary mechanism through which MLLMs interpret visual information; however, its limited localization capability ofte

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

ResearchDGX agent

arXiv:2512.20340v3 Announce Type: replace Abstract: Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing metho

The Unseen Adversaries: Robust and Generalized Defense Against Adversarial Patches

Model ReleasesDGX agent

arXiv:2604.26317v1 Announce Type: new Abstract: The vulnerabilities of deep neural networks against singularities have raised serious concerns regarding their deployment in the physical world. One of

Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation

AgentsDGX agent

arXiv:2604.26946v1 Announce Type: new Abstract: Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These

Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

SafetyDGX agent

arXiv:2604.26370v1 Announce Type: new Abstract: Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language lea

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution

TutorialsDGX agent

arXiv:2509.23980v2 Announce Type: replace Abstract: Diffusion models have recently shown promising results for video super-resolution (VSR). However, directly adapting generative diffusion models to V

U-FaceBP: Uncertainty-aware Bayesian Ensemble Deep Learning for Face Video-based Blood Pressure Estimation

ResearchDGX agent

arXiv:2412.10679v3 Announce Type: replace Abstract: Blood pressure (BP) measurement is crucial for daily health assessment. Remote photoplethysmography (rPPG), which extracts pulse waves from face vid

Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis

SafetyDGX agent

arXiv:2506.16742v3 Announce Type: replace Abstract: To be adopted in safety-critical domains like medical image analysis, AI systems must provide human-interpretable decisions. Variational Information

Uncertainty-Aware Pedestrian Attribute Recognition via Evidential Deep Learning

ApplicationsDGX agent

arXiv:2604.26873v1 Announce Type: new Abstract: We propose UAPAR, an Uncertainty-Aware Pedestrian Attribute Recognition framework. To the best of our knowledge, this is the first EDL-based uncertainty

ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection

SafetyDGX agent

arXiv:2604.26218v1 Announce Type: new Abstract: Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward vis

Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery

HardwareDGX agent

arXiv:2603.05811v2 Announce Type: replace Abstract: Current video generation models suffer from high computational latency, making real-time applications prohibitively costly. In this paper, we addres

Virtual-reality based patient-specific simulation of spine surgical procedures: A fast, highly automated and high-fidelity system for surgical education and planning

ApplicationsDGX agent

arXiv:2604.26781v1 Announce Type: new Abstract: Surgical training involves didactic teaching, mentor-led learning, surgical skills laboratories, and direct exposure to surgery; however, increasing cli

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

SafetyDGX agent

arXiv:2505.20032v3 Announce Type: replace Abstract: Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite rec

Which Face and Whose Identity? Solving the Dual Challenge of Deepfake Proactive Forensics in Multi-Face Scenarios

ApplicationsDGX agent

arXiv:2604.26342v1 Announce Type: new Abstract: Unlike single-face forgeries, deepfakes in complex multi-person interaction scenarios (such as group photos and multi-person meetings) more closely refl

Why Domain Matters: A Preliminary Study of Domain Effects in Underwater Object Detection

ApplicationsDGX agent

arXiv:2604.26174v1 Announce Type: new Abstract: Domain shift, where deviations between training and deployment data distributions degrade model performance, is a key challenge in underwater environmen

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

ResearchDGX agent

arXiv:2604.26934v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that

29 Apr 2026

8DNA: 8D Neural Asset Light Transport by Distribution Learning

ResearchDGX agent

arXiv:2604.25129v1 Announce Type: cross Abstract: High-fidelity 3D assets exhibit intriguing global illumination effects like subsurface scattering, glossy interreflections, and fine-scale fiber scatt

A Comparative Study in Surgical AI: Datasets, Foundation Models, and Barriers to Med-AGI

Model ReleasesDGX agent

arXiv:2603.27341v2 Announce Type: replace-cross Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but su

A graph generation pipeline for critical infrastructures based on heuristics, images and depth data

ResearchDGX agent

arXiv:2512.07269v2 Announce Type: replace Abstract: Virtual representations of physical critical infrastructures, such as water or energy plants, are used for simulations and digital twins to ensure r

A New Kind of Network? Review and Reference Implementation of Neural Cellular Automata

TutorialsDGX agent

arXiv:2604.24990v1 Announce Type: new Abstract: Stephen Wolfram proclaimed in his 2003 seminal work 'A New Kind Of Science' that simple recursive programs in the form of Cellular Automata (CA) are a p

A Systematic Post-Train Framework for Video Generation

SafetyDGX agent

arXiv:2604.25427v1 Announce Type: new Abstract: While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a signif

Accuracy Improvement of Cell Image Segmentation Using Feedback Former

ResearchDGX agent

arXiv:2408.12974v4 Announce Type: replace Abstract: Semantic segmentation of microscopy cell images by deep learning is a significant technique. We considered that the Transformers, which have recentl

AdaTooler-V: Adaptive Tool-Use for Images and Videos

Model ReleasesDGX agent

arXiv:2512.16918v3 Announce Type: replace Abstract: Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool

Agentic AI for Remote Sensing: Technical Challenges and Research Directions

AgentsDGX agent

arXiv:2604.24919v1 Announce Type: new Abstract: Earth Observation (EO) is moving beyond static prediction toward multi-step analytical workflows that require coordinated reasoning over data, tools, an

AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization

AgentsDGX agent

arXiv:2410.24116v3 Announce Type: replace Abstract: Image labeling is a critical bottleneck in the development of computer vision technologies, often constraining machine learning performance due to t

Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

Model ReleasesDGX agent

arXiv:2602.23069v2 Announce Type: replace Abstract: Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are fa

ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs

ResearchDGX agent

arXiv:2410.24214v3 Announce Type: replace-cross Abstract: Mixed precision quantization has become an important technique for optimizing the execution of deep neural networks (DNNs). Certified robustne

Assessment of the quantitative impact of occlusal positioning splints on temporomandibular joint conditions

TutorialsDGX agent

arXiv:2604.25322v1 Announce Type: new Abstract: A computational method for quantitative analysis of temporomandibular joint (TMJ) configuration using occlusal positioning splints is proposed and demon

Automated detection of pediatric congenital heart disease from phonocardiograms using deep and handcrafted feature fusion

ResearchDGX agent

arXiv:2604.24767v1 Announce Type: cross Abstract: Congenital heart disease (CHD) is the most common type of birth defect, impacting about 1% of live births worldwide. Echocardiography, the gold-standa

Benchmarking and Improving GUI Agents in High-Dynamic Environments

Model ReleasesDGX agent

arXiv:2604.25380v1 Announce Type: new Abstract: Recent advancements in Graphical User Interface (GUI) agents have predominantly focused on training paradigms like supervised fine-tuning (SFT) and rein

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings

Model ReleasesDGX agent

arXiv:2604.25358v1 Announce Type: new Abstract: Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescri

Benchmarking OCR Pipelines with Adaptive Enhancement for Multi-Domain Retail Bill Digitization

Model ReleasesDGX agent

arXiv:2604.25176v1 Announce Type: new Abstract: The digitization of multi-domain retail billing documents remains a challenging task due to variability in scan quality, layout heterogeneity, and domai

← Previous
1…165166167168169…209
Next →