AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness

DGX agent

arXiv:2604.26341v1 Announce Type: new Abstract: Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image

researcharxiv-cs-cv
30 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading

DGX agent

arXiv:2604.26614v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measuremen

model-releasesarxiv-cs-cv
30 Apr 2026
Agents

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

DGX agent

arXiv:2508.01875v4 Announce Type: replace Abstract: Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offl

agentsarxiv-cs-cv
30 Apr 2026
Model Releases

TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

DGX agent

arXiv:2604.26772v1 Announce Type: new Abstract: Recent methods demonstrate that large-scale pretrained models, such as CLIP vision transformers, effectively detect AI-generated images (AIGIs) from uns

model-releasesarxiv-cs-cv
30 Apr 2026
Local Ai

Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention

DGX agent

arXiv:2511.20032v3 Announce Type: replace Abstract: Visual attention serves as the primary mechanism through which MLLMs interpret visual information; however, its limited localization capability ofte

local-aiarxiv-cs-cv
30 Apr 2026
Research

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

DGX agent

arXiv:2512.20340v3 Announce Type: replace Abstract: Although diffusion transformer (DiT)-based video virtual try-on (VVT) has made significant progress in synthesizing realistic videos, existing metho

researcharxiv-cs-cv
30 Apr 2026
Model Releases

The Unseen Adversaries: Robust and Generalized Defense Against Adversarial Patches

DGX agent

arXiv:2604.26317v1 Announce Type: new Abstract: The vulnerabilities of deep neural networks against singularities have raised serious concerns regarding their deployment in the physical world. One of

model-releasesarxiv-cs-cv
30 Apr 2026
Agents

Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation

DGX agent

arXiv:2604.26946v1 Announce Type: new Abstract: Breakthrough progress in vision-based navigation through unknown environments has been achieved by using multimodal large language models (MLLMs). These

agentsarxiv-cs-cv
30 Apr 2026
Safety

Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

DGX agent

arXiv:2604.26370v1 Announce Type: new Abstract: Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language lea

safetyarxiv-cs-cv
30 Apr 2026
Tutorials

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution

DGX agent

arXiv:2509.23980v2 Announce Type: replace Abstract: Diffusion models have recently shown promising results for video super-resolution (VSR). However, directly adapting generative diffusion models to V

tutorialsarxiv-cs-cv
30 Apr 2026
Research

U-FaceBP: Uncertainty-aware Bayesian Ensemble Deep Learning for Face Video-based Blood Pressure Estimation

DGX agent

arXiv:2412.10679v3 Announce Type: replace Abstract: Blood pressure (BP) measurement is crucial for daily health assessment. Remote photoplethysmography (rPPG), which extracts pulse waves from face vid

researcharxiv-cs-cv
30 Apr 2026
Safety

Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis

DGX agent

arXiv:2506.16742v3 Announce Type: replace Abstract: To be adopted in safety-critical domains like medical image analysis, AI systems must provide human-interpretable decisions. Variational Information

safetyarxiv-cs-cv
30 Apr 2026
Applications

Uncertainty-Aware Pedestrian Attribute Recognition via Evidential Deep Learning

DGX agent

arXiv:2604.26873v1 Announce Type: new Abstract: We propose UAPAR, an Uncertainty-Aware Pedestrian Attribute Recognition framework. To the best of our knowledge, this is the first EDL-based uncertainty

applicationsarxiv-cs-cv
30 Apr 2026
Safety

ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection

DGX agent

arXiv:2604.26218v1 Announce Type: new Abstract: Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward vis

safetyarxiv-cs-cv
30 Apr 2026
Hardware

Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery

DGX agent

arXiv:2603.05811v2 Announce Type: replace Abstract: Current video generation models suffer from high computational latency, making real-time applications prohibitively costly. In this paper, we addres

hardwarearxiv-cs-cv
30 Apr 2026
Applications

Virtual-reality based patient-specific simulation of spine surgical procedures: A fast, highly automated and high-fidelity system for surgical education and planning

DGX agent

arXiv:2604.26781v1 Announce Type: new Abstract: Surgical training involves didactic teaching, mentor-led learning, surgical skills laboratories, and direct exposure to surgery; however, increasing cli

applicationsarxiv-cs-cv
30 Apr 2026
Safety

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

DGX agent

arXiv:2505.20032v3 Announce Type: replace Abstract: Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite rec

safetyarxiv-cs-cv
30 Apr 2026
Applications

Which Face and Whose Identity? Solving the Dual Challenge of Deepfake Proactive Forensics in Multi-Face Scenarios

DGX agent

arXiv:2604.26342v1 Announce Type: new Abstract: Unlike single-face forgeries, deepfakes in complex multi-person interaction scenarios (such as group photos and multi-person meetings) more closely refl

applicationsarxiv-cs-cv
30 Apr 2026
Applications

Why Domain Matters: A Preliminary Study of Domain Effects in Underwater Object Detection

DGX agent

arXiv:2604.26174v1 Announce Type: new Abstract: Domain shift, where deviations between training and deployment data distributions degrade model performance, is a key challenge in underwater environmen

applicationsarxiv-cs-cv
30 Apr 2026
Research

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

DGX agent

arXiv:2604.26934v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that

researcharxiv-cs-cv
30 Apr 2026
Research

8DNA: 8D Neural Asset Light Transport by Distribution Learning

DGX agent

arXiv:2604.25129v1 Announce Type: cross Abstract: High-fidelity 3D assets exhibit intriguing global illumination effects like subsurface scattering, glossy interreflections, and fine-scale fiber scatt

researcharxiv-cs-cv
29 Apr 2026
Model Releases

A Comparative Study in Surgical AI: Datasets, Foundation Models, and Barriers to Med-AGI

DGX agent

arXiv:2603.27341v2 Announce Type: replace-cross Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but su

model-releasesarxiv-cs-cv
29 Apr 2026
Research

A graph generation pipeline for critical infrastructures based on heuristics, images and depth data

DGX agent

arXiv:2512.07269v2 Announce Type: replace Abstract: Virtual representations of physical critical infrastructures, such as water or energy plants, are used for simulations and digital twins to ensure r

researcharxiv-cs-cv
29 Apr 2026
Tutorials

A New Kind of Network? Review and Reference Implementation of Neural Cellular Automata

DGX agent

arXiv:2604.24990v1 Announce Type: new Abstract: Stephen Wolfram proclaimed in his 2003 seminal work 'A New Kind Of Science' that simple recursive programs in the form of Cellular Automata (CA) are a p

tutorialsarxiv-cs-cv
29 Apr 2026
Safety

A Systematic Post-Train Framework for Video Generation

DGX agent

arXiv:2604.25427v1 Announce Type: new Abstract: While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a signif

safetyarxiv-cs-cv
29 Apr 2026
Research

Accuracy Improvement of Cell Image Segmentation Using Feedback Former

DGX agent

arXiv:2408.12974v4 Announce Type: replace Abstract: Semantic segmentation of microscopy cell images by deep learning is a significant technique. We considered that the Transformers, which have recentl

researcharxiv-cs-cv
29 Apr 2026
Model Releases

AdaTooler-V: Adaptive Tool-Use for Images and Videos

DGX agent

arXiv:2512.16918v3 Announce Type: replace Abstract: Recent advances have shown that multimodal large language models (MLLMs) benefit from multimodal interleaved chain-of-thought (CoT) with vision tool

model-releasesarxiv-cs-cv
29 Apr 2026
Agents

Agentic AI for Remote Sensing: Technical Challenges and Research Directions

DGX agent

arXiv:2604.24919v1 Announce Type: new Abstract: Earth Observation (EO) is moving beyond static prediction toward multi-step analytical workflows that require coordinated reasoning over data, tools, an

agentsarxiv-cs-cv
29 Apr 2026
Agents

AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization

DGX agent

arXiv:2410.24116v3 Announce Type: replace Abstract: Image labeling is a critical bottleneck in the development of computer vision technologies, often constraining machine learning performance due to t

agentsarxiv-cs-cv
29 Apr 2026
Model Releases

Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

DGX agent

arXiv:2602.23069v2 Announce Type: replace Abstract: Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are fa

model-releasesarxiv-cs-cv
29 Apr 2026
Research

ARQ: A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs

DGX agent

arXiv:2410.24214v3 Announce Type: replace-cross Abstract: Mixed precision quantization has become an important technique for optimizing the execution of deep neural networks (DNNs). Certified robustne

researcharxiv-cs-cv
29 Apr 2026
Tutorials

Assessment of the quantitative impact of occlusal positioning splints on temporomandibular joint conditions

DGX agent

arXiv:2604.25322v1 Announce Type: new Abstract: A computational method for quantitative analysis of temporomandibular joint (TMJ) configuration using occlusal positioning splints is proposed and demon

tutorialsarxiv-cs-cv
29 Apr 2026
Research

Automated detection of pediatric congenital heart disease from phonocardiograms using deep and handcrafted feature fusion

DGX agent

arXiv:2604.24767v1 Announce Type: cross Abstract: Congenital heart disease (CHD) is the most common type of birth defect, impacting about 1% of live births worldwide. Echocardiography, the gold-standa

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Benchmarking and Improving GUI Agents in High-Dynamic Environments

DGX agent

arXiv:2604.25380v1 Announce Type: new Abstract: Recent advancements in Graphical User Interface (GUI) agents have predominantly focused on training paradigms like supervised fine-tuning (SFT) and rein

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings

DGX agent

arXiv:2604.25358v1 Announce Type: new Abstract: Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescri

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Benchmarking OCR Pipelines with Adaptive Enhancement for Multi-Domain Retail Bill Digitization

DGX agent

arXiv:2604.25176v1 Announce Type: new Abstract: The digitization of multi-domain retail billing documents remains a challenging task due to variability in scan quality, layout heterogeneity, and domai

model-releasesarxiv-cs-cv
29 Apr 2026
Agents

BEVal: A Cross-dataset Evaluation Study of BEV Segmentation Models for Autonomous Driving

DGX agent

arXiv:2408.16322v4 Announce Type: replace Abstract: Current research in semantic bird's-eye view segmentation for autonomous driving focuses solely on optimizing neural network models using a single d

agentsarxiv-cs-cv
29 Apr 2026
Safety

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

DGX agent

arXiv:2604.25072v1 Announce Type: new Abstract: Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evalua

safetyarxiv-cs-cv
29 Apr 2026
Research

Beyond Fidelity: Semantic Similarity Assessment in Low-Level Image Processing

DGX agent

arXiv:2604.25408v1 Announce Type: new Abstract: Low-level image processing has long been evaluated mainly from the perspective of visual fidelity. However, with the rise of deep learning and generativ

researcharxiv-cs-cv
29 Apr 2026
Model Releases

BifDet: A 3D Bifurcation Detection Dataset for Airway-Tree Modeling

DGX agent

arXiv:2604.24999v1 Announce Type: new Abstract: Thoracic Computed Tomography (CT) scans offer detailed insights into the intricate branching network of the airway tree, which is essential for understa

model-releasesarxiv-cs-cv
29 Apr 2026
Tutorials

C3G: Learning Compact 3D Representations with 2K Gaussians

DGX agent

arXiv:2512.04021v2 Announce Type: replace Abstract: Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. R

tutorialsarxiv-cs-cv
29 Apr 2026
Research

Can We Change the Stroke Size for Easier Diffusion?

DGX agent

arXiv:2603.26783v2 Announce Type: replace Abstract: Diffusion models can be challenged in the low signal-to-noise regime, where they have to make pixel-level predictions despite the presence of high n

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

DGX agent

arXiv:2604.25273v1 Announce Type: new Abstract: Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus

model-releasesarxiv-cs-cv
29 Apr 2026
Local Ai

COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization

DGX agent

arXiv:2604.25388v1 Announce Type: new Abstract: Architectural floor plans are widely available priors which contain not only geometry but also the semantic information of the environment, yet existing

local-aiarxiv-cs-cv
29 Apr 2026
Agents

Control Your Queries: Heterogeneous Query Interaction for Camera-Radar Fusion

DGX agent

arXiv:2604.25574v1 Announce Type: new Abstract: In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, f

agentsarxiv-cs-cv
29 Apr 2026
Model Releases

CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation

DGX agent

arXiv:2604.25376v1 Announce Type: new Abstract: Accurate brain lesion segmentation in MRI is vital for effective clinical diagnosis and treatment planning. Due to high annotation costs and strict data

model-releasesarxiv-cs-cv
29 Apr 2026
Research

CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images

DGX agent

arXiv:2604.24793v1 Announce Type: cross Abstract: We present CRC-SAM, a unified framework for colorectal cancer segmentation across colonoscopy, CT, and histopathology images. Unlike prior single-moda

researcharxiv-cs-cv
29 Apr 2026
Tutorials

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

DGX agent

arXiv:2604.25477v1 Announce Type: new Abstract: Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance t

tutorialsarxiv-cs-cv
29 Apr 2026
← Previous
1…207208209210211…261
Next →