AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
10 Apr 2026

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

Model ReleasesDGX agent

arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe

Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting

TutorialsDGX agent

arXiv:2503.09640v2 Announce Type: replace-cross Abstract: Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applicat

PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2503.24135v3 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) methods allow training models to classify images and localize ROIs. WSOL only requires low-cost image-c

Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models

SafetyDGX agent

arXiv:2604.07779v1 Announce Type: new Abstract: Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagn

PLUME: Latent Reasoning Based Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

SafetyDGX agent

arXiv:2604.08125v1 Announce Type: new Abstract: Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are l

Preventing Overfitting in Deep Image Prior for Hyperspectral Image Denoising

ResearchDGX agent

arXiv:2604.08272v1 Announce Type: new Abstract: Deep image prior (DIP) is an unsupervised deep learning framework that has been successfully applied to a variety of inverse imaging problems. However,

Privacy Attacks on Image AutoRegressive Models

ResearchDGX agent

arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m

PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation

Model ReleasesDGX agent

arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech

Pseudo-Expert Regularized Offline RL for End-to-End Autonomous Driving in Photorealistic Closed-Loop Environments

AgentsDGX agent

arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Model ReleasesDGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

Quantifying Explanation Consistency: The C-Score Metric for CAM-Based Explainability in Medical Image Classification

Local AiDGX agent

arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi

RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing

HardwareDGX agent

arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r

Reading Recognition in the Wild

TutorialsDGX agent

arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning

SafetyDGX agent

arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

ResearchDGX agent

arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

ResearchDGX agent

arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

Model ReleasesDGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

AgentsDGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

Revisiting Radar Perception With Spectral Point Clouds

Model ReleasesDGX agent

arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp

RewardFlow: Generate Images by Optimizing What You Reward

SafetyDGX agent

arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

SafetyDGX agent

arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli

Rotation Equivariant Convolutions in Deformable Registration of Brain MRI

SafetyDGX agent

arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a

RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection

AgentsDGX agent

arXiv:2505.17732v2 Announce Type: replace Abstract: Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have eme

Sampling-Aware 3D Spatial Analysis in Multiplexed Imaging

ResearchDGX agent

arXiv:2604.07890v1 Announce Type: new Abstract: Highly multiplexed microscopy enables rich spatial characterization of tissues at single-cell resolution, yet most analyses rely on two-dimensional sect

SAT: Selective Aggregation Transformer for Image Super-Resolution

Local AiDGX agent

arXiv:2604.07994v1 Announce Type: new Abstract: Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational comple

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

Local AiDGX agent

arXiv:2604.08542v1 Announce Type: new Abstract: This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown pro

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

AgentsDGX agent

arXiv:2604.08366v1 Announce Type: cross Abstract: Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, th

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

Model ReleasesDGX agent

arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection

Model ReleasesDGX agent

arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research

SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

ResearchDGX agent

arXiv:2604.03134v2 Announce Type: replace Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

Self-Improving 4D Perception via Self-Distillation

ResearchDGX agent

arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

SafetyDGX agent

arXiv:2604.08147v1 Announce Type: cross Abstract: Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, j

SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP

SafetyDGX agent

arXiv:2509.26036v3 Announce Type: replace Abstract: While Contrastive Language-Image Pretraining (CLIP) excels at zero-shot tasks by aligning image and text embeddings, its performance in few-shot cla

Shortcut Learning in Glomerular AI: Adversarial Penalties Hurt, Entropy Helps

SafetyDGX agent

arXiv:2604.07936v1 Announce Type: new Abstract: Stain variability is a pervasive source of distribution shift and potential shortcut learning in renal pathology AI. We ask whether lupus nephritis glom

SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds

SafetyDGX agent

arXiv:2604.08544v1 Announce Type: cross Abstract: Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in

SMFD-UNet: Semantic Face Mask Is The Only Thing You Need To Deblur Faces

ApplicationsDGX agent

arXiv:2604.07477v1 Announce Type: new Abstract: For applications including facial identification, forensic analysis, photographic improvement, and medical imaging diagnostics, facial image deblurring

SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models

SafetyDGX agent

arXiv:2504.13378v2 Announce Type: replace-cross Abstract: Generating high-quality, photorealistic textures for 3D human avatars remains a fundamental yet challenging task in computer vision and multim

SonoSelect: Efficient Ultrasound Perception via Active Probe Exploration

ResearchDGX agent

arXiv:2604.05933v3 Announce Type: replace Abstract: Ultrasound perception typically requires multiple scan views through probe movement to reduce diagnostic ambiguity, mitigate acoustic occlusions, an

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility

Model ReleasesDGX agent

arXiv:2512.23365v3 Announce Type: replace Abstract: The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning.

Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

Model ReleasesDGX agent

arXiv:2604.07923v1 Announce Type: new Abstract: Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing

SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

Model ReleasesDGX agent

arXiv:2602.22683v2 Announce Type: replace Abstract: The rapid advancement of AI-powered smart glasses-one of the hottest wearable devices-has unlocked new frontiers for multimodal interaction, with Vi

SurfelSplat: Learning Efficient and Generalizable Gaussian Surfel Representations for Sparse-View Surface Reconstruction

ResearchDGX agent

arXiv:2604.08370v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has demonstrated impressive performance in 3D scene reconstruction. Beyond novel view synthesis, it shows great potential f

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation

ResearchDGX agent

arXiv:2412.10437v3 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

ResearchDGX agent

arXiv:2604.08405v1 Announce Type: new Abstract: Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinf

T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation

Model ReleasesDGX agent

arXiv:2604.08167v1 Announce Type: new Abstract: Medical image segmentation traditionally relies on fully supervised 3D architectures that demand a large amount of dense, voxel-level annotations from c

Tabular GANs for uneven distribution

Model ReleasesDGX agent

arXiv:2010.00638v2 Announce Type: replace-cross Abstract: Generative models for tabular data have evolved rapidly beyond Generative Adversarial Networks (GANs). While GANs pioneered synthetic tabular

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

ResearchDGX agent

arXiv:2604.07916v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and

Tensor-Augmented Convolutional Neural Networks: Enhancing Expressivity with Generic Tensor Kernels

Model ReleasesDGX agent

arXiv:2604.08072v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) excel at extracting local features hierarchically, but their performance in capturing complex correlations hinges h

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models

ResearchDGX agent

arXiv:2511.11435v3 Announce Type: replace Abstract: The ambiguity between generalization and memorization in TTI diffusion models becomes pronounced when prompts invoke culturally shared visual refere

The Weaponization of Computer Vision: Tracing Military-Surveillance Ties through Conference Sponsorship

ResearchDGX agent

arXiv:2604.07803v1 Announce Type: cross Abstract: Computer vision, a core domain of artificial intelligence (AI), is the field that enables the computational analysis, understanding, and generation of

Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding

ResearchDGX agent

arXiv:2503.10183v4 Announce Type: replace Abstract: Existing vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not groun

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

Model ReleasesDGX agent

arXiv:2512.08410v2 Announce Type: replace Abstract: Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose

Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)

ResearchDGX agent

arXiv:2604.07522v1 Announce Type: new Abstract: Positional encoding has become the de facto standard for grounding deep neural networks on discrete point-wise positions, and it has achieved remarkable

U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations

ResearchDGX agent

arXiv:2604.08295v1 Announce Type: cross Abstract: As AI models grow more complex, explainability is essential for building trust, yet concept-based counterfactual methods still face a trade-off betwee

Understanding Task Transfer in Vision-Language Models

TutorialsDGX agent

arXiv:2511.18787v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like dep

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

ResearchDGX agent

arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

← Previous
1…208209210211
Next →