AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
10 Apr 2026

On the Uphill Battle of Image frequency Analysis

ResearchDGX agent

arXiv:2604.07563v1 Announce Type: new Abstract: This work is a follow up on the newly proposed clustering algorithm called The Inverse Square Mean Shift Algorithm. In this paper a special case of algo

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

Model ReleasesDGX agent

arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int

OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation

ResearchDGX agent

arXiv:2512.03532v2 Announce Type: replace Abstract: Generalizing open-vocabulary 3D instance segmentation (OV-3DIS) to diverse, unstructured, and mesh-free environments is crucial for robotics and AR/


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models

Model ReleasesDGX agent

arXiv:2604.08266v1 Announce Type: new Abstract: Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems

oslash Source Models Leak What They Shouldn't nrightarrow: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization

Model ReleasesDGX agent

arXiv:2604.08238v1 Announce Type: new Abstract: The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may

OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation

ResearchDGX agent

arXiv:2604.08110v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance

SafetyDGX agent

arXiv:2604.08461v1 Announce Type: new Abstract: Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based a

OxEnsemble: Fair Ensembles for Low-Data Classification

SafetyDGX agent

arXiv:2512.09665v2 Announce Type: replace Abstract: We address the problem of fair classification in settings where data is scarce and unbalanced across demographic groups. Such low-data regimes are c

PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

ResearchDGX agent

arXiv:2602.06912v2 Announce Type: replace Abstract: Unsupervised segmentation from self-supervised ViT patches holds promise but lacks robustness: multi-object scenes confound saliency cues, and low-s

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation

ResearchDGX agent

arXiv:2604.07901v1 Announce Type: new Abstract: 360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications,

ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models

AgentsDGX agent

arXiv:2604.07912v1 Announce Type: new Abstract: Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant ent

ParseBench: A Document Parsing Benchmark for AI Agents

Model ReleasesDGX agent

arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and

Part^{2}GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting

SafetyDGX agent

arXiv:2506.17212v2 Announce Type: replace Abstract: Articulated objects are common in the real world, yet modeling their structure and motion remains a challenging task for 3D reconstruction methods.

Personalizing Text-to-Image Generation to Individual Taste

SafetyDGX agent

arXiv:2604.07427v1 Announce Type: new Abstract: Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models opt

Phantasia: Context-Adaptive Backdoors in Vision Language Models

ResearchDGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

ApplicationsDGX agent

arXiv:2604.08503v1 Announce Type: new Abstract: Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

Model ReleasesDGX agent

arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

Model ReleasesDGX agent

arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe

Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting

TutorialsDGX agent

arXiv:2503.09640v2 Announce Type: replace-cross Abstract: Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applicat

PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization

Local AiDGX agent

arXiv:2503.24135v3 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) methods allow training models to classify images and localize ROIs. WSOL only requires low-cost image-c

Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models

SafetyDGX agent

arXiv:2604.07779v1 Announce Type: new Abstract: Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagn

PLUME: Latent Reasoning Based Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

SafetyDGX agent

arXiv:2604.08125v1 Announce Type: new Abstract: Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are l

Preventing Overfitting in Deep Image Prior for Hyperspectral Image Denoising

ResearchDGX agent

arXiv:2604.08272v1 Announce Type: new Abstract: Deep image prior (DIP) is an unsupervised deep learning framework that has been successfully applied to a variety of inverse imaging problems. However,

Privacy Attacks on Image AutoRegressive Models

ResearchDGX agent

arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m

PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation

Model ReleasesDGX agent

arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech

Pseudo-Expert Regularized Offline RL for End-to-End Autonomous Driving in Photorealistic Closed-Loop Environments

AgentsDGX agent

arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Model ReleasesDGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

Quantifying Explanation Consistency: The C-Score Metric for CAM-Based Explainability in Medical Image Classification

Local AiDGX agent

arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi

RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing

HardwareDGX agent

arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r

Reading Recognition in the Wild

TutorialsDGX agent

arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning

SafetyDGX agent

arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

ResearchDGX agent

arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

ResearchDGX agent

arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

Model ReleasesDGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

AgentsDGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

Revisiting Radar Perception With Spectral Point Clouds

Model ReleasesDGX agent

arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp

RewardFlow: Generate Images by Optimizing What You Reward

SafetyDGX agent

arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

SafetyDGX agent

arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli

Rotation Equivariant Convolutions in Deformable Registration of Brain MRI

SafetyDGX agent

arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a

RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection

AgentsDGX agent

arXiv:2505.17732v2 Announce Type: replace Abstract: Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have eme

Sampling-Aware 3D Spatial Analysis in Multiplexed Imaging

ResearchDGX agent

arXiv:2604.07890v1 Announce Type: new Abstract: Highly multiplexed microscopy enables rich spatial characterization of tissues at single-cell resolution, yet most analyses rely on two-dimensional sect

SAT: Selective Aggregation Transformer for Image Super-Resolution

Local AiDGX agent

arXiv:2604.07994v1 Announce Type: new Abstract: Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational comple

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

Local AiDGX agent

arXiv:2604.08542v1 Announce Type: new Abstract: This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown pro

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

AgentsDGX agent

arXiv:2604.08366v1 Announce Type: cross Abstract: Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, th

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

Model ReleasesDGX agent

arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection

Model ReleasesDGX agent

arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research

SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

ResearchDGX agent

arXiv:2604.03134v2 Announce Type: replace Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

Self-Improving 4D Perception via Self-Distillation

ResearchDGX agent

arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

SafetyDGX agent

arXiv:2604.08147v1 Announce Type: cross Abstract: Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, j

SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP

SafetyDGX agent

arXiv:2509.26036v3 Announce Type: replace Abstract: While Contrastive Language-Image Pretraining (CLIP) excels at zero-shot tasks by aligning image and text embeddings, its performance in few-shot cla

Shortcut Learning in Glomerular AI: Adversarial Penalties Hurt, Entropy Helps

SafetyDGX agent

arXiv:2604.07936v1 Announce Type: new Abstract: Stain variability is a pervasive source of distribution shift and potential shortcut learning in renal pathology AI. We ask whether lupus nephritis glom

SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds

SafetyDGX agent

arXiv:2604.08544v1 Announce Type: cross Abstract: Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in

SMFD-UNet: Semantic Face Mask Is The Only Thing You Need To Deblur Faces

ApplicationsDGX agent

arXiv:2604.07477v1 Announce Type: new Abstract: For applications including facial identification, forensic analysis, photographic improvement, and medical imaging diagnostics, facial image deblurring

SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models

SafetyDGX agent

arXiv:2504.13378v2 Announce Type: replace-cross Abstract: Generating high-quality, photorealistic textures for 3D human avatars remains a fundamental yet challenging task in computer vision and multim

SonoSelect: Efficient Ultrasound Perception via Active Probe Exploration

ResearchDGX agent

arXiv:2604.05933v3 Announce Type: replace Abstract: Ultrasound perception typically requires multiple scan views through probe movement to reduce diagnostic ambiguity, mitigate acoustic occlusions, an

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility

Model ReleasesDGX agent

arXiv:2512.23365v3 Announce Type: replace Abstract: The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning.

Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

Model ReleasesDGX agent

arXiv:2604.07923v1 Announce Type: new Abstract: Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing

← Previous
1…206207208209
Next →