AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
10 Apr 2026

Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks

Model ReleasesDGX agent

arXiv:2510.15770v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by predicting human-understandable concepts as intermediate representations. However, exis

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework

Model ReleasesDGX agent

arXiv:2509.23322v2 Announce Type: replace Abstract: With the continuous expansion of Large Language Models (LLMs) and advances in reinforcement learning, LLMs have demonstrated exceptional reasoning c

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2412.20718v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (LVLMs) into critical domains necessitates comprehensive moral evaluation to ensure their alig

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

AgentsDGX agent

arXiv:2604.08516v1 Announce Type: new Abstract: Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with t

Monocular Depth Estimation From the Perspective of Feature Restoration: A Diffusion Enhanced Depth Restoration Approach

Model ReleasesDGX agent

arXiv:2604.07664v1 Announce Type: new Abstract: Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods empl

MonoUNet: A Robust Tiny Neural Network for Automated Knee Cartilage Segmentation on Point-of-Care Ultrasound Devices

SafetyDGX agent

arXiv:2604.07780v1 Announce Type: cross Abstract: Objective: To develop a robust and compact deep learning model for automated knee cartilage segmentation on point-of-care ultrasound (POCUS) devices.

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models

SafetyDGX agent

arXiv:2604.07991v1 Announce Type: new Abstract: Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation f

MSCT: Differential Cross-Modal Attention for Deepfake Detection

SafetyDGX agent

arXiv:2604.07741v1 Announce Type: new Abstract: Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily ex

MSGL-Transformer: A Multi-Scale Global-Local Transformer for Rodent Social Behavior Recognition

Local AiDGX agent

arXiv:2604.07578v1 Announce Type: new Abstract: Recognition of rodent behavior is important for understanding neural and behavioral mechanisms. Traditional manual scoring is time-consuming and prone t

MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation

ApplicationsDGX agent

arXiv:2603.11633v2 Announce Type: replace Abstract: Recent unified 3D generation models have made remarkable progress in producing high-quality 3D assets from a single image. Notably, layout-aware app

MVOS_HSI: A Python Library for Preprocessing Agricultural Crop Hyperspectral Data

ResearchDGX agent

arXiv:2604.07656v1 Announce Type: cross Abstract: Hyperspectral imaging (HSI) allows researchers to study plant traits non-destructively. By capturing hundreds of narrow spectral bands per pixel, it r

Nearest Neighbor Projection Removal Adversarial Training

ResearchDGX agent

arXiv:2509.07673v4 Announce Type: replace Abstract: Deep neural networks have exhibited impressive performance in image classification tasks but remain vulnerable to adversarial examples. Standard adv

Needle in a Haystack -- One-Class Representation Learning for Detecting Rare Malignant Cells in Computational Cytology

TutorialsDGX agent

arXiv:2604.07722v1 Announce Type: new Abstract: In computational cytology, detecting malignancy on whole-slide images is difficult because malignant cells are morphologically diverse yet vanishingly r

Novel View Synthesis as Video Completion

ResearchDGX agent

arXiv:2604.08500v1 Announce Type: new Abstract: We tackle the problem of sparse novel view synthesis (NVS) using video diffusion models; given K (approx 5) multi-view images of a scene and their

NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results

Model ReleasesDGX agent

arXiv:2604.06945v2 Announce Type: replace Abstract: This paper reports on the NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration (BSCVR). The challenge aims to advance research on recoverin

Object-Centric Stereo Ranging for Autonomous Driving: From Dense Disparity to Census-Based Template Matching

HardwareDGX agent

arXiv:2604.07980v1 Announce Type: new Abstract: Accurate depth estimation is critical for autonomous driving perception systems, particularly for long range vehicle detection on highways. Traditional

OceanMAE: A Foundation Model for Ocean Remote Sensing

TutorialsDGX agent

arXiv:2604.08171v1 Announce Type: new Abstract: Accurate ocean mapping is essential for applications such as bathymetry estimation, seabed characterization, marine litter detection, and ecosystem moni

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

ResearchDGX agent

arXiv:2604.08209v1 Announce Type: new Abstract: To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative

On the Global Photometric Alignment for Low-Level Vision

SafetyDGX agent

arXiv:2604.08172v1 Announce Type: new Abstract: Supervised low-level vision models rely on pixel-wise losses against paired references, yet paired training sets exhibit per-pair photometric inconsiste

On the Uphill Battle of Image frequency Analysis

ResearchDGX agent

arXiv:2604.07563v1 Announce Type: new Abstract: This work is a follow up on the newly proposed clustering algorithm called The Inverse Square Mean Shift Algorithm. In this paper a special case of algo

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

Model ReleasesDGX agent

arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int

OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation

ResearchDGX agent

arXiv:2512.03532v2 Announce Type: replace Abstract: Generalizing open-vocabulary 3D instance segmentation (OV-3DIS) to diverse, unstructured, and mesh-free environments is crucial for robotics and AR/

Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models

Model ReleasesDGX agent

arXiv:2604.08266v1 Announce Type: new Abstract: Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems

oslash Source Models Leak What They Shouldn't nrightarrow: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization

Model ReleasesDGX agent

arXiv:2604.08238v1 Announce Type: new Abstract: The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may

OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation

ResearchDGX agent

arXiv:2604.08110v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance

SafetyDGX agent

arXiv:2604.08461v1 Announce Type: new Abstract: Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based a

OxEnsemble: Fair Ensembles for Low-Data Classification

SafetyDGX agent

arXiv:2512.09665v2 Announce Type: replace Abstract: We address the problem of fair classification in settings where data is scarce and unbalanced across demographic groups. Such low-data regimes are c

PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

ResearchDGX agent

arXiv:2602.06912v2 Announce Type: replace Abstract: Unsupervised segmentation from self-supervised ViT patches holds promise but lacks robustness: multi-object scenes confound saliency cues, and low-s

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation

ResearchDGX agent

arXiv:2604.07901v1 Announce Type: new Abstract: 360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications,

ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models

AgentsDGX agent

arXiv:2604.07912v1 Announce Type: new Abstract: Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant ent

ParseBench: A Document Parsing Benchmark for AI Agents

Model ReleasesDGX agent

arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and

Part^{2}GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting

SafetyDGX agent

arXiv:2506.17212v2 Announce Type: replace Abstract: Articulated objects are common in the real world, yet modeling their structure and motion remains a challenging task for 3D reconstruction methods.

Personalizing Text-to-Image Generation to Individual Taste

SafetyDGX agent

arXiv:2604.07427v1 Announce Type: new Abstract: Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models opt

Phantasia: Context-Adaptive Backdoors in Vision Language Models

ResearchDGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

ApplicationsDGX agent

arXiv:2604.08503v1 Announce Type: new Abstract: Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

Model ReleasesDGX agent

arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

Model ReleasesDGX agent

arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe

Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting

TutorialsDGX agent

arXiv:2503.09640v2 Announce Type: replace-cross Abstract: Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applicat

PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization

Local AiDGX agent

arXiv:2503.24135v3 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) methods allow training models to classify images and localize ROIs. WSOL only requires low-cost image-c

Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models

SafetyDGX agent

arXiv:2604.07779v1 Announce Type: new Abstract: Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagn

PLUME: Latent Reasoning Based Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

SafetyDGX agent

arXiv:2604.08125v1 Announce Type: new Abstract: Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are l

Preventing Overfitting in Deep Image Prior for Hyperspectral Image Denoising

ResearchDGX agent

arXiv:2604.08272v1 Announce Type: new Abstract: Deep image prior (DIP) is an unsupervised deep learning framework that has been successfully applied to a variety of inverse imaging problems. However,

Privacy Attacks on Image AutoRegressive Models

ResearchDGX agent

arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m

PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation

Model ReleasesDGX agent

arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech

Pseudo-Expert Regularized Offline RL for End-to-End Autonomous Driving in Photorealistic Closed-Loop Environments

AgentsDGX agent

arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Model ReleasesDGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

Quantifying Explanation Consistency: The C-Score Metric for CAM-Based Explainability in Medical Image Classification

Local AiDGX agent

arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi

RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing

HardwareDGX agent

arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r

Reading Recognition in the Wild

TutorialsDGX agent

arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning

SafetyDGX agent

arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

ResearchDGX agent

arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

ResearchDGX agent

arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

Model ReleasesDGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

AgentsDGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

Revisiting Radar Perception With Spectral Point Clouds

Model ReleasesDGX agent

arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp

RewardFlow: Generate Images by Optimizing What You Reward

SafetyDGX agent

arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

SafetyDGX agent

arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli

Rotation Equivariant Convolutions in Deformable Registration of Brain MRI

SafetyDGX agent

arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a

← Previous
1…204205206207
Next →