AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2604.08110v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging

researcharxiv-cs-cv
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance

DGX agent

arXiv:2604.08461v1 Announce Type: new Abstract: Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based a

safetyarxiv-cs-cv
10 Apr 2026
Safety

OxEnsemble: Fair Ensembles for Low-Data Classification

DGX agent

arXiv:2512.09665v2 Announce Type: replace Abstract: We address the problem of fair classification in settings where data is scarce and unbalanced across demographic groups. Such low-data regimes are c

safetyarxiv-cs-cv
10 Apr 2026
Research

PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

DGX agent

arXiv:2602.06912v2 Announce Type: replace Abstract: Unsupervised segmentation from self-supervised ViT patches holds promise but lacks robustness: multi-object scenes confound saliency cues, and low-s

researcharxiv-cs-cv
10 Apr 2026
Research

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation

DGX agent

arXiv:2604.07901v1 Announce Type: new Abstract: 360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications,

researcharxiv-cs-cv
10 Apr 2026
Agents

ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models

DGX agent

arXiv:2604.07912v1 Announce Type: new Abstract: Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant ent

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

ParseBench: A Document Parsing Benchmark for AI Agents

DGX agent

arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

Part^{2}GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting

DGX agent

arXiv:2506.17212v2 Announce Type: replace Abstract: Articulated objects are common in the real world, yet modeling their structure and motion remains a challenging task for 3D reconstruction methods.

safetyarxiv-cs-cv
10 Apr 2026
Safety

Personalizing Text-to-Image Generation to Individual Taste

DGX agent

arXiv:2604.07427v1 Announce Type: new Abstract: Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models opt

safetyarxiv-cs-cv
10 Apr 2026
Research

Phantasia: Context-Adaptive Backdoors in Vision Language Models

DGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

researcharxiv-cs-cv
10 Apr 2026
Applications

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

DGX agent

arXiv:2604.08503v1 Announce Type: new Abstract: Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

DGX agent

arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

DGX agent

arXiv:2603.23286v3 Announce Type: replace Abstract: Physical knot classification is a fine-grained task in which the intended cue is rope crossing structure, but high accuracy may still come from appe

model-releasesarxiv-cs-cv
10 Apr 2026
Tutorials

Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting

DGX agent

arXiv:2503.09640v2 Announce Type: replace-cross Abstract: Rendering realistic human-object interactions (HOIs) from sparse-view inputs is a challenging yet crucial task for various real-world applicat

tutorialsarxiv-cs-cv
10 Apr 2026
Local Ai

PixelCAM: Pixel Class Activation Mapping for Histology Image Classification and ROI Localization

DGX agent

arXiv:2503.24135v3 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) methods allow training models to classify images and localize ROIs. WSOL only requires low-cost image-c

local-aiarxiv-cs-cv
10 Apr 2026
Safety

Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models

DGX agent

arXiv:2604.07779v1 Announce Type: new Abstract: Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagn

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

PLUME: Latent Reasoning Based Universal Multimodal Embedding

DGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models

DGX agent

arXiv:2604.08340v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved remarkable progress in static visual understanding, their deployment in complex 3D embodied environmen

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

DGX agent

arXiv:2604.08125v1 Announce Type: new Abstract: Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are l

safetyarxiv-cs-cv
10 Apr 2026
Research

Preventing Overfitting in Deep Image Prior for Hyperspectral Image Denoising

DGX agent

arXiv:2604.08272v1 Announce Type: new Abstract: Deep image prior (DIP) is an unsupervised deep learning framework that has been successfully applied to a variety of inverse imaging problems. However,

researcharxiv-cs-cv
10 Apr 2026
Research

Privacy Attacks on Image AutoRegressive Models

DGX agent

arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m

researcharxiv-cs-cv
10 Apr 2026
Model Releases

PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation

DGX agent

arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech

model-releasesarxiv-cs-cv
10 Apr 2026
Agents

Pseudo-Expert Regularized Offline RL for End-to-End Autonomous Driving in Photorealistic Closed-Loop Environments

DGX agent

arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

DGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

model-releasesarxiv-cs-cv
10 Apr 2026
Local Ai

Quantifying Explanation Consistency: The C-Score Metric for CAM-Based Explainability in Medical Image Classification

DGX agent

arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi

local-aiarxiv-cs-cv
10 Apr 2026
Hardware

RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing

DGX agent

arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r

hardwarearxiv-cs-cv
10 Apr 2026
Tutorials

Reading Recognition in the Wild

DGX agent

arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,

tutorialsarxiv-cs-cv
10 Apr 2026
Safety

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning

DGX agent

arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc

safetyarxiv-cs-cv
10 Apr 2026
Research

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

DGX agent

arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p

researcharxiv-cs-cv
10 Apr 2026
Research

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

DGX agent

arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

DGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

model-releasesarxiv-cs-cv
10 Apr 2026
Agents

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

DGX agent

arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

Revisiting Radar Perception With Spectral Point Clouds

DGX agent

arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

RewardFlow: Generate Images by Optimizing What You Reward

DGX agent

arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La

safetyarxiv-cs-cv
10 Apr 2026
Safety

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

DGX agent

arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli

safetyarxiv-cs-cv
10 Apr 2026
Safety

Rotation Equivariant Convolutions in Deformable Registration of Brain MRI

DGX agent

arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a

safetyarxiv-cs-cv
10 Apr 2026
Agents

RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection

DGX agent

arXiv:2505.17732v2 Announce Type: replace Abstract: Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have eme

agentsarxiv-cs-cv
10 Apr 2026
Research

Sampling-Aware 3D Spatial Analysis in Multiplexed Imaging

DGX agent

arXiv:2604.07890v1 Announce Type: new Abstract: Highly multiplexed microscopy enables rich spatial characterization of tissues at single-cell resolution, yet most analyses rely on two-dimensional sect

researcharxiv-cs-cv
10 Apr 2026
Local Ai

SAT: Selective Aggregation Transformer for Image Super-Resolution

DGX agent

arXiv:2604.07994v1 Announce Type: new Abstract: Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational comple

local-aiarxiv-cs-cv
10 Apr 2026
Local Ai

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

DGX agent

arXiv:2604.08542v1 Announce Type: new Abstract: This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown pro

local-aiarxiv-cs-cv
10 Apr 2026
Agents

Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

DGX agent

arXiv:2604.08366v1 Announce Type: cross Abstract: Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, th

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

DGX agent

arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection

DGX agent

arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research

model-releasesarxiv-cs-cv
10 Apr 2026
Research

SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

DGX agent

arXiv:2604.03134v2 Announce Type: replace Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing

researcharxiv-cs-cv
10 Apr 2026
Model Releases

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

DGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Self-Improving 4D Perception via Self-Distillation

DGX agent

arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr

researcharxiv-cs-cv
10 Apr 2026
Safety

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

DGX agent

arXiv:2604.08147v1 Announce Type: cross Abstract: Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, j

safetyarxiv-cs-cv
10 Apr 2026
Safety

SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP

DGX agent

arXiv:2509.26036v3 Announce Type: replace Abstract: While Contrastive Language-Image Pretraining (CLIP) excels at zero-shot tasks by aligning image and text embeddings, its performance in few-shot cla

safetyarxiv-cs-cv
10 Apr 2026
← Previous
1…260261262263
Next →