AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

G^2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

DGX agent

arXiv:2605.12309v1 Announce Type: new Abstract: The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. I

researcharxiv-cs-cv
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization

DGX agent

arXiv:2605.12431v1 Announce Type: new Abstract: Conventional gait de-identification methods often encounter an inherent trade-off: they either provide insufficient identity suppression or introduce sp

researcharxiv-cs-cv
13 May 2026
Research

GATA2Floor: Graph attention for floor counting in street-view facades

DGX agent

arXiv:2605.11863v1 Announce Type: new Abstract: Automated analysis of building facades from street-level imagery has great potential for urban analytics, energy assessment, and emergency planning. How

researcharxiv-cs-cv
13 May 2026
Safety

Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences

DGX agent

arXiv:2605.11276v1 Announce Type: new Abstract: Highway construction workers face a high risk of serious injury or death. Image-based training materials depicting hazardous scenarios are essential for

safetyarxiv-cs-cv
13 May 2026
Research

GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction

DGX agent

arXiv:2605.12399v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to sever

researcharxiv-cs-cv
13 May 2026
Model Releases

GeoR-Bench: Evaluating Geoscience Visual Reasoning

DGX agent

arXiv:2605.11541v1 Announce Type: new Abstract: Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains s

model-releasesarxiv-cs-cv
13 May 2026
Safety

Gradient-Free Noise Optimization for Reward Alignment in Generative Models

DGX agent

arXiv:2605.11347v1 Announce Type: cross Abstract: Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to determi

safetyarxiv-cs-cv
13 May 2026
Safety

GRASP: Guided Residual Adapters with Sample-wise Partitioning

DGX agent

arXiv:2512.01675v2 Announce Type: replace Abstract: Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting thei

safetyarxiv-cs-cv
13 May 2026
Local Ai

Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances

DGX agent

arXiv:2605.11616v1 Announce Type: new Abstract: Functional affordance grounding requires more than recognizing an object: an agent must localize the specific region that supports an interaction, such

local-aiarxiv-cs-cv
13 May 2026
Research

h-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement

DGX agent

arXiv:2605.11871v1 Announce Type: new Abstract: Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video suppl

researcharxiv-cs-cv
13 May 2026
Research

H2G: Hierarchy-Aware Hyperbolic Grouping for 3D Scenes

DGX agent

arXiv:2605.11967v1 Announce Type: new Abstract: Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on sema

researcharxiv-cs-cv
13 May 2026
Local Ai

H3D-MarNet: Wavelet-Guided Dual-Path Learning for Metal Artifact Suppression and CT Modality Transformation for Radiotherapy Workflows

DGX agent

arXiv:2605.12252v1 Announce Type: new Abstract: Metal artifacts in computed tomography (CT) severely degrade image quality, compromising diagnostic accuracy and radiotherapy planning, especially in ca

local-aiarxiv-cs-cv
13 May 2026
Applications

HamBR: Active Decision Boundary Restoration Based on Hamiltonian Dynamics for Learning with Noisy Labels

DGX agent

arXiv:2605.11383v1 Announce Type: new Abstract: In large-scale visual recognition and data mining tasks, the presence of noisy labels severely undermines the generalization capability of deep neural n

applicationsarxiv-cs-cv
13 May 2026
Model Releases

Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation

DGX agent

arXiv:2605.11208v1 Announce Type: new Abstract: Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain chall

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

DGX agent

arXiv:2605.11061v1 Announce Type: new Abstract: The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In

model-releasesarxiv-cs-cv
13 May 2026
Research

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation

DGX agent

arXiv:2605.11596v1 Announce Type: new Abstract: Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR

researcharxiv-cs-cv
13 May 2026
Research

Hyperbolic Concept Bottleneck Models

DGX agent

arXiv:2605.06440v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier input

researcharxiv-cs-cv
13 May 2026
Safety

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

DGX agent

arXiv:2605.12305v1 Announce Type: new Abstract: While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggl

safetyarxiv-cs-cv
13 May 2026
Applications

Instruct-ICL: Instruction-Guided In-Context Learning for Post-Disaster Damage Assessment

DGX agent

arXiv:2605.11439v1 Announce Type: new Abstract: Rapid and accurate situational awareness is essential for effective response during natural disasters, where delays in analysis can significantly hinder

applicationsarxiv-cs-cv
13 May 2026
Research

Interactive Mars Image Content-Based Search with Interpretable Machine Learning

DGX agent

arXiv:2402.16860v2 Announce Type: replace Abstract: The NASA Planetary Data System (PDS) hosts millions of images of planets, moons, and other bodies collected throughout many missions. The ever-expan

researcharxiv-cs-cv
13 May 2026
Research

Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

DGX agent

arXiv:2605.11934v1 Announce Type: new Abstract: Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality indepen

researcharxiv-cs-cv
13 May 2026
Safety

JACoP: Joint Alignment for Compliant Multi-Agent Prediction

DGX agent

arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e

safetyarxiv-cs-cv
13 May 2026
Model Releases

KAN-CL: Per-Knot Importance Regularization for Continual Learning with Kolmogorov-Arnold Networks

DGX agent

arXiv:2605.12306v1 Announce Type: cross Abstract: Catastrophic forgetting remains the central obstacle in continual learning (CL): parameters shared across tasks interfere with one another, and existi

model-releasesarxiv-cs-cv
13 May 2026
Local Ai

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

DGX agent

arXiv:2605.11605v1 Announce Type: new Abstract: Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, ma

local-aiarxiv-cs-cv
13 May 2026
Tutorials

L2P: Unlocking Latent Potential for Pixel Generation

DGX agent

arXiv:2605.12013v1 Announce Type: new Abstract: Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohib

tutorialsarxiv-cs-cv
13 May 2026
Safety

LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition

DGX agent

arXiv:2603.29057v2 Announce Type: replace Abstract: Skeleton-based isolated sign language recognition (ISLR) demands fine-grained understanding of articulated motion across multiple spatial scales, fr

safetyarxiv-cs-cv
13 May 2026
Model Releases

Large-Small Model Collaboration for Farmland Semantic Change Detection

DGX agent

arXiv:2605.12282v1 Announce Type: new Abstract: Farmland Semantic Change Detection (SCD) is essential for cultivated land protection, yet existing benchmarks and models remain insufficient for fine-gr

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR

DGX agent

arXiv:2605.11115v1 Announce Type: new Abstract: High Dynamic Range (HDR) generation remains challenging for generative models, which are largely limited to low dynamic range outputs. Recent diffusionb

model-releasesarxiv-cs-cv
13 May 2026
Agents

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs

DGX agent

arXiv:2605.11477v1 Announce Type: new Abstract: Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token bud

agentsarxiv-cs-cv
13 May 2026
Safety

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

DGX agent

arXiv:2605.11931v1 Announce Type: new Abstract: Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acqui

safetyarxiv-cs-cv
13 May 2026
Model Releases

Learnable Multi-level Discrete Wavelet Transforms for 3D Gaussian Splatting Frequency Modulation

DGX agent

arXiv:2602.14199v2 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful approach for novel view synthesis. However, the number of Gaussian primitives often gro

model-releasesarxiv-cs-cv
13 May 2026
Research

Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

DGX agent

arXiv:2605.12218v1 Announce Type: new Abstract: Bird's-eye-view (BEV) representations derived from multi-camera input have become a central interface for online high-definition (HD) map construction.

researcharxiv-cs-cv
13 May 2026
Model Releases

Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data

DGX agent

arXiv:2605.11881v1 Announce Type: new Abstract: The high-dimensional features extracted from large-scale unlabeled data via various pretrained models with diverse architectures are referred to as hete

model-releasesarxiv-cs-cv
13 May 2026
Safety

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

DGX agent

arXiv:2605.11444v1 Announce Type: new Abstract: All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent me

safetyarxiv-cs-cv
13 May 2026
Model Releases

LiBrA-Net: Lie-Algebraic Bilateral Affine Fields for Real-Time 4K Video Dehazing

DGX agent

arXiv:2605.11508v1 Announce Type: new Abstract: Currently, there is a gap in the field of ultra-high-definition (UHD) video dehazing due to the lack of a benchmark for evaluation. Furthermore, existin

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction

DGX agent

arXiv:2605.11354v1 Announce Type: new Abstract: Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering str

model-releasesarxiv-cs-cv
13 May 2026
Safety

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration

DGX agent

arXiv:2605.11591v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in multi-image cross-modal retrieval, yet suffer from severe position bias, where

safetyarxiv-cs-cv
13 May 2026
Model Releases

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs

DGX agent

arXiv:2509.19207v2 Announce Type: replace Abstract: Contrastive vision-language models (VLMs) have made significant progress in binding visual and textual information, yet understanding long, composit

model-releasesarxiv-cs-cv
13 May 2026
Safety

Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making

DGX agent

arXiv:2602.07668v2 Announce Type: replace Abstract: The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state

safetyarxiv-cs-cv
13 May 2026
Safety

LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer

DGX agent

arXiv:2509.22414v4 Announce Type: replace Abstract: Image restoration (IR) aims to recover images degraded by unknown mixtures while preserving semanticsconditions under which discriminative restorers

safetyarxiv-cs-cv
13 May 2026
Safety

LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution

DGX agent

arXiv:2603.05947v3 Announce Type: replace Abstract: Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs

safetyarxiv-cs-cv
13 May 2026
Agents

LychSim: A Controllable and Interactive Simulation Framework for Vision Research

DGX agent

arXiv:2605.12449v1 Announce Type: new Abstract: While self-supervised pretraining has reduced vision systems' reliance on synthetic data, simulation remains an indispensable tool for closed-loop optim

agentsarxiv-cs-cv
13 May 2026
Applications

M^4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

DGX agent

arXiv:2605.11760v1 Announce Type: new Abstract: The Segment Anything Model 2 (SAM2) has emerged as a foundation model for universal segmentation. Owing to its generalizable visual representations, SAM

applicationsarxiv-cs-cv
13 May 2026
Research

MieDB-100k: A Comprehensive Dataset for Medical Image Editing

DGX agent

arXiv:2602.09587v2 Announce Type: replace Abstract: The scarcity of high-quality data remains a primary bottleneck in adapting multimodal generative models for medical image editing. Existing medical

researcharxiv-cs-cv
13 May 2026
Local Ai

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement

DGX agent

arXiv:2605.11808v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance on diverse vision-language tasks. However, LVLMs still suffer from hallucinati

local-aiarxiv-cs-cv
13 May 2026
Research

Mobile Traffic Camera Calibration from Road Geometry for UAV-Based Traffic Surveillance

DGX agent

arXiv:2605.11900v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) can provide flexible traffic surveillance where fixed roadside cameras are unavailable, costly, or impractical. However,

researcharxiv-cs-cv
13 May 2026
Safety

MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics

DGX agent

arXiv:2605.12119v1 Announce Type: new Abstract: Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view chan

safetyarxiv-cs-cv
13 May 2026
Research

Monitoring access to piped water and sanitation infrastructure in Africa at disaggregated scales using satellite imagery and self-supervised learning

DGX agent

arXiv:2411.19093v4 Announce Type: replace Abstract: Access to drinking water and sanitation services is essential for health and well-being, yet large global disparities persist. Sustainable Developme

researcharxiv-cs-cv
13 May 2026
← Previous
1…176177178179180…263
Next →