AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
4 May 2026

Information-geometric adaptive sampling for graph diffusion

ResearchDGX agent

arXiv:2605.00250v1 Announce Type: cross Abstract: Standard diffusion models for graph generation typically rely on uniform time-stepping, an approach that overlooks the non-homogeneous dynamics of dis

InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization

SafetyDGX agent

arXiv:2605.00664v1 Announce Type: new Abstract: We present a training-free approach for controllable 3D inpainting based on initial noise optimization. In the structured 3D latent diffusion framework,

Instance-Aware Pseudo-Labeling and Class-Focused Contrastive Learning for Weakly Supervised Domain Adaptive Segmentation of Electron Microscopy

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.16450v2 Announce Type: replace Abstract: Annotation-efficient segmentation of the numerous mitochondria instances from various electron microscopy (EM) images is highly valuable for biologi

Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision-Language Models

ResearchDGX agent

arXiv:2605.00591v1 Announce Type: new Abstract: Contrastive vision-language models like CLIP exhibit remarkable zero-shot generalization. However, prompt tuning remains highly sensitive to label noise

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

ResearchDGX agent

arXiv:2601.00090v2 Announce Type: replace Abstract: Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prom

Jailbreaking Vision-Language Models Through the Visual Modality

Model ReleasesDGX agent

arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta

LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping

ResearchDGX agent

arXiv:2511.08156v2 Announce Type: replace Abstract: Land Use and Land Cover (LULC) mapping is a fundamental task in Earth Observation (EO). However, current LULC models are typically developed for a s

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

SafetyDGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

Learning Coarse-to-Fine Osteoarthritis Representations under Noisy Hierarchical Labels

SafetyDGX agent

arXiv:2605.00718v1 Announce Type: new Abstract: Knee osteoarthritis (OA) assessment involves a natural but often underused label hierarchy: a coarse binary OA decision and a fine-grained Kellgren--Law

Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

Model ReleasesDGX agent

arXiv:2605.00448v1 Announce Type: new Abstract: The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetri

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation

Model ReleasesDGX agent

arXiv:2605.00051v1 Announce Type: new Abstract: Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

Model ReleasesDGX agent

arXiv:2412.00452v2 Announce Type: replace-cross Abstract: Conventioanl federated learning (FL) heavily depends on high-quality labels, which are often impractical in the real world, leading to the fed

Learning physically grounded traffic accident reconstruction from public accident reports

SafetyDGX agent

arXiv:2605.00050v1 Announce Type: cross Abstract: Traffic accidents are routinely documented in textual reports, yet physically grounded accident reconstruction remains difficult because detailed scen

Let ViT Speak: Generative Language-Image Pre-training

ResearchDGX agent

arXiv:2605.00809v1 Announce Type: new Abstract: In this paper, we present extbf{Gen}erative extbf{L}anguage-extbf{I}mage extbf{P}re-training (GenLIP), a minimalist generative pretraining framework for

Leveraging Vision-Language Models as Weak Annotators in Active Learning

ResearchDGX agent

arXiv:2605.00480v1 Announce Type: new Abstract: Active learning aims to reduce annotation cost by selectively querying informative samples for supervision under a limited labeling budget. In this work

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

ApplicationsDGX agent

arXiv:2605.00434v1 Announce Type: new Abstract: Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation

Local AiDGX agent

arXiv:2605.00244v1 Announce Type: cross Abstract: We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking multi-modal data to train real-world robotic systems. At th

MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video

TutorialsDGX agent

arXiv:2605.00242v1 Announce Type: new Abstract: Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation. However, existing methods typically rely

Make Your LVLM KV Cache More Lightweight

Model ReleasesDGX agent

arXiv:2605.00789v1 Announce Type: new Abstract: Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency

Map2World: Segment Map Conditioned Text to 3D World Generation

AgentsDGX agent

arXiv:2605.00781v1 Announce Type: new Abstract: 3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world gener

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

ApplicationsDGX agent

arXiv:2605.00495v1 Announce Type: cross Abstract: Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound producti

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

ResearchDGX agent

arXiv:2605.00431v1 Announce Type: cross Abstract: Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model ro

Modeling Subjective Urban Perception with Human Gaze

ResearchDGX agent

arXiv:2605.00764v1 Announce Type: new Abstract: Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computationa

MSACT: Multistage Spatial Alignment for Stable Low-Latency Fine Manipulation

Local AiDGX agent

arXiv:2605.00475v1 Announce Type: cross Abstract: Real-world fine manipulation, particularly in bimanual manipulation, typically requires low-latency control and stable visual localization, while coll

Multi-frame Restoration for High-rate Lissajous Confocal Laser Endomicroscopy

Model ReleasesDGX agent

arXiv:2605.00527v1 Announce Type: cross Abstract: Lissajous confocal laser endomicroscopy (CLE) is a promising solution for high speed in vivo optical biopsy for handheld scenarios. However, Lissajous

Online Self-Calibration Against Hallucination in Vision-Language Models

SafetyDGX agent

arXiv:2605.00323v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image.

Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement

Model ReleasesDGX agent

arXiv:2605.00634v1 Announce Type: cross Abstract: We introduce Paired-CSLiDAR (CSLiDAR), a cross-source aerial-ground LiDAR benchmark for single-scan pose refinement: refining a ground-scan pose withi

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

Model ReleasesDGX agent

arXiv:2605.00814v1 Announce Type: new Abstract: While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a 'Visual Signal Dilution' p

PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning

ResearchDGX agent

arXiv:2605.00707v1 Announce Type: new Abstract: Image editing instructions are heterogeneous: a color swap, an object insertion, and a physical-action edit all demand different spatial coverage and di

PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation

ResearchDGX agent

arXiv:2605.00517v1 Announce Type: new Abstract: Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Nota

Pose-Aware Diffusion for 3D Generation

SafetyDGX agent

arXiv:2605.00345v1 Announce Type: new Abstract: Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rota

Possibilistic Predictive Uncertainty for Deep Learning

ResearchDGX agent

arXiv:2605.00600v1 Announce Type: cross Abstract: Deep neural networks achieve impressive results across diverse applications, yet their overconfidence on unseen inputs necessitates reliable epistemic

Posterior Augmented Flow Matching

ResearchDGX agent

arXiv:2605.00825v1 Announce Type: new Abstract: Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-di

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

SafetyDGX agent

arXiv:2411.02327v4 Announce Type: replace Abstract: In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long vi

Prediction of Alzheimer's Disease Risk Factors from Retinal Images via Deep Learning: Development and Validation of Biologically Relevant Morphological Associations in the UK Biobank

ApplicationsDGX agent

arXiv:2605.00665v1 Announce Type: new Abstract: The systemic, metabolic, lifestyle factors have established associations with Alzheimer's Disease (AD) through epidemiologic and AD-specific biomarker s

Prefer-DAS: Learning from Local Preferences and Sparse Prompts for Domain Adaptive Segmentation of Electron Microscopy

Local AiDGX agent

arXiv:2602.19423v3 Announce Type: replace Abstract: Domain adaptive segmentation (DAS) is a promising paradigm for delineating intracellular structures from various large-scale electron microscopy (EM

Quantum Gradient-Based Approach for Edge and Corner Detection Using Sobel Kernels

ResearchDGX agent

arXiv:2605.00744v1 Announce Type: new Abstract: Edge detection refers to identifying points in a digital image where intensity changes sharply, indicating object boundaries or structural features. Cor

Real-Time Frame- and Event-based Object Detection with Spiking Neural Networks on Edge Neuromorphic Hardware: Design, Deployment and Benchmark

Model ReleasesDGX agent

arXiv:2605.00146v1 Announce Type: new Abstract: Real-time object detection on energy-constrained platforms is critical for applications such as UAV-based inspection, autonomous navigation, and mobile

REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception

ResearchDGX agent

arXiv:2605.00271v1 Announce Type: new Abstract: Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to ex

Remote SAMsing: From Segment Anything to Segment Everything

Model ReleasesDGX agent

arXiv:2605.00256v1 Announce Type: new Abstract: SAM2 produces high-quality zero-shot segmentation on natural images, but applying it to large remote sensing scenes exposes two problems: (1) its mask g

Robust Fusion of Object-Level V2X for Learned 3D Object Detection

ApplicationsDGX agent

arXiv:2605.00595v1 Announce Type: new Abstract: Perception for automated driving is largely based on onboard environmental sensors, such as cameras and radar, which are cost-effective but limited by l

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

Model ReleasesDGX agent

arXiv:2605.00392v1 Announce Type: new Abstract: DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundan

Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems

ResearchDGX agent

arXiv:2605.00510v1 Announce Type: cross Abstract: Complex physical systems, from supersonic turbulence to the macroscopic structure of the universe, are governed by continuous multiscale dynamics. Whi

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

SafetyDGX agent

arXiv:2605.00444v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded percept

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

Model ReleasesDGX agent

arXiv:2602.14276v2 Announce Type: replace Abstract: Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain

SIMON: Saliency-aware Integrative Multi-view Object-centric Neural Decoding

SafetyDGX agent

arXiv:2605.00401v1 Announce Type: new Abstract: Recent EEG-to-image retrieval methods leverage pretrained vision encoders and foveation-inspired priors, but typically assume a fixed, center-focused vi

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation

ResearchDGX agent

arXiv:2505.18875v4 Announce Type: replace Abstract: Diffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention. B

Static and Dynamic Graph Alignment Network for Temporal Video Grounding

Model ReleasesDGX agent

arXiv:2605.00684v1 Announce Type: new Abstract: Temporal Video Grounding (TVG) aims to localize temporal moments in an untrimmed video that semantically correspond to given natural language queries. R

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas

ResearchDGX agent

arXiv:2603.28980v2 Announce Type: replace Abstract: The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with

Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis

ResearchDGX agent

arXiv:2512.01116v3 Announce Type: replace Abstract: The integration of histology images and gene profiles has shown great promise for improving survival prediction in cancer. However, current approach

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor

ResearchDGX agent

arXiv:2601.09896v4 Announce Type: replace-cross Abstract: Visual generative AI models are trained using a one-size-fits-all measure of aesthetic appeal. However, what is deemed 'aesthetic' is inextric

The Determinism of Randomness: Latent Space Degeneracy in Diffusion Model

SafetyDGX agent

arXiv:2511.07756v4 Announce Type: replace Abstract: Diffusion models initialize generation from an isotropic Gaussian latent, yet changing only the random seed can substantially alter prompt faithfuln

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

AgentsDGX agent

arXiv:2602.06037v4 Announce Type: replace Abstract: Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs

ResearchDGX agent

arXiv:2506.11989v3 Announce Type: replace Abstract: Test-time scaling offers a promising way to improve the reasoning performance of vision-language large models (VLLMs) without additional training. I

Time-series Meets Complex Motion Modeling: Robust and Computational-effective Motion Predictor for Multi-object Tracking

AgentsDGX agent

arXiv:2605.00362v1 Announce Type: new Abstract: Multi-object tracking (MOT) is critical in numerous real-world applications, including surveillance, autonomous driving, and robotics. Accurately predic

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

ApplicationsDGX agent

arXiv:2605.00015v1 Announce Type: cross Abstract: Time Series Foundation Models (TSFMs) advance generalization and data efficiency in time series forecasting by unified large-scale pretraining. But TS

Two-View Accumulation as the Primary Training Lever for Hybrid-Capture Gaussian Splatting: A Variance-Decomposition View of When Gradient Surgery Helps

ResearchDGX agent

arXiv:2605.00052v1 Announce Type: new Abstract: Hybrid-capture novel view synthesis combines images at substantially different camera distances (e.g., aerial drone and ground-level views). Standard 3D

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

SafetyDGX agent

arXiv:2605.00658v1 Announce Type: new Abstract: Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often tr

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards

SafetyDGX agent

arXiv:2510.00072v2 Announce Type: replace Abstract: Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. W

Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategy

SafetyDGX agent

arXiv:2605.00719v1 Announce Type: new Abstract: Unsupervised deraining has attracted attention for its ability to learn the real-world distribution of rain without paired supervision. However, the lac

← Previous
1…161162163164165…209
Next →