AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Multi-frame Restoration for High-rate Lissajous Confocal Laser Endomicroscopy

DGX agent

arXiv:2605.00527v1 Announce Type: cross Abstract: Lissajous confocal laser endomicroscopy (CLE) is a promising solution for high speed in vivo optical biopsy for handheld scenarios. However, Lissajous

model-releasesarxiv-cs-cv
4 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Online Self-Calibration Against Hallucination in Vision-Language Models

DGX agent

arXiv:2605.00323v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image.

safetyarxiv-cs-cv
4 May 2026
Model Releases

Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement

DGX agent

arXiv:2605.00634v1 Announce Type: cross Abstract: We introduce Paired-CSLiDAR (CSLiDAR), a cross-source aerial-ground LiDAR benchmark for single-scan pose refinement: refining a ground-scan pose withi

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

DGX agent

arXiv:2605.00814v1 Announce Type: new Abstract: While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a 'Visual Signal Dilution' p

model-releasesarxiv-cs-cv
4 May 2026
Research

PhysEdit: Physically-Consistent Region-Aware Image Editing via Adaptive Spatio-Temporal Reasoning

DGX agent

arXiv:2605.00707v1 Announce Type: new Abstract: Image editing instructions are heterogeneous: a color swap, an object insertion, and a physical-action edit all demand different spatial coverage and di

researcharxiv-cs-cv
4 May 2026
Research

PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation

DGX agent

arXiv:2605.00517v1 Announce Type: new Abstract: Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Nota

researcharxiv-cs-cv
4 May 2026
Safety

Pose-Aware Diffusion for 3D Generation

DGX agent

arXiv:2605.00345v1 Announce Type: new Abstract: Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rota

safetyarxiv-cs-cv
4 May 2026
Research

Possibilistic Predictive Uncertainty for Deep Learning

DGX agent

arXiv:2605.00600v1 Announce Type: cross Abstract: Deep neural networks achieve impressive results across diverse applications, yet their overconfidence on unseen inputs necessitates reliable epistemic

researcharxiv-cs-cv
4 May 2026
Research

Posterior Augmented Flow Matching

DGX agent

arXiv:2605.00825v1 Announce Type: new Abstract: Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-di

researcharxiv-cs-cv
4 May 2026
Safety

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

DGX agent

arXiv:2411.02327v4 Announce Type: replace Abstract: In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long vi

safetyarxiv-cs-cv
4 May 2026
Applications

Prediction of Alzheimer's Disease Risk Factors from Retinal Images via Deep Learning: Development and Validation of Biologically Relevant Morphological Associations in the UK Biobank

DGX agent

arXiv:2605.00665v1 Announce Type: new Abstract: The systemic, metabolic, lifestyle factors have established associations with Alzheimer's Disease (AD) through epidemiologic and AD-specific biomarker s

applicationsarxiv-cs-cv
4 May 2026
Local Ai

Prefer-DAS: Learning from Local Preferences and Sparse Prompts for Domain Adaptive Segmentation of Electron Microscopy

DGX agent

arXiv:2602.19423v3 Announce Type: replace Abstract: Domain adaptive segmentation (DAS) is a promising paradigm for delineating intracellular structures from various large-scale electron microscopy (EM

local-aiarxiv-cs-cv
4 May 2026
Research

Quantum Gradient-Based Approach for Edge and Corner Detection Using Sobel Kernels

DGX agent

arXiv:2605.00744v1 Announce Type: new Abstract: Edge detection refers to identifying points in a digital image where intensity changes sharply, indicating object boundaries or structural features. Cor

researcharxiv-cs-cv
4 May 2026
Model Releases

Real-Time Frame- and Event-based Object Detection with Spiking Neural Networks on Edge Neuromorphic Hardware: Design, Deployment and Benchmark

DGX agent

arXiv:2605.00146v1 Announce Type: new Abstract: Real-time object detection on energy-constrained platforms is critical for applications such as UAV-based inspection, autonomous navigation, and mobile

model-releasesarxiv-cs-cv
4 May 2026
Research

REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception

DGX agent

arXiv:2605.00271v1 Announce Type: new Abstract: Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to ex

researcharxiv-cs-cv
4 May 2026
Model Releases

Remote SAMsing: From Segment Anything to Segment Everything

DGX agent

arXiv:2605.00256v1 Announce Type: new Abstract: SAM2 produces high-quality zero-shot segmentation on natural images, but applying it to large remote sensing scenes exposes two problems: (1) its mask g

model-releasesarxiv-cs-cv
4 May 2026
Applications

Robust Fusion of Object-Level V2X for Learned 3D Object Detection

DGX agent

arXiv:2605.00595v1 Announce Type: new Abstract: Perception for automated driving is largely based on onboard environmental sensors, such as cameras and radar, which are cost-effective but limited by l

applicationsarxiv-cs-cv
4 May 2026
Model Releases

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

DGX agent

arXiv:2605.00392v1 Announce Type: new Abstract: DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundan

model-releasesarxiv-cs-cv
4 May 2026
Research

Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems

DGX agent

arXiv:2605.00510v1 Announce Type: cross Abstract: Complex physical systems, from supersonic turbulence to the macroscopic structure of the universe, are governed by continuous multiscale dynamics. Whi

researcharxiv-cs-cv
4 May 2026
Safety

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

DGX agent

arXiv:2605.00444v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded percept

safetyarxiv-cs-cv
4 May 2026
Model Releases

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

DGX agent

arXiv:2602.14276v2 Announce Type: replace Abstract: Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain

model-releasesarxiv-cs-cv
4 May 2026
Safety

SIMON: Saliency-aware Integrative Multi-view Object-centric Neural Decoding

DGX agent

arXiv:2605.00401v1 Announce Type: new Abstract: Recent EEG-to-image retrieval methods leverage pretrained vision encoders and foveation-inspired priors, but typically assume a fixed, center-focused vi

safetyarxiv-cs-cv
4 May 2026
Research

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation

DGX agent

arXiv:2505.18875v4 Announce Type: replace Abstract: Diffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention. B

researcharxiv-cs-cv
4 May 2026
Model Releases

Static and Dynamic Graph Alignment Network for Temporal Video Grounding

DGX agent

arXiv:2605.00684v1 Announce Type: new Abstract: Temporal Video Grounding (TVG) aims to localize temporal moments in an untrimmed video that semantically correspond to given natural language queries. R

model-releasesarxiv-cs-cv
4 May 2026
Research

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas

DGX agent

arXiv:2603.28980v2 Announce Type: replace Abstract: The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with

researcharxiv-cs-cv
4 May 2026
Research

Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis

DGX agent

arXiv:2512.01116v3 Announce Type: replace Abstract: The integration of histology images and gene profiles has shown great promise for improving survival prediction in cancer. However, current approach

researcharxiv-cs-cv
4 May 2026
Research

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor

DGX agent

arXiv:2601.09896v4 Announce Type: replace-cross Abstract: Visual generative AI models are trained using a one-size-fits-all measure of aesthetic appeal. However, what is deemed 'aesthetic' is inextric

researcharxiv-cs-cv
4 May 2026
Safety

The Determinism of Randomness: Latent Space Degeneracy in Diffusion Model

DGX agent

arXiv:2511.07756v4 Announce Type: replace Abstract: Diffusion models initialize generation from an isotropic Gaussian latent, yet changing only the random seed can substantially alter prompt faithfuln

safetyarxiv-cs-cv
4 May 2026
Agents

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

DGX agent

arXiv:2602.06037v4 Announce Type: replace Abstract: Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However

agentsarxiv-cs-cv
4 May 2026
Research

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs

DGX agent

arXiv:2506.11989v3 Announce Type: replace Abstract: Test-time scaling offers a promising way to improve the reasoning performance of vision-language large models (VLLMs) without additional training. I

researcharxiv-cs-cv
4 May 2026
Agents

Time-series Meets Complex Motion Modeling: Robust and Computational-effective Motion Predictor for Multi-object Tracking

DGX agent

arXiv:2605.00362v1 Announce Type: new Abstract: Multi-object tracking (MOT) is critical in numerous real-world applications, including surveillance, autonomous driving, and robotics. Accurately predic

agentsarxiv-cs-cv
4 May 2026
Applications

TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

DGX agent

arXiv:2605.00015v1 Announce Type: cross Abstract: Time Series Foundation Models (TSFMs) advance generalization and data efficiency in time series forecasting by unified large-scale pretraining. But TS

applicationsarxiv-cs-cv
4 May 2026
Research

Two-View Accumulation as the Primary Training Lever for Hybrid-Capture Gaussian Splatting: A Variance-Decomposition View of When Gradient Surgery Helps

DGX agent

arXiv:2605.00052v1 Announce Type: new Abstract: Hybrid-capture novel view synthesis combines images at substantially different camera distances (e.g., aerial drone and ground-level views). Standard 3D

researcharxiv-cs-cv
4 May 2026
Safety

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

DGX agent

arXiv:2605.00658v1 Announce Type: new Abstract: Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often tr

safetyarxiv-cs-cv
4 May 2026
Safety

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards

DGX agent

arXiv:2510.00072v2 Announce Type: replace Abstract: Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. W

safetyarxiv-cs-cv
4 May 2026
Safety

Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategy

DGX agent

arXiv:2605.00719v1 Announce Type: new Abstract: Unsupervised deraining has attracted attention for its ability to learn the real-world distribution of rain without paired supervision. However, the lac

safetyarxiv-cs-cv
4 May 2026
Research

Unsupervised Denoising of Real Clinical Low Dose Liver CT with Perceptual Attention Networks

DGX agent

arXiv:2605.00793v1 Announce Type: cross Abstract: With the development of deep learning, medical image processing has been widely used to assist clinical research. This paper focuses on the denoising

researcharxiv-cs-cv
4 May 2026
Research

VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image

DGX agent

arXiv:2602.04349v2 Announce Type: replace Abstract: 3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches predomina

researcharxiv-cs-cv
4 May 2026
Model Releases

Vesselpose: Vessel Graph Reconstruction from Learned Voxel-wise Direction Vectors in 3D Vascular Images

DGX agent

arXiv:2605.00538v1 Announce Type: new Abstract: Blood vessel segmentation and -tracing are essential tasks in many medical imaging applications. Although numerous methods exist, the prevailing segment

model-releasesarxiv-cs-cv
4 May 2026
Research

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding

DGX agent

arXiv:2603.22285v2 Announce Type: replace Abstract: Long video understanding remains challenging for multimodal large language models (MLLMs) due to limited context windows, which necessitate identify

researcharxiv-cs-cv
4 May 2026
Local Ai

VkSplat: High-Performance 3DGS Training in Vulkan Compute

DGX agent

arXiv:2605.00219v1 Announce Type: new Abstract: We present VkSplat, a high-performance, cross-vendor 3D Gaussian Splatting (3DGS) training pipeline implemented fully in Vulkan compute, addressing perf

local-aiarxiv-cs-cv
4 May 2026
Tutorials

When Do Diffusion Models learn to Generate Multiple Objects?

DGX agent

arXiv:2605.00273v1 Announce Type: new Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical ev

tutorialsarxiv-cs-cv
4 May 2026
Research

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery

DGX agent

arXiv:2602.13305v2 Announce Type: replace Abstract: Wildfires are a growing threat to ecosystems, human lives, and infrastructure, with their frequency and intensity rising due to climate change and h

researcharxiv-cs-cv
4 May 2026
Safety

World Model for Robot Learning: A Comprehensive Survey

DGX agent

arXiv:2605.00080v1 Announce Type: cross Abstract: World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They s

safetyarxiv-cs-cv
4 May 2026
Applications

3D Reconstruction Techniques in the Manufacturing Domain: Applications, Research Opportunities and Use Cases

DGX agent

arXiv:2604.28064v1 Announce Type: new Abstract: This comprehensive review examines the evolution and the current state of the art in three-dimensional (3D) reconstruction techniques in manufacturing a

applicationsarxiv-cs-cv
1 May 2026
Research

3D-ReGen: A Unified 3D Geometry Regeneration Framework

DGX agent

arXiv:2604.28134v1 Announce Type: new Abstract: We consider the problem of regenerating 3D objects from 2D images and initial 3D shapes. Most 3D generators operate in a one-shot fashion, converting te

researcharxiv-cs-cv
1 May 2026
Tutorials

A generalised pre-training strategy for deep learning networks in semantic segmentation of remotely sensed images

DGX agent

arXiv:2604.27704v1 Announce Type: new Abstract: In the segmentation of remotely sensed images, deep learning models are typically pre-trained using large image databases like ImageNet before fine-tune

tutorialsarxiv-cs-cv
1 May 2026
Research

A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

DGX agent

arXiv:2604.27383v1 Announce Type: cross Abstract: Nasotracheal intubation (NTI) is a critical clinical procedure for establishing and maintaining patient airway patency. Machine-assisted NTI has emerg

researcharxiv-cs-cv
1 May 2026
← Previous
1…202203204205206…261
Next →