AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Applications

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

DGX agent

arXiv:2608.01906v1 Announce Type: new Abstract: Rapid and accurate post-disaster building damage assessment is essential, yet remains a challenging task. Unmanned Aerial Vehicle (UAV) imagery offers a

applicationsarxiv-cs-cv
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

DGX agent

arXiv:2608.02006v1 Announce Type: new Abstract: Dynamic 3D scene reconstruction has achieved remarkable success under the assumption of strictly synchronized multi-camera inputs. However, in real-worl

safetyarxiv-cs-cv
4 Aug 2026
Tutorials

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

DGX agent

arXiv:2608.01276v1 Announce Type: new Abstract: Full-body capture from unconstrained photographs requires global correspondence across arbitrary views, poses, crops, and occlusions. Yet pose, geometry

tutorialsarxiv-cs-cv
4 Aug 2026
Research

Attention-Steered Vision-Language Models for Sign Language Translation

DGX agent

arXiv:2608.00235v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice

DGX agent

arXiv:2608.00145v1 Announce Type: cross Abstract: Accurate estimation of left ventricular (LV) orientation is essential for cardiac magnetic resonance (CMR) imaging and downstream analysis. Existing m

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Belief-Space Perception Routing under Coupled Sensor Faults and Compute Contention

DGX agent

arXiv:2608.00322v1 Announce Type: cross Abstract: A robot that has to see and react on a fixed clock runs into two problems at once. Its cameras degrade in rain, mud, fog, and darkness. And the single

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

DGX agent

arXiv:2608.00077v1 Announce Type: new Abstract: Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this proto

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Beyond Edge Maps: Wavelet-Domain Conditioning for Multi-Adapter Map-to-Satellite Diffusion

DGX agent

arXiv:2608.00083v1 Announce Type: new Abstract: Commercial mapping partnerships are often unavailable in low-resource regions, leaving satellite basemaps stale and motivating synthesis of satellite im

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

DGX agent

arXiv:2608.02016v1 Announce Type: new Abstract: Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active

local-aiarxiv-cs-cv
4 Aug 2026
Research

Beyond Illumination: A Conditional Mutual Information-Guided Network for Low-Light Image Enhancement

DGX agent

arXiv:2608.01886v1 Announce Type: new Abstract: Low-light image enhancement (LLIE) seeks to restore structural fidelity, natural color rendition, and proper exposure from images captured under inadequ

researcharxiv-cs-cv
4 Aug 2026
Safety

Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging

DGX agent

arXiv:2608.00073v1 Announce Type: new Abstract: Rigorous dataset partitioning is a foundational, yet frequently overlooked, prerequisite for reliable deep learning in longitudinal medical imaging. Nai

safetyarxiv-cs-cv
4 Aug 2026
Local Ai

Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection

DGX agent

arXiv:2608.00442v1 Announce Type: new Abstract: Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and modalities. Exi

local-aiarxiv-cs-cv
4 Aug 2026
Research

Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection

DGX agent

arXiv:2608.01302v1 Announce Type: new Abstract: State-of-the-art RGB-Event detectors improve the detection of small, fast-moving objects by combining complementary features from RGB and Event data, ye

researcharxiv-cs-cv
4 Aug 2026
Research

Beyond Token-Level Cross-Entropy: Frechet Distributional Post-Training for Autoregressive Image Generation

DGX agent

arXiv:2608.00562v1 Announce Type: new Abstract: Autoregressive image generators are commonly pretrained with token-level cross-entropy under teacher forcing, yet evaluated by the distributional qualit

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

DGX agent

arXiv:2608.00415v1 Announce Type: new Abstract: Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various en

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection

DGX agent

arXiv:2601.02837v2 Announce Type: replace Abstract: Infrared small target detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, extremely small target sizes, and complex cl

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

DGX agent

arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, ev

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

Breaking the Statistical Similarity Trap in Extreme Convection Detection

DGX agent

arXiv:2509.09195v2 Announce Type: replace-cross Abstract: Current evaluation metrics for deep learning weather models create a 'Statistical Similarity Trap', rewarding blurry predictions while missing

safetyarxiv-cs-cv
4 Aug 2026
Research

BRIC-Net: Boundary-Reliable Illumination-Color Interaction for Remote Sensing Image Deshadowing

DGX agent

arXiv:2608.00682v1 Announce Type: new Abstract: Shadows in remote sensing images obscure surface appearance and disrupt radiometric continuity, reducing the reliability of visual interpretation and do

researcharxiv-cs-cv
4 Aug 2026
Model Releases

CADENA: Stepwise CAD Reverse Engineering

DGX agent

arXiv:2608.00799v1 Announce Type: new Abstract: Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. M

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

DGX agent

arXiv:2608.02309v1 Announce Type: new Abstract: We present CalibBEV, a novel Bird's Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shar

safetyarxiv-cs-cv
4 Aug 2026
Research

Calibrated Similarity and Graph Clustering for Open-Set Animal Re-Identification

DGX agent

arXiv:2608.02469v1 Announce Type: new Abstract: AnimalCLEF26 addresses discovery-oriented animal re-identification, where systems must both attach query images to known individuals and discover unseen

researcharxiv-cs-cv
4 Aug 2026
Applications

Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit

DGX agent

arXiv:2608.01753v1 Announce Type: new Abstract: Addressing urban blight has seen increased focus in the past 15 years. Assessing urban blight is essential for guiding urban planning, targeting rehabil

applicationsarxiv-cs-cv
4 Aug 2026
Model Releases

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

DGX agent

arXiv:2608.02589v1 Announce Type: new Abstract: Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the c

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views

DGX agent

arXiv:2506.18164v2 Announce Type: replace Abstract: Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such

local-aiarxiv-cs-cv
4 Aug 2026
Research

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation

DGX agent

arXiv:2602.10880v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have shown promise in generating plotting code from chart images, yet achieving structural fidelity remains challengin

researcharxiv-cs-cv
4 Aug 2026
Safety

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

DGX agent

arXiv:2608.00769v1 Announce Type: new Abstract: One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

CHOW-SLAM: Compact Hybrid Representation with Complementary Overlap Window Optimization for RGB-D SLAM

DGX agent

arXiv:2608.01914v1 Announce Type: new Abstract: Simultaneous localization and mapping (SLAM) based on Neural Radiance Fields (NeRF) enables dense, continuous scene reconstruction. However, existing sy

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

CLEAR: Conflict-aware Learning via Evidence-guided Adaptive Routing for Unified Sparse-View 3D Gaussian Super-Resolution

DGX agent

arXiv:2608.02206v1 Announce Type: new Abstract: Sparse-view 3D Gaussian Splatting Super-resolution is highly challenging since the sparse and low-resolution (LR) inputs lack sufficient geometric and h

applicationsarxiv-cs-cv
4 Aug 2026
Research

Clear-Weighted Bit Allocation for Satellite Downlinks

DGX agent

arXiv:2608.01457v1 Announce Type: cross Abstract: Earth-observation satellites capture more imagery than intermittent ground contacts can transmit. Onboard systems threshold a cloud detector, discard

researcharxiv-cs-cv
4 Aug 2026
Research

Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild

DGX agent

arXiv:2608.02331v1 Announce Type: new Abstract: The same body posture can convey entirely different emotions depending on its surrounding context, yet most methods for recognising bodily emotions trea

researcharxiv-cs-cv
4 Aug 2026
Research

CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds

DGX agent

arXiv:2608.00674v1 Announce Type: new Abstract: Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subj

researcharxiv-cs-cv
4 Aug 2026
Model Releases

CORTIVA: Candidate-Score Fusion of Complementary Visual Teachers for EEG- and MEG-to-Image Retrieval

DGX agent

arXiv:2608.01355v1 Announce Type: new Abstract: Decoding visual experience from non-invasive brain activity is central to neuroscience and brain-computer interfaces. Functional magnetic resonance imag

model-releasesarxiv-cs-cv
4 Aug 2026
Tutorials

CoT-Edit: Let CoT Guide Instruction Video Editing

DGX agent

arXiv:2608.01113v1 Announce Type: new Abstract: Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relation

tutorialsarxiv-cs-cv
4 Aug 2026
Safety

Counting the Cost of War Under Satellite Embargo: Zero-Shot Estimation of Impacted Infrastructure

DGX agent

arXiv:2608.00119v1 Announce Type: new Abstract: Rapid estimation of impacted structures - critical for conflict-zone humanitarian response - is frequently hindered by post-strike satellite data embarg

safetyarxiv-cs-cv
4 Aug 2026
Tutorials

Coverage-Driven Adaptive Keyframe Selection for Video Understanding

DGX agent

arXiv:2608.00714v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of fram

tutorialsarxiv-cs-cv
4 Aug 2026
Model Releases

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

DGX agent

arXiv:2608.01644v1 Announce Type: new Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the pre

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

DGX agent

arXiv:2608.01055v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-obj

researcharxiv-cs-cv
4 Aug 2026
Research

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

DGX agent

arXiv:2608.01588v1 Announce Type: new Abstract: Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time

researcharxiv-cs-cv
4 Aug 2026
Safety

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

DGX agent

arXiv:2608.01821v1 Announce Type: new Abstract: Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual c

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

DeCLIP: Decoupled Prompting for Multi-Label Class-Incremental Learning with CLIP

DGX agent

arXiv:2509.23335v3 Announce Type: replace Abstract: Multi-label class-incremental learning (MLCIL) continuously expands the label space while recognizing multiple co-occurring categories, making catas

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Decoding Children's Gait Behavior

DGX agent

arXiv:2608.00371v1 Announce Type: new Abstract: We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We speci

researcharxiv-cs-cv
4 Aug 2026
Agents

DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing

DGX agent

arXiv:2608.01761v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong

agentsarxiv-cs-cv
4 Aug 2026
Model Releases

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

DGX agent

arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks

model-releasesarxiv-cs-cv
4 Aug 2026
Tutorials

Deep Learning CNN and Recurrence Analysis for Alpha Gamma EEG Biomarkers in Fragile X Syndrome

DGX agent

arXiv:2608.00835v1 Announce Type: cross Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disr

tutorialsarxiv-cs-cv
4 Aug 2026
Model Releases

Deep Learning for Retinal Degeneration Assessment: A Comprehensive Analysis of the MARIO Challenge

DGX agent

arXiv:2506.02976v4 Announce Type: replace Abstract: The MARIO challenge, held at MICCAI 2024, focused on advancing the automated detection and monitoring of age-related macular degeneration (AMD) thro

model-releasesarxiv-cs-cv
4 Aug 2026
Tutorials

Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion

DGX agent

arXiv:2608.02092v1 Announce Type: new Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusi

tutorialsarxiv-cs-cv
4 Aug 2026
Agents

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

DGX agent

arXiv:2608.01827v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to

agentsarxiv-cs-cv
4 Aug 2026
← Previous
1…1920212223…261
Next →