AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
5 Aug 2026

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

Model ReleasesDGX agent

arXiv:2608.03649v1 Announce Type: new Abstract: Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead,

XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation

ApplicationsDGX agent

arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computation

4 Aug 2026

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ResearchDGX agent

arXiv:2608.01185v1 Announce Type: new Abstract: Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial rea

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

Model ReleasesDGX agent

arXiv:2608.01258v1 Announce Type: new Abstract: The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years.

A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology

Model ReleasesDGX agent

arXiv:2608.02300v1 Announce Type: new Abstract: Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology re

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

Model ReleasesDGX agent

arXiv:2608.02505v1 Announce Type: cross Abstract: Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothes

ABRA: Teleporting Fine-Tuned Knowledge Across Domains for Open-Vocabulary Object Detection

ResearchDGX agent

arXiv:2603.12409v2 Announce Type: replace Abstract: Although recent Open-Vocabulary Object Detection architectures, such as Grounding DINO, demonstrate strong zero-shot capabilities, their performance

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks lar

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

ApplicationsDGX agent

arXiv:2608.02471v1 Announce Type: new Abstract: Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to

AdaForensics: Learning A Characteristic-aware Adaptive Deepfake Detector

TutorialsDGX agent

arXiv:2608.02160v1 Announce Type: new Abstract: In this paper, we propose a characteristic-aware adaptive network named AdaForensics for deepfake detection. Most existing methods learn a fixed network

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

SafetyDGX agent

arXiv:2608.01980v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether

AeroLLE: Constrained Pseudo-Supervision for Nighttime Aerial Image Enhancement with the AeroNight-1.5K Benchmark

Model ReleasesDGX agent

arXiv:2608.00702v1 Announce Type: new Abstract: Nighttime aerial image enhancement is challenged by spatially nonuniform exposure, mixed illumination, and weak structural evidence, while registered no

AIMold: An Autonomous AI-based Pipeline for Complex Mold Design

AgentsDGX agent

arXiv:2608.00800v1 Announce Type: new Abstract: Injection molding is the cornerstone of mass-producing plastic components. While current algorithms can automate mold design for basic geometries using

AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting

Local AiDGX agent

arXiv:2603.25129v2 Announce Type: replace Abstract: While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero-shot capabilities in visual geometry estimation, their direct applicati

An Accessible Solution for Deformable Image Registration Compared with Learning-Based Approaches

Model ReleasesDGX agent

arXiv:2608.02248v1 Announce Type: cross Abstract: Deformable image registration (DIR) is a core problem in medical image analysis; but, unlike labeling decision problems such as classification and seg

ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection

ApplicationsDGX agent

arXiv:2607.02252v2 Announce Type: replace Abstract: The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which

ARM: Detector-Agnostic Changepoint Attribution with Finite-Sample Error Control

ResearchDGX agent

arXiv:2608.01691v1 Announce Type: cross Abstract: Detecting a change in a multivariate series answers only the first of two questions; the operational question is which coordinates changed. Existing a

Artificial Intelligence for the Characterization of Particles and Fibers by Optical Microscopy

ResearchDGX agent

arXiv:2608.00361v1 Announce Type: new Abstract: Optical microscopy of particle and fiber dispersions involves interpreting subtle visual cues influenced by specimen morphology, chemical composition, m

Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data

Model ReleasesDGX agent

arXiv:2608.01336v1 Announce Type: new Abstract: Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism,

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

ApplicationsDGX agent

arXiv:2608.01906v1 Announce Type: new Abstract: Rapid and accurate post-disaster building damage assessment is essential, yet remains a challenging task. Unmanned Aerial Vehicle (UAV) imagery offers a

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

SafetyDGX agent

arXiv:2608.02006v1 Announce Type: new Abstract: Dynamic 3D scene reconstruction has achieved remarkable success under the assumption of strictly synchronized multi-camera inputs. However, in real-worl

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

TutorialsDGX agent

arXiv:2608.01276v1 Announce Type: new Abstract: Full-body capture from unconstrained photographs requires global correspondence across arbitrary views, poses, crops, and occlusions. Yet pose, geometry

Attention-Steered Vision-Language Models for Sign Language Translation

ResearchDGX agent

arXiv:2608.00235v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language

Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice

Model ReleasesDGX agent

arXiv:2608.00145v1 Announce Type: cross Abstract: Accurate estimation of left ventricular (LV) orientation is essential for cardiac magnetic resonance (CMR) imaging and downstream analysis. Existing m

Belief-Space Perception Routing under Coupled Sensor Faults and Compute Contention

Model ReleasesDGX agent

arXiv:2608.00322v1 Announce Type: cross Abstract: A robot that has to see and react on a fixed clock runs into two problems at once. Its cameras degrade in rain, mud, fog, and darkness. And the single

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

Model ReleasesDGX agent

arXiv:2608.00077v1 Announce Type: new Abstract: Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this proto

Beyond Edge Maps: Wavelet-Domain Conditioning for Multi-Adapter Map-to-Satellite Diffusion

Model ReleasesDGX agent

arXiv:2608.00083v1 Announce Type: new Abstract: Commercial mapping partnerships are often unavailable in low-resource regions, leaving satellite basemaps stale and motivating synthesis of satellite im

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

Local AiDGX agent

arXiv:2608.02016v1 Announce Type: new Abstract: Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active

Beyond Illumination: A Conditional Mutual Information-Guided Network for Low-Light Image Enhancement

ResearchDGX agent

arXiv:2608.01886v1 Announce Type: new Abstract: Low-light image enhancement (LLIE) seeks to restore structural fidelity, natural color rendition, and proper exposure from images captured under inadequ

Beyond Random Partitioning: Unsupervised Spatio-Temporal Stratification for Cohort Balancing in Longitudinal Medical Imaging

SafetyDGX agent

arXiv:2608.00073v1 Announce Type: new Abstract: Rigorous dataset partitioning is a foundational, yet frequently overlooked, prerequisite for reliable deep learning in longitudinal medical imaging. Nai

Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection

Local AiDGX agent

arXiv:2608.00442v1 Announce Type: new Abstract: Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and modalities. Exi

Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection

ResearchDGX agent

arXiv:2608.01302v1 Announce Type: new Abstract: State-of-the-art RGB-Event detectors improve the detection of small, fast-moving objects by combining complementary features from RGB and Event data, ye

Beyond Token-Level Cross-Entropy: Frechet Distributional Post-Training for Autoregressive Image Generation

ResearchDGX agent

arXiv:2608.00562v1 Announce Type: new Abstract: Autoregressive image generators are commonly pretrained with token-level cross-entropy under teacher forcing, yet evaluated by the distributional qualit

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

Model ReleasesDGX agent

arXiv:2608.00415v1 Announce Type: new Abstract: Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various en

Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection

ResearchDGX agent

arXiv:2601.02837v2 Announce Type: replace Abstract: Infrared small target detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, extremely small target sizes, and complex cl

Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2608.00678v1 Announce Type: new Abstract: Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, ev

Breaking the Statistical Similarity Trap in Extreme Convection Detection

SafetyDGX agent

arXiv:2509.09195v2 Announce Type: replace-cross Abstract: Current evaluation metrics for deep learning weather models create a 'Statistical Similarity Trap', rewarding blurry predictions while missing

BRIC-Net: Boundary-Reliable Illumination-Color Interaction for Remote Sensing Image Deshadowing

ResearchDGX agent

arXiv:2608.00682v1 Announce Type: new Abstract: Shadows in remote sensing images obscure surface appearance and disrupt radiometric continuity, reducing the reliability of visual interpretation and do

CADENA: Stepwise CAD Reverse Engineering

Model ReleasesDGX agent

arXiv:2608.00799v1 Announce Type: new Abstract: Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. M

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

SafetyDGX agent

arXiv:2608.02309v1 Announce Type: new Abstract: We present CalibBEV, a novel Bird's Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shar

Calibrated Similarity and Graph Clustering for Open-Set Animal Re-Identification

ResearchDGX agent

arXiv:2608.02469v1 Announce Type: new Abstract: AnimalCLEF26 addresses discovery-oriented animal re-identification, where systems must both attach query images to known individuals and discover unseen

Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit

ApplicationsDGX agent

arXiv:2608.01753v1 Announce Type: new Abstract: Addressing urban blight has seen increased focus in the past 15 years. Assessing urban blight is essential for guiding urban planning, targeting rehabil

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Model ReleasesDGX agent

arXiv:2608.02589v1 Announce Type: new Abstract: Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the c

CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views

Local AiDGX agent

arXiv:2506.18164v2 Announce Type: replace Abstract: Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation

ResearchDGX agent

arXiv:2602.10880v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have shown promise in generating plotting code from chart images, yet achieving structural fidelity remains challengin

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

SafetyDGX agent

arXiv:2608.00769v1 Announce Type: new Abstract: One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes

CHOW-SLAM: Compact Hybrid Representation with Complementary Overlap Window Optimization for RGB-D SLAM

Model ReleasesDGX agent

arXiv:2608.01914v1 Announce Type: new Abstract: Simultaneous localization and mapping (SLAM) based on Neural Radiance Fields (NeRF) enables dense, continuous scene reconstruction. However, existing sy

CLEAR: Conflict-aware Learning via Evidence-guided Adaptive Routing for Unified Sparse-View 3D Gaussian Super-Resolution

ApplicationsDGX agent

arXiv:2608.02206v1 Announce Type: new Abstract: Sparse-view 3D Gaussian Splatting Super-resolution is highly challenging since the sparse and low-resolution (LR) inputs lack sufficient geometric and h

Clear-Weighted Bit Allocation for Satellite Downlinks

ResearchDGX agent

arXiv:2608.01457v1 Announce Type: cross Abstract: Earth-observation satellites capture more imagery than intermittent ground contacts can transmit. Onboard systems threshold a cloud detector, discard

Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild

ResearchDGX agent

arXiv:2608.02331v1 Announce Type: new Abstract: The same body posture can convey entirely different emotions depending on its surrounding context, yet most methods for recognising bodily emotions trea

CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds

ResearchDGX agent

arXiv:2608.00674v1 Announce Type: new Abstract: Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subj

CORTIVA: Candidate-Score Fusion of Complementary Visual Teachers for EEG- and MEG-to-Image Retrieval

Model ReleasesDGX agent

arXiv:2608.01355v1 Announce Type: new Abstract: Decoding visual experience from non-invasive brain activity is central to neuroscience and brain-computer interfaces. Functional magnetic resonance imag

CoT-Edit: Let CoT Guide Instruction Video Editing

TutorialsDGX agent

arXiv:2608.01113v1 Announce Type: new Abstract: Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relation

Counting the Cost of War Under Satellite Embargo: Zero-Shot Estimation of Impacted Infrastructure

SafetyDGX agent

arXiv:2608.00119v1 Announce Type: new Abstract: Rapid estimation of impacted structures - critical for conflict-zone humanitarian response - is frequently hindered by post-strike satellite data embarg

Coverage-Driven Adaptive Keyframe Selection for Video Understanding

TutorialsDGX agent

arXiv:2608.00714v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of fram

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.01644v1 Announce Type: new Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the pre

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

ResearchDGX agent

arXiv:2608.01055v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-obj

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

ResearchDGX agent

arXiv:2608.01588v1 Announce Type: new Abstract: Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time

DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models

SafetyDGX agent

arXiv:2608.01821v1 Announce Type: new Abstract: Diffusion vision-language models (dVLMs) iteratively denoise masked responses while conditioning each denoising step on visual evidence, making visual c

DeCLIP: Decoupled Prompting for Multi-Label Class-Incremental Learning with CLIP

Model ReleasesDGX agent

arXiv:2509.23335v3 Announce Type: replace Abstract: Multi-label class-incremental learning (MLCIL) continuously expands the label space while recognizing multiple co-occurring categories, making catas

← Previous
1…1314151617…207
Next →