AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
24 Jul 2026

Unified Video Dense Prediction from Disjoint Data

ResearchDGX agent

arXiv:2607.21592v1 Announce Type: new Abstract: Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragment

Unsupervised Metal Artifact Reduction in Dental CBCT using Fine-tuned Cycle-Consistent Adversarial Networks

ResearchDGX agent

arXiv:2607.20977v1 Announce Type: new Abstract: Metal artifacts generated by dental implants significantly degrade cone-beam computed tomography (CBCT) volumes, obscuring critical anatomical structure

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Model ReleasesDGX agent

arXiv:2607.20868v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental ab


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

WAT3R: Feedforward Underwater 3D Reconstruction

ResearchDGX agent

arXiv:2607.21023v1 Announce Type: new Abstract: Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality a

Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

Model ReleasesDGX agent

arXiv:2607.20874v1 Announce Type: new Abstract: Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learn

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

Model ReleasesDGX agent

arXiv:2607.20883v1 Announce Type: new Abstract: Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing

23 Jul 2026

4DGS360: 360{eg} Gaussian Reconstruction of Dynamic Objects from a Single Video

Model ReleasesDGX agent

arXiv:2603.21618v2 Announce Type: replace Abstract: We introduce 4DGS360, a diffusion-free framework for 360^{irc} dynamic object reconstruction from casual monocular video. Existing methods often fai

A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities

Model ReleasesDGX agent

arXiv:2607.19716v1 Announce Type: new Abstract: Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assessment is essential for effective clinical m

A Unified Variational Framework for Deep Weakly Supervised Image Segmentation

ResearchDGX agent

arXiv:2607.19669v1 Announce Type: new Abstract: We propose a unified variational framework for image segmentation under sparse pixel-level supervision. Our method is based on a simplex-constrained Pot

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

Local AiDGX agent

arXiv:2607.19191v1 Announce Type: new Abstract: We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data i

Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

ResearchDGX agent

arXiv:2602.01345v2 Announce Type: replace Abstract: Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention

TutorialsDGX agent

arXiv:2607.19086v1 Announce Type: new Abstract: Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, cli

Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

Model ReleasesDGX agent

arXiv:2601.15368v3 Announce Type: replace Abstract: Generative image inpainting can produce realistic results even with large, irregular masks, but existing methods still suffer from two common proble

An Exploratory Analysis of Pain Localization via Explainable Computational Modeling

Local AiDGX agent

arXiv:2607.19726v1 Announce Type: new Abstract: Automatic pain localization, which involves identifying the anatomical origin of pain from peripheral physiological signals without patient self-report,

Anatomy-Aware 3D Mesh Refinement of Pericardium Segmentations on Computed Tomography

HardwareDGX agent

arXiv:2607.19210v1 Announce Type: new Abstract: Accurate delineation of the pericardium in a cardiac CT scan is essential for quantifying epicardial adipose tissue, yet it remains one of the most chal

AniGS: Bridging Rendering and Diffusion Prior for 3D Scene Animation

ApplicationsDGX agent

arXiv:2607.18539v1 Announce Type: new Abstract: Novel view rendering of large and complex reconstructed scenes is becoming increasingly photorealistic. However, most reconstructions remain static and

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Local AiDGX agent

arXiv:2607.19344v1 Announce Type: new Abstract: Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identiti

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

HardwareDGX agent

arXiv:2607.20417v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in u

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs

ResearchDGX agent

arXiv:2607.18577v1 Announce Type: new Abstract: Attention and saliency heatmaps are widely used to explain medical Vision-Language Model (VLM) outputs on chest X-rays, yet whether they truly highlight

Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

ResearchDGX agent

arXiv:2607.18695v1 Announce Type: new Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the result

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

ResearchDGX agent

arXiv:2607.18997v1 Announce Type: new Abstract: Information Extraction (IE) from Architecture, Engineering, and Construction (AEC) drawings remains hindered by manual inefficiency, while Layout Detect

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics

ApplicationsDGX agent

arXiv:2607.19515v1 Announce Type: cross Abstract: Continuous surveillance video creates a growing storage, transmission, and inference burden for enterprise video analytics systems. While modern codec

Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks

ResearchDGX agent

arXiv:2607.18767v1 Announce Type: new Abstract: The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalability. However

Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

Local AiDGX agent

arXiv:2602.19863v3 Announce Type: replace Abstract: Foundation models are transforming Earth Observation (EO), yet the diversity of EO sensors and modalities makes a single universal model unrealistic

CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders

ResearchDGX agent

arXiv:2607.18451v1 Announce Type: cross Abstract: Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Nearest-neighbor

CGMap: A Geospatially Aware Deep Learning Framework for Crop Gap Mapping Using UAV

ResearchDGX agent

arXiv:2607.18779v1 Announce Type: new Abstract: In India, crop germination is primarily monitored by visual inspection and manual counting, which are prone to errors, despite their crucial role in det

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning

ResearchDGX agent

arXiv:2607.19547v1 Announce Type: new Abstract: Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approac

CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation

Local AiDGX agent

arXiv:2607.16577v2 Announce Type: replace Abstract: This paper presents a latent-space 3D shape editing framework built upon a coupled neural shape (CNS) representation and a neural feature volume opt

Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency

SafetyDGX agent

arXiv:2607.19194v2 Announce Type: cross Abstract: High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely infer

CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement

AgentsDGX agent

arXiv:2607.19036v1 Announce Type: new Abstract: V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple col

Confidence-Gated Vision-Only Heading Alignment for UAV-UGV Cooperative Systems

SafetyDGX agent

arXiv:2607.18713v1 Announce Type: cross Abstract: Vision-based heading prediction is useful for UAV--UGV cooperation, but accurate prediction alone does not guarantee that every predicted heading shou

Context-structured Video Anomaly Detection with Large Vision-Language Models

ResearchDGX agent

arXiv:2607.19077v1 Announce Type: new Abstract: Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vis

Continual Video-MLLM Adaptation over Evolving Domains

Model ReleasesDGX agent

arXiv:2607.18716v1 Announce Type: new Abstract: Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remain

Contour Errors: Ego-Centric Matching for 3D Multi-Object Tracking Performance Evaluation

SafetyDGX agent

arXiv:2506.04122v4 Announce Type: replace Abstract: Open-loop performance evaluation of 3D multi-object tracking in autonomous driving requires matching criteria that effectively penalize translationa

Contrastive On-Policy Distillation

SafetyDGX agent

arXiv:2607.19046v1 Announce Type: new Abstract: On-policy Distillation (OPD) supervises a student model on trajectories sampled from its own policy by minimizing the divergence between the output dist

CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval

Model ReleasesDGX agent

arXiv:2607.19115v1 Announce Type: new Abstract: Edit-conditioned 3D scene retrieval pairs a reference 3D room with a natural-language modification and retrieves rooms from a corpus that satisfy the ed

CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation

Model ReleasesDGX agent

arXiv:2506.10890v2 Announce Type: replace Abstract: Graphic design plays a crucial role in both commercial and personal contexts, yet creating high-quality, editable, and aesthetically pleasing graphi

Cross-Dataset Generalization in Breast MRI Tumor Classification via Class-Wise Dataset Mixing

SafetyDGX agent

arXiv:2607.18678v1 Announce Type: new Abstract: Breast MRI is highly sensitive for detecting breast tumors, but exams contain many slices and require substantial reading time. Deep learning models oft

Cross-Modal UAV Object Tracking: State-Aware Representation Learning and A Unified Benchmark

Model ReleasesDGX agent

arXiv:2607.18768v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) object tracking has emerged as a popular research field with broad practical applications. Modern UAVs are increasingly eq

Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

SafetyDGX agent

arXiv:2607.19517v1 Announce Type: new Abstract: Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex sce

Current Injection Spiking Neural Network for Infrared and Visible Image Fusion

Model ReleasesDGX agent

arXiv:2607.19879v1 Announce Type: new Abstract: Infrared and visible image fusion (IVIF) integrates the complementary information of two modalities into a single image with richer scene content. While

Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection

Local AiDGX agent

arXiv:2607.18700v1 Announce Type: new Abstract: The FathomNetCLEF 2026 competition combines underwater object detection and fine-grained marine species classification under a positive-unlabeled evalua

Deep Learning Estimation of Sex, Age, Height, and Weight from CT-derived Digitally Reconstructed Radiographs

ResearchDGX agent

arXiv:2607.18638v1 Announce Type: new Abstract: Purpose: To develop and validate a deep learning ensemble for estimating adult sex, age, height, and weight from coronal digitally reconstructed radiogr

DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

Local AiDGX agent

arXiv:2607.18664v1 Announce Type: new Abstract: Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be describe

Delineate Anything v2: A Global Foundation Model for Field Delineation

Model ReleasesDGX agent

arXiv:2607.19069v1 Announce Type: new Abstract: Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounti

Denoising Monte Carlo Renders with Diffusion Models

ResearchDGX agent

arXiv:2404.00491v3 Announce Type: replace Abstract: Physically-based renderings contain Monte Carlo noise, with variance that increases as the number of rays per pixel decreases. This noise, while zer

Detect Early, Escalate Rarely: Anytime Detection of AI-Generated Video from the Compressed Bitstream

ResearchDGX agent

arXiv:2607.19476v1 Announce Type: new Abstract: Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasingly by a large vision-language model. Dete

Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images

Local AiDGX agent

arXiv:2607.20087v1 Announce Type: new Abstract: Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure, function, and pathology, but requires sub

Diverse-Intent Multi-Turn Fashion Image Retrieval

Model ReleasesDGX agent

arXiv:2607.20291v1 Announce Type: new Abstract: Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictiv

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

Model ReleasesDGX agent

arXiv:2607.18988v1 Announce Type: new Abstract: Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completenes

Domain Shift in Echocardiography: Interpretable Quantification and Prediction of Cross-Dataset Left Ventricular Segmentation

ResearchDGX agent

arXiv:2607.19643v1 Announce Type: new Abstract: Cross-dataset generalisation remains a major barrier to clinical deployment of echocardiographic left ventricular segmentation, yet the sources of this

DRGBT-1K: A Large-scale High-quality Benchmark for Dynamic RGBT Tracking

Model ReleasesDGX agent

arXiv:2607.19772v1 Announce Type: new Abstract: Dynamic RGBT (DRGBT) tracking aims to continuously localize a target when the available sensing modalities and observation platforms vary over time. Com

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model

SafetyDGX agent

arXiv:2607.18958v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulne

Dual-Edged Homogeneous-Modality Similarity: Towards Visible-Infrared Modality-Incomplete Person Re-Identification with Modality Adaptive Matching

ResearchDGX agent

arXiv:2607.18688v1 Announce Type: new Abstract: Visible-Infrared Person Re-Identification (VI-ReID) operates under a closed-world assumption, where queries and galleries are from heterogeneous modalit

DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer

ResearchDGX agent

arXiv:2607.18510v1 Announce Type: new Abstract: Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this info

EAR-Net: Pursuing End-to-End Absolute Rotations from Multi-View Images

ResearchDGX agent

arXiv:2310.10051v3 Announce Type: replace Abstract: Absolute rotation estimation is an important topic in 3D computer vision. Existing works in literature generally employ a multi-stage (at least two-

ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization

Model ReleasesDGX agent

arXiv:2607.18466v1 Announce Type: new Abstract: Recent advances in differentiable Gaussian splatting have highlighted the potential of primitive-based approaches as alternative scene representations f

Efficient Tracking and Understanding Object Transformations

ApplicationsDGX agent

arXiv:2607.19743v1 Announce Type: new Abstract: Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensi

EGRNet: A Lightweight Semantic Segmentation Network with Edge-Gated Refinement and Adversarial Sensing

SafetyDGX agent

arXiv:2607.19617v1 Announce Type: new Abstract: As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semant

ERank in Latent Space as an Image-Complexity and Richness Measure

ResearchDGX agent

arXiv:2607.19315v1 Announce Type: new Abstract: We propose the effective rank (ERank) of the channel covariance of an image's deep feature map as a per-sample, label-free measure of visual richness, c

← Previous
1…3536373839…209
Next →