AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

DGX agent

arXiv:2510.03550v4 Announce Type: replace Abstract: Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensu

researcharxiv-cs-cv
29 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Structure-Aware Text Recognition for Ancient Greek Critical Editions

DGX agent

arXiv:2603.02803v2 Announce Type: replace Abstract: Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex

model-releasesarxiv-cs-cv
29 May 2026
Research

Subcortical Shape Variations and Their Associations with Cognition Across the 8th Decade of Life. A Study in the Lothian Birth Cohort 1936

DGX agent

arXiv:2605.29703v1 Announce Type: cross Abstract: The study of brain morphology changes in normal individuals may capture aspects of functionally-relevant brain aging not fully indicated by gross volu

researcharxiv-cs-cv
29 May 2026
Agents

Supercharging Thermal Gaussian Splatting with Depth Estimation

DGX agent

arXiv:2605.30328v1 Announce Type: new Abstract: Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content f

agentsarxiv-cs-cv
29 May 2026
Research

SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

DGX agent

arXiv:2605.29655v1 Announce Type: new Abstract: Autoregressive multimodal large language models (MLLMs) enable 3D generation but struggle to scale to high-resolution shapes due to inadequate 3D tokeni

researcharxiv-cs-cv
29 May 2026
Applications

SurfFill: Completion of LiDAR Point Clouds via Gaussian Surfel Splatting

DGX agent

arXiv:2512.03010v2 Announce Type: replace Abstract: LiDAR-captured point clouds are often considered the gold standard in active 3D reconstruction. While their accuracy is exceptional in flat regions,

applicationsarxiv-cs-cv
29 May 2026
Model Releases

SwInception -- Local Attention Meets Convolutions

DGX agent

arXiv:2605.29954v1 Announce Type: new Abstract: Sparse vision transformers have gained popularity as efficient encoders for medical volumetric segmentation, with Swin emerging as a prominent choice. S

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

TAE: Target-aware enhancer for nighttime UAV tracking

DGX agent

arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object

model-releasesarxiv-cs-cv
29 May 2026
Research

Towards Consistent Video Geometry Estimation

DGX agent

arXiv:2605.30060v1 Announce Type: new Abstract: This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built

researcharxiv-cs-cv
29 May 2026
Research

Towards the automated segmentation of epicardial and mediastinal fats: A multi-manufacturer approach using intersubject registration and random forest

DGX agent

arXiv:2605.29217v1 Announce Type: new Abstract: The amount of fat on the surroundings of the heart is correlated to several health risk factors such as carotid stiffness, coronary artery calcification

researcharxiv-cs-cv
29 May 2026
Safety

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

DGX agent

arXiv:2605.29894v1 Announce Type: new Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual task

safetyarxiv-cs-cv
29 May 2026
Research

Trajectory Constraints for Imaging Inverse Problems

DGX agent

arXiv:2605.29012v1 Announce Type: new Abstract: Diffusion-based and iterative methods have become effective tools for solving imaging inverse problems. Their reconstruction process naturally forms a t

researcharxiv-cs-cv
29 May 2026
Research

Treatment-Conditioned Diffusion for Forecasting Neurodegenerative Disease Progression

DGX agent

arXiv:2605.29932v1 Announce Type: cross Abstract: Forecasting the progression of neurodegenerative diseases, such as Parkinson's disease, is essential for effective long-term planning and personalized

researcharxiv-cs-cv
29 May 2026
Research

Turbulence-Robust Dynamic Object Segmentation with Multi-Signal Priors and SAM2 Refinement

DGX agent

arXiv:2605.29292v1 Announce Type: new Abstract: This technical report presents our solution for the CVPR 2026 UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence (DOST). We design a trai

researcharxiv-cs-cv
29 May 2026
Research

Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Field

DGX agent

arXiv:2605.30342v1 Announce Type: new Abstract: We present Gaussian Splatting Anisotropic Visibility Field (GAVIS), a novel framework for uncertainty quantification and active mapping in 3DGS. Our key

researcharxiv-cs-cv
29 May 2026
Tutorials

Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection

DGX agent

arXiv:2605.29455v1 Announce Type: new Abstract: Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. Wh

tutorialsarxiv-cs-cv
29 May 2026
Local Ai

UniNote: A Unified Embedding Model for Multimodal Representation and Ranking

DGX agent

arXiv:2605.29287v1 Announce Type: cross Abstract: Item-to-Item (I2I) retrieval is a fundamental part of modern content platforms, supporting critical industrial workflows from recommendation engines t

local-aiarxiv-cs-cv
29 May 2026
Local Ai

Unsupervised Semantic Segmentation Facilitates Model Understanding

DGX agent

arXiv:2605.29691v1 Announce Type: new Abstract: Self-supervised learning (SSL) has produced a diverse landscape of vision transformers (ViTs) whose pretrained representations support a wide range of d

local-aiarxiv-cs-cv
29 May 2026
Research

Unveiling the Visual Counting Bottleneck in Vision-Language Models

DGX agent

arXiv:2605.30170v1 Announce Type: cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visu

researcharxiv-cs-cv
29 May 2026
Safety

V2XCrafter: Learning to Generate Driving Scene Across Agents

DGX agent

arXiv:2605.29471v1 Announce Type: new Abstract: Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet

safetyarxiv-cs-cv
29 May 2026
Research

Veda: Scalable Video Diffusion via Distilled Sparse Attention

DGX agent

arXiv:2605.30325v1 Announce Type: new Abstract: Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse atte

researcharxiv-cs-cv
29 May 2026
Research

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement

DGX agent

arXiv:2605.29302v1 Announce Type: new Abstract: The digital media landscape has seen a pervasive shift toward short-form video advertising on TV, social media and e-commerce platforms. The present stu

researcharxiv-cs-cv
29 May 2026
Model Releases

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

DGX agent

arXiv:2601.12500v2 Announce Type: replace Abstract: Counting and tracking dense crowds in large-scale scenes is a highly practical yet challenging problem. Existing methods mostly rely on fixed-camera

model-releasesarxiv-cs-cv
29 May 2026
Research

Visual Spatial Learning: Single-Field Spatial Interpolation Using Convolutional Neural Networks

DGX agent

arXiv:2605.30167v1 Announce Type: cross Abstract: Predicting a complete spatially correlated field from sparse observations is a fundamental challenge in spatial statistics and environmental modelling

researcharxiv-cs-cv
29 May 2026
Model Releases

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

DGX agent

arXiv:2605.30317v1 Announce Type: new Abstract: Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time,

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

DGX agent

arXiv:2605.30161v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D und

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

DGX agent

arXiv:2605.30346v1 Announce Type: new Abstract: As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistica

model-releasesarxiv-cs-cv
29 May 2026
Tutorials

Zero-shot CT Super-Resolution using Diffusion-based 2D Projection Priors and Signed 3D Gaussians

DGX agent

arXiv:2508.15151v3 Announce Type: replace-cross Abstract: Computed tomography (CT) is important in clinical diagnosis, but acquiring high-resolution (HR) CT is constrained by radiation exposure risks.

tutorialsarxiv-cs-cv
29 May 2026
Research

A Multiscale Kinetic Framework for Image Segmentation: From Particle Systems to Continuum Models

DGX agent

arXiv:2605.28619v1 Announce Type: new Abstract: In this work, we present a multiscale kinetic framework for consensus-based image segmentation. By interpreting an image as a system of interacting part

researcharxiv-cs-cv
28 May 2026
Research

A novel ordinal multi-view aggregation scheme for oak defoliation

DGX agent

arXiv:2605.28151v1 Announce Type: new Abstract: Forest decline driven by climate and biotic stressors threatens ecosystem functioning, making accurate monitoring of tree health essential. In this work

researcharxiv-cs-cv
28 May 2026
Research

A Patient-Specific Pulmonary Arterial Tree Digital Twin to Extract Pulmonary Embolism Biomarkers

DGX agent

arXiv:2605.28217v1 Announce Type: new Abstract: Pulmonary embolism, the obstruction of a pulmonary artery by a blood clot, is one of the leading causes of acute cardiovascular syndrome. In clinical pr

researcharxiv-cs-cv
28 May 2026
Research

A Road-Conditioned Traffic Movie Prediction Network with Spatiotemporal and Structure-Consistent Learning

DGX agent

arXiv:2605.27884v1 Announce Type: new Abstract: City-wide traffic forecasting is important for congestion management, route guidance, and intelligent transportation systems, but accurate prediction re

researcharxiv-cs-cv
28 May 2026
Applications

A self-supervised learning approach to deep filter banks for texture recognition

DGX agent

arXiv:2605.27843v1 Announce Type: new Abstract: An important challenge in texture recognition is the limited amount of data for training frequently found in real-world applications. In computer vision

applicationsarxiv-cs-cv
28 May 2026
Research

A Survey on Event-based Optical Marker Systems

DGX agent

arXiv:2504.20736v2 Announce Type: replace-cross Abstract: The advent of event-based cameras, with their low latency, high dynamic range, and reduced power consumption, marked a turning point in machin

researcharxiv-cs-cv
28 May 2026
Research

ABot-OCR Technical Report

DGX agent

arXiv:2605.27978v1 Announce Type: new Abstract: We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing

researcharxiv-cs-cv
28 May 2026
Research

Accelerating Diffusion Sampling via Exploiting Local Transition Coherence

DGX agent

arXiv:2503.09675v3 Announce Type: replace Abstract: Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the

researcharxiv-cs-cv
28 May 2026
Research

Adaptive Temporal Gating of Longitudinal Magnetic Resonance Imaging for Alzheimer's Prediction

DGX agent

arXiv:2605.28397v1 Announce Type: new Abstract: Predicting conversion from Mild Cognitive Impairment (MCI) to Alzheimer's Disease (AD) is critical for early intervention. Current deep learning paradig

researcharxiv-cs-cv
28 May 2026
Research

Alterbute: Editing Intrinsic Attributes of Objects in Images

DGX agent

arXiv:2601.10714v2 Announce Type: replace Abstract: We introduce Alterbute, a diffusion-based method for editing an object's intrinsic attributes in an image. We allow changing color, texture, materia

researcharxiv-cs-cv
28 May 2026
Research

An analytic theory of convolutional neural network inverse problems solvers

DGX agent

arXiv:2601.10334v2 Announce Type: replace Abstract: Supervised convolutional neural networks (CNNs) are widely used to solve imaging inverse problems, achieving state-of-the-art performance in numerou

researcharxiv-cs-cv
28 May 2026
Research

An Empirical Study on Variance-based MC Dropout Uncertainty-Error Correlation in 2D Brain Tumor Segmentation

DGX agent

arXiv:2510.15541v2 Announce Type: replace-cross Abstract: Accurate brain tumor segmentation from MRI is vital for diagnosis and treatment planning. Although Monte Carlo (MC) Dropout is widely used to

researcharxiv-cs-cv
28 May 2026
Model Releases

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

DGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

model-releasesarxiv-cs-cv
28 May 2026
Safety

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

DGX agent

arXiv:2605.28809v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing

safetyarxiv-cs-cv
28 May 2026
Safety

Artemis: Structured Visual Reasoning for Perception Policy Learning

DGX agent

arXiv:2512.01988v2 Announce Type: replace Abstract: Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural languag

safetyarxiv-cs-cv
28 May 2026
Model Releases

Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction

DGX agent

arXiv:2605.27726v1 Announce Type: new Abstract: Frequent cloud cover severely limits the usability of Sentinel-2 (S2) optical time series for Earth surface monitoring. Sentinel-1 (S1) SAR provides all

model-releasesarxiv-cs-cv
28 May 2026
Safety

Automated Estimation of Impact Time, Impact Location, and Shuttlecock Speed in Badminton Smashes Using Event Cameras

DGX agent

arXiv:2605.28011v1 Announce Type: new Abstract: Quantifying impact phenomena in badminton smashes is important for evaluating both athletic performance and equipment; however, conventional measurement

safetyarxiv-cs-cv
28 May 2026
Tutorials

Automatic Pruning Discovery for Large Language Models

DGX agent

arXiv:2511.15390v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance on a wide range of tasks, hindering real-world deployment due to their massive siz

tutorialsarxiv-cs-cv
28 May 2026
Model Releases

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

DGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

model-releasesarxiv-cs-cv
28 May 2026
← Previous
1…137138139140141…263
Next →