AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlog
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
49+ results
Syntheses

Synthesis: Arxiv-Cs-Cv

DGX agent

Auto-generated synthesis of 874 entries about arxiv-cs-cv

synthesisarxiv-cs-cvauto-generated
16 Apr 2026
Agents

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

DGX agent

arXiv:2608.10107v1 Announce Type: new Abstract: Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and

agents
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
arxiv-cs-cv
12 Aug 2026
Safety

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

DGX agent

arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

DGX agent

arXiv:2608.10978v1 Announce Type: new Abstract: Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form s

model-releasesarxiv-cs-cv
12 Aug 2026
Applications

A second-order theory of texture for depth from focus

DGX agent

arXiv:2608.10411v1 Announce Type: new Abstract: We present a theory of textured appearance of optically rough surfaces based on wave optics, emphasizing the role of texture for passive depth from focu

applicationsarxiv-cs-cv
12 Aug 2026
Safety

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

DGX agent

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

safetyarxiv-cs-cv
12 Aug 2026
Research

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

DGX agent

arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for

researcharxiv-cs-cv
12 Aug 2026
Research

Algorithmic statistics of retinal images

DGX agent

arXiv:2608.09989v1 Announce Type: cross Abstract: There has been a tremendous amount of image processing and machine learning research to measure and classify disease progression from live optical coh

researcharxiv-cs-cv
12 Aug 2026
Safety

APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

DGX agent

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

safetyarxiv-cs-cv
12 Aug 2026
Research

Beyond Pixels: From Video Priors to 4D Worlds

DGX agent

arXiv:2608.10744v1 Announce Type: new Abstract: 4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a sepa

researcharxiv-cs-cv
12 Aug 2026
Safety

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

DGX agent

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

safetyarxiv-cs-cv
12 Aug 2026
Research

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

DGX agent

arXiv:2608.10479v1 Announce Type: new Abstract: Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling lar

researcharxiv-cs-cv
12 Aug 2026
Model Releases

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

DGX agent

arXiv:2608.10680v1 Announce Type: new Abstract: Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignmen

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

DGX agent

arXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (

model-releasesarxiv-cs-cv
12 Aug 2026
Research

Capturing Uncertainty in Human Motion for Representation Learning in Soccer

DGX agent

arXiv:2608.11203v1 Announce Type: new Abstract: This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion

researcharxiv-cs-cv
12 Aug 2026
Applications

CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images

DGX agent

arXiv:2608.10345v1 Announce Type: new Abstract: Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and mo

applicationsarxiv-cs-cv
12 Aug 2026
Research

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

DGX agent

arXiv:2608.11150v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle

researcharxiv-cs-cv
12 Aug 2026
Agents

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

DGX agent

arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomo

agentsarxiv-cs-cv
12 Aug 2026
Model Releases

Chartography: A Benchmark for Professional Chart Understanding

DGX agent

arXiv:2608.10677v1 Announce Type: new Abstract: Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benc

model-releasesarxiv-cs-cv
12 Aug 2026
Research

Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging

DGX agent

arXiv:2608.10712v1 Announce Type: new Abstract: 3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this

researcharxiv-cs-cv
12 Aug 2026
Safety

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

DGX agent

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

safetyarxiv-cs-cv
12 Aug 2026
Research

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

DGX agent

arXiv:2608.11093v1 Announce Type: cross Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field

researcharxiv-cs-cv
12 Aug 2026
Research

DoseBridge: Denoising Diffusion Bridge Model for Dose Prediction in Lung Intensity-Modulated Proton Therapy

DGX agent

arXiv:2608.10173v1 Announce Type: new Abstract: Most radiotherapy dose-prediction models use only CT images and anatomical structures, although intensity-modulated proton therapy (IMPT) dose also depe

researcharxiv-cs-cv
12 Aug 2026
Model Releases

DreamOmni3: Scribble-based Editing and Generation

DGX agent

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

DGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

model-releasesarxiv-cs-cv
12 Aug 2026
Tutorials

DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars

DGX agent

arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis

tutorialsarxiv-cs-cv
12 Aug 2026
Tutorials

DynaPPI: A Large-scale Dynamic Protein Dataset for AI-driven Advances in Protein Interactomics

DGX agent

arXiv:2608.10435v1 Announce Type: new Abstract: Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven bi

tutorialsarxiv-cs-cv
12 Aug 2026
Model Releases

E^3mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

DGX agent

arXiv:2608.10796v1 Announce Type: new Abstract: Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interact

model-releasesarxiv-cs-cv
12 Aug 2026
Local Ai

Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive Perception

DGX agent

arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa

local-aiarxiv-cs-cv
12 Aug 2026
Local Ai

Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition

DGX agent

arXiv:2608.10684v1 Announce Type: new Abstract: Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods e

local-aiarxiv-cs-cv
12 Aug 2026
Local Ai

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

DGX agent

arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex

local-aiarxiv-cs-cv
12 Aug 2026
Local Ai

ENCORE: Efficient Noise Context-Aware Representation for Low-Dose CT Denoising

DGX agent

arXiv:2608.10343v1 Announce Type: new Abstract: While deep learning-based denoising has become widely adopted in low-dose CT, conventional models use generic architectures designed for natural images,

local-aiarxiv-cs-cv
12 Aug 2026
Research

Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery

DGX agent

arXiv:2608.10801v1 Announce Type: new Abstract: Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segm

researcharxiv-cs-cv
12 Aug 2026
Research

Every Packet Counts: Dispersing Information for Loss-Resilient Learned Image Compression

DGX agent

arXiv:2608.11096v1 Announce Type: new Abstract: Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss,

researcharxiv-cs-cv
12 Aug 2026
Model Releases

Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking

DGX agent

arXiv:2507.21606v2 Announce Type: replace Abstract: The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous hu

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

DGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

model-releasesarxiv-cs-cv
12 Aug 2026
Research

FARCLUSS: Fuzzy Adaptive Rebalancing and Contrastive Uncertainty Learning for Semi-Supervised Semantic Segmentation

DGX agent

arXiv:2506.11142v3 Announce Type: replace Abstract: Semi-supervised semantic segmentation (SSSS) faces persistent challenges in effectively leveraging unlabeled data, such as ineffective utilization o

researcharxiv-cs-cv
12 Aug 2026
Model Releases

Flex-pi: A Multi-Stream World-Action Model with Compute Flexibility

DGX agent

arXiv:2608.10860v1 Announce Type: cross Abstract: World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction,

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

DGX agent

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

model-releasesarxiv-cs-cv
12 Aug 2026
Research

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

DGX agent

arXiv:2608.11076v1 Announce Type: new Abstract: Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotrac

researcharxiv-cs-cv
12 Aug 2026
Hardware

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

DGX agent

arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection.

hardwarearxiv-cs-cv
12 Aug 2026
Tutorials

Gaussian Sculpting: End-to-End Controllable Surface Reconstruction via Field Optimization

DGX agent

arXiv:2608.10602v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled real-time novel view synthesis with impressive quality. However, it struggles to recover accurate surf

tutorialsarxiv-cs-cv
12 Aug 2026
Model Releases

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

DGX agent

arXiv:2608.10426v1 Announce Type: new Abstract: Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories sp

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

DGX agent

arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how

model-releasesarxiv-cs-cv
12 Aug 2026
Local Ai

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

DGX agent

arXiv:2608.10723v1 Announce Type: new Abstract: Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher i

local-aiarxiv-cs-cv
12 Aug 2026
Research

GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids

DGX agent

arXiv:2603.25467v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without c

researcharxiv-cs-cv
12 Aug 2026
Research

GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting

DGX agent

arXiv:2608.10938v1 Announce Type: new Abstract: Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generaliza

researcharxiv-cs-cv
12 Aug 2026
Model Releases

HNDiff: Haze-Noise Diffusion for Image Dehazing

DGX agent

arXiv:2608.10995v1 Announce Type: new Abstract: Existing diffusion-based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formatio

model-releasesarxiv-cs-cv
12 Aug 2026
← Previous
123…259
Next →
12,414 results