AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
61+ results
16 Apr 2026

Synthesis: Arxiv-Cs-Cv

SynthesesDGX agent

Auto-generated synthesis of 874 entries about arxiv-cs-cv

12 Aug 2026

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

AgentsDGX agent

arXiv:2608.10107v1 Announce Type: new Abstract: Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

SafetyDGX agent

arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

Model ReleasesDGX agent

arXiv:2608.10978v1 Announce Type: new Abstract: Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form s

A second-order theory of texture for depth from focus

ApplicationsDGX agent

arXiv:2608.10411v1 Announce Type: new Abstract: We present a theory of textured appearance of optically rough surfaces based on wave optics, emphasizing the role of texture for passive depth from focu

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

SafetyDGX agent

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

ResearchDGX agent

arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for

Algorithmic statistics of retinal images

ResearchDGX agent

arXiv:2608.09989v1 Announce Type: cross Abstract: There has been a tremendous amount of image processing and machine learning research to measure and classify disease progression from live optical coh

APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

SafetyDGX agent

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

Beyond Pixels: From Video Priors to 4D Worlds

ResearchDGX agent

arXiv:2608.10744v1 Announce Type: new Abstract: 4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a sepa

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

SafetyDGX agent

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

ResearchDGX agent

arXiv:2608.10479v1 Announce Type: new Abstract: Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling lar

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

Model ReleasesDGX agent

arXiv:2608.10680v1 Announce Type: new Abstract: Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignmen

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

Model ReleasesDGX agent

arXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (

Capturing Uncertainty in Human Motion for Representation Learning in Soccer

ResearchDGX agent

arXiv:2608.11203v1 Announce Type: new Abstract: This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion

CasDeblurGS: Cascaded 2D-to-3D Multi-View Consistency for 3D Gaussian Splatting from Two Blurry Images

ApplicationsDGX agent

arXiv:2608.10345v1 Announce Type: new Abstract: Free-viewpoint 3D scene media is increasingly important for immersive applications, yet practical capture often suffers from severe view sparsity and mo

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

ResearchDGX agent

arXiv:2608.11150v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

AgentsDGX agent

arXiv:2608.10278v1 Announce Type: new Abstract: Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomo

Chartography: A Benchmark for Professional Chart Understanding

Model ReleasesDGX agent

arXiv:2608.10677v1 Announce Type: new Abstract: Professionals across medicine, engineering, finance, manufacturing, and the sciences often make consequential decisions from charts. Existing chart benc

Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging

ResearchDGX agent

arXiv:2608.10712v1 Announce Type: new Abstract: 3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

SafetyDGX agent

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

ResearchDGX agent

arXiv:2608.11093v1 Announce Type: cross Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field

DoseBridge: Denoising Diffusion Bridge Model for Dose Prediction in Lung Intensity-Modulated Proton Therapy

ResearchDGX agent

arXiv:2608.10173v1 Announce Type: new Abstract: Most radiotherapy dose-prediction models use only CT images and anatomical structures, although intensity-modulated proton therapy (IMPT) dose also depe

DreamOmni3: Scribble-based Editing and Generation

Model ReleasesDGX agent

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars

TutorialsDGX agent

arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis

DynaPPI: A Large-scale Dynamic Protein Dataset for AI-driven Advances in Protein Interactomics

TutorialsDGX agent

arXiv:2608.10435v1 Announce Type: new Abstract: Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven bi

E^3mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

Model ReleasesDGX agent

arXiv:2608.10796v1 Announce Type: new Abstract: Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interact

Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive Perception

Local AiDGX agent

arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa

Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition

Local AiDGX agent

arXiv:2608.10684v1 Announce Type: new Abstract: Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods e

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

Local AiDGX agent

arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex

ENCORE: Efficient Noise Context-Aware Representation for Low-Dose CT Denoising

Local AiDGX agent

arXiv:2608.10343v1 Announce Type: new Abstract: While deep learning-based denoising has become widely adopted in low-dose CT, conventional models use generic architectures designed for natural images,

Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery

ResearchDGX agent

arXiv:2608.10801v1 Announce Type: new Abstract: Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segm

Every Packet Counts: Dispersing Information for Loss-Resilient Learned Image Compression

ResearchDGX agent

arXiv:2608.11096v1 Announce Type: new Abstract: Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss,

Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking

Model ReleasesDGX agent

arXiv:2507.21606v2 Announce Type: replace Abstract: The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous hu

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

Model ReleasesDGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

FARCLUSS: Fuzzy Adaptive Rebalancing and Contrastive Uncertainty Learning for Semi-Supervised Semantic Segmentation

ResearchDGX agent

arXiv:2506.11142v3 Announce Type: replace Abstract: Semi-supervised semantic segmentation (SSSS) faces persistent challenges in effectively leveraging unlabeled data, such as ineffective utilization o

Flex-pi: A Multi-Stream World-Action Model with Compute Flexibility

Model ReleasesDGX agent

arXiv:2608.10860v1 Announce Type: cross Abstract: World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction,

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

Model ReleasesDGX agent

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets

ResearchDGX agent

arXiv:2608.11076v1 Announce Type: new Abstract: Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotrac

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

HardwareDGX agent

arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection.

Gaussian Sculpting: End-to-End Controllable Surface Reconstruction via Field Optimization

TutorialsDGX agent

arXiv:2608.10602v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled real-time novel view synthesis with impressive quality. However, it struggles to recover accurate surf

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Model ReleasesDGX agent

arXiv:2608.10426v1 Announce Type: new Abstract: Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories sp

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

Model ReleasesDGX agent

arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

Local AiDGX agent

arXiv:2608.10723v1 Announce Type: new Abstract: Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher i

GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids

ResearchDGX agent

arXiv:2603.25467v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without c

GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting

ResearchDGX agent

arXiv:2608.10938v1 Announce Type: new Abstract: Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generaliza

HNDiff: Haze-Noise Diffusion for Image Dehazing

Model ReleasesDGX agent

arXiv:2608.10995v1 Announce Type: new Abstract: Existing diffusion-based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formatio

HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models

ResearchDGX agent

arXiv:2512.05746v3 Announce Type: replace Abstract: Diffusion models have demonstrated significant applications in the field of image generation. However, their high computational and memory costs pos

HUI360: A 360{eg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

Model ReleasesDGX agent

arXiv:2608.11051v1 Announce Type: new Abstract: As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware beh

Human versus Computer Vision

TutorialsDGX agent

arXiv:2608.10181v1 Announce Type: new Abstract: Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in

Implicit representations are dead. Long live explicit primitives!

Model ReleasesDGX agent

arXiv:2608.10001v1 Announce Type: cross Abstract: Continuous parameterization of medical data has emerged as a powerful paradigm for resolution-independent image representation. While Implicit Neural

InterPruner: Interactive Structured Pruning via Taylor-Implicit Criterion and Language-Prior Modulator for Multimodal Object Detection

ResearchDGX agent

arXiv:2608.10724v1 Announce Type: new Abstract: Multimodal object detection proves effective in remote sensing, especially the RGB-Infrared paradigm. The parallel feature extractors provide rich multi

Introspective Attention Modulation for Safe Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2608.11135v1 Announce Type: new Abstract: Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in r

Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension

ResearchDGX agent

arXiv:2608.10566v1 Announce Type: cross Abstract: How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping

Learning Gaussian Structure: Intervention-Guided Density Control for Feed-Forward Driving Reconstruction

Local AiDGX agent

arXiv:2608.11077v1 Announce Type: new Abstract: Feed-forward Gaussian reconstruction has recently emerged as an efficient approach for driving scene reconstruction. However, prevailing LiDAR-based met

Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

SafetyDGX agent

arXiv:2509.06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robot

LEGO: Leveled Language Gaussian Splatting

ResearchDGX agent

arXiv:2608.10057v1 Announce Type: new Abstract: We introduce LEGO for advanced open-vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsi

Lesion-Aware Adaptive Fourier Neural Operator for CT-to-PSMA PET Synthesis in Prostate Cancer

Local AiDGX agent

arXiv:2608.10429v1 Announce Type: new Abstract: Deep learning models that synthesize PET from CT or MRI can reduce patient dose and scanner demand, but are typically optimized with global losses such

← Previous
1
Next →
12,414 results
← Previous
123…207
Next →