AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration

DGX agent

arXiv:2608.06981v1 Announce Type: new Abstract: Diffusion models have demonstrated remarkable effectiveness in image restoration tasks. However, when guiding image reconstruction, existing Diffusion M

researcharxiv-cs-cv
10 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

DGX agent

arXiv:2608.07463v1 Announce Type: new Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging

model-releasesarxiv-cs-cv
10 Aug 2026
Research

MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding

DGX agent

arXiv:2506.07136v2 Announce Type: replace Abstract: First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the app

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Multiple Hypothesis Flow Estimation for Video Frame Interpolation under Matching Ambiguity

DGX agent

arXiv:2608.07120v1 Announce Type: new Abstract: Many flow-based video frame interpolation (VFI) methods synthesize an intermediate frame by estimating optical flow fields, warping the two input frames

model-releasesarxiv-cs-cv
10 Aug 2026
Tutorials

MuST-VAD: Mutual Structured Learning for Video Anomaly Detection

DGX agent

arXiv:2608.06913v1 Announce Type: new Abstract: In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detec

tutorialsarxiv-cs-cv
10 Aug 2026
Research

oldsymbol{lambda}-Orthogonality Regularization for Compatible Representation Learning

DGX agent

arXiv:2509.16664v2 Announce Type: cross Abstract: Retrieval systems rely on representations learned by increasingly powerful models. However, due to the high training cost and inconsistencies in learn

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots

DGX agent

arXiv:2603.13108v2 Announce Type: replace-cross Abstract: Panoramic imagery provides holistic 360{eg} visual coverage for environmental perception in quadruped robots. However, existing occupancy pred

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

DGX agent

arXiv:2608.06794v1 Announce Type: new Abstract: While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objecti

safetyarxiv-cs-cv
10 Aug 2026
Research

Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings

DGX agent

arXiv:2511.23204v2 Announce Type: replace Abstract: Pathology foundation models (FMs) have driven significant progress in computational pathology. However, these high-performing models can easily exce

researcharxiv-cs-cv
10 Aug 2026
Research

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

DGX agent

arXiv:2608.06901v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidl

researcharxiv-cs-cv
10 Aug 2026
Safety

R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim

DGX agent

arXiv:2608.06827v1 Announce Type: cross Abstract: Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-envi

safetyarxiv-cs-cv
10 Aug 2026
Model Releases

RegionDet: A Benchmark for Region Detection Beyond Object Instances

DGX agent

arXiv:2608.06850v1 Announce Type: new Abstract: Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bo

model-releasesarxiv-cs-cv
10 Aug 2026
Research

Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT

DGX agent

arXiv:2512.12236v2 Announce Type: replace-cross Abstract: Sparse-view Computed Tomography (CT) reconstructs images from a limited number of X-ray projections to reduce radiation and scanning time, whi

researcharxiv-cs-cv
10 Aug 2026
Safety

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

DGX agent

arXiv:2608.06914v1 Announce Type: new Abstract: Rib fractures are common, clinically significant, and time-consuming to localize on computed tomography (CT). We ask whether fractures detected in two o

safetyarxiv-cs-cv
10 Aug 2026
Research

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

DGX agent

arXiv:2608.07012v1 Announce Type: new Abstract: Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a

researcharxiv-cs-cv
10 Aug 2026
Agents

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

DGX agent

arXiv:2608.07468v1 Announce Type: new Abstract: World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods requir

agentsarxiv-cs-cv
10 Aug 2026
Model Releases

SkySeaLand: A Wide-Format Satellite Transportation Benchmark with an Ultra-Lightweight Detection Baseline

DGX agent

arXiv:2608.07382v1 Announce Type: new Abstract: Satellite object detection is challenged by small targets and wide-format scenes that lose detail under standard square-input resizing. We introduce Sky

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

SparseVoxelDet: Fully Sparse Voxel Networks for Efficient Event-Based Drone Detection

DGX agent

arXiv:2603.21638v2 Announce Type: replace Abstract: Event cameras excel at detecting small, fast drones, but today's detectors give away their key advantage: they convert the sparse event stream into

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

DGX agent

arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing

model-releasesarxiv-cs-cv
10 Aug 2026
Research

SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching

DGX agent

arXiv:2608.06408v1 Announce Type: cross Abstract: Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchroniza

researcharxiv-cs-cv
10 Aug 2026
Hardware

Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

DGX agent

arXiv:2608.06959v1 Announce Type: new Abstract: Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downl

hardwarearxiv-cs-cv
10 Aug 2026
Model Releases

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

DGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Symbolic Graphics Programming with Large Language Models

DGX agent

arXiv:2509.05208v2 Announce Type: replace Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise vis

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

DGX agent

arXiv:2608.07314v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement lea

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition

DGX agent

arXiv:2608.06467v1 Announce Type: new Abstract: Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

Toward surface-based registration of a virtual preoperative cutting guide onto the mandible for reconstruction surgery

DGX agent

arXiv:2608.06599v1 Announce Type: new Abstract: Mandibular reconstruction restores facial continuity and oral function after segmental resection. Patient-specific cutting guides transfer a computed to

safetyarxiv-cs-cv
10 Aug 2026
Model Releases

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

DGX agent

arXiv:2608.06404v1 Announce Type: new Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and manage

model-releasesarxiv-cs-cv
10 Aug 2026
Research

Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection

DGX agent

arXiv:2608.07015v1 Announce Type: new Abstract: Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains an

researcharxiv-cs-cv
10 Aug 2026
Research

UniCycleFlow: Bidirectional Unpaired Image Translation with a Shared Rectified Flow

DGX agent

arXiv:2608.06784v1 Announce Type: new Abstract: Bidirectional unpaired image translation must preserve source-specific structure while learning coherent transformations in both directions without pair

researcharxiv-cs-cv
10 Aug 2026
Tutorials

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

DGX agent

arXiv:2608.07409v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent s

tutorialsarxiv-cs-cv
10 Aug 2026
Model Releases

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

DGX agent

arXiv:2511.01295v3 Announce Type: replace Abstract: Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still str

model-releasesarxiv-cs-cv
10 Aug 2026
Research

Vernata: Self-Supervised Learning of LiDAR Point Representations

DGX agent

arXiv:2608.06919v1 Announce Type: new Abstract: LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain

researcharxiv-cs-cv
10 Aug 2026
Research

Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning

DGX agent

arXiv:2608.06934v1 Announce Type: cross Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and prefere

researcharxiv-cs-cv
10 Aug 2026
Applications

WaveFreqAnchor: Wave-Structural Anchoring and Frequency Correction Diffusion for Training-Free Face Restoration

DGX agent

arXiv:2608.06717v1 Announce Type: new Abstract: Diffusion-based face restoration that adjusts the sampling trajectory of pre-trained diffusion models has achieved remarkable progress. However, existin

applicationsarxiv-cs-cv
10 Aug 2026
Research

When One Modality Is Not Enough: Multimodal Sex and Life-Stage Classification of Red Deer from Aerial RGB-Thermal Video

DGX agent

arXiv:2608.06973v1 Announce Type: new Abstract: Aerial drone surveys increasingly support wildlife population estimation, yet a useful census is more than a count: population dynamics are defined by s

researcharxiv-cs-cv
10 Aug 2026
Research

When Semantics Saturate or Emerge: Adaptation-Conditional Semantic Utility in Source-Free Cross-Domain Few-Shot Learning

DGX agent

arXiv:2608.06673v1 Announce Type: new Abstract: Language descriptions in source-free cross-domain few-shot learning (SF-CDFSL) are often selected according to zero-shot accuracy obtained with a frozen

researcharxiv-cs-cv
10 Aug 2026
Model Releases

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

DGX agent

arXiv:2608.07051v1 Announce Type: new Abstract: Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous op

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

DGX agent

arXiv:2510.06612v2 Announce Type: replace Abstract: Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in En

safetyarxiv-cs-cv
7 Aug 2026
Research

A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets

DGX agent

arXiv:2608.05471v1 Announce Type: cross Abstract: Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy-restricted, and hard-to-annotate datasets

researcharxiv-cs-cv
7 Aug 2026
Research

A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence

DGX agent

arXiv:2608.05570v1 Announce Type: cross Abstract: 360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardw

researcharxiv-cs-cv
7 Aug 2026
Model Releases

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

DGX agent

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

model-releasesarxiv-cs-cv
7 Aug 2026
Research

A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization

DGX agent

arXiv:2608.05217v1 Announce Type: cross Abstract: Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets.

researcharxiv-cs-cv
7 Aug 2026
Applications

Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery

DGX agent

arXiv:2608.05840v1 Announce Type: new Abstract: Accurate vehicle localization from monocular roadside surveillance cameras is important for intelligent transportation systems, traffic monitoring, and

applicationsarxiv-cs-cv
7 Aug 2026
Research

Active View Selection for Scene-level Multi-view Crowd Counting and Localization with Limited Labeling Budget

DGX agent

arXiv:2509.16684v2 Announce Type: replace Abstract: Multi-view crowd counting and localization fuse the input multi-views for estimating the crowd number or locations on the ground. Existing methods m

researcharxiv-cs-cv
7 Aug 2026
Model Releases

Adapting Vision Foundation Models with Cascaded Semantics

DGX agent

arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapt

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

DGX agent

arXiv:2503.03556v3 Announce Type: replace Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and

model-releasesarxiv-cs-cv
7 Aug 2026
Tutorials

ALTER: Modeling Longitudinal Changes via Regional Differencing for 3D CT Report Generation

DGX agent

arXiv:2608.05615v1 Announce Type: new Abstract: Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow-up, yet automatically generating accurate and complete radiology

tutorialsarxiv-cs-cv
7 Aug 2026
Tutorials

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture

DGX agent

arXiv:2608.06062v1 Announce Type: new Abstract: Bar charts are commonly used in data visualization, and while they are easily understood by humans, it is non-trivial to extract the underlying data com

tutorialsarxiv-cs-cv
7 Aug 2026
← Previous
1…89101112…259
Next →