AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
10 Aug 2026

HazeSpikeMamba: Coupling Spiking-Inspired and State-Space Features for Self-Supervised Real-World Dehazing

Local AiDGX agent

arXiv:2608.06886v1 Announce Type: new Abstract: Dehazing networks are commonly trained on synthetic hazy-clear pairs, but their performance often drops on real photographs. Synthetic haze generated us

HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models

ResearchDGX agent

arXiv:2608.07003v1 Announce Type: new Abstract: Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primari

Human-AI Perceptual Alignment by Playing Hues and Cues

Local AiDGX agent

arXiv:2608.07141v1 Announce Type: new Abstract: Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks tha


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

IceHorizon: A Dataset for Horizon Detection in Ice-Covered Maritime Environments and Comparative Evaluation of Detection Methods

ResearchDGX agent

arXiv:2608.07018v1 Announce Type: cross Abstract: Horizon detection in images of ice-covered waters is a challenging problem for maritime navigation due to low contrast between water and sky, cluttere

Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation

SafetyDGX agent

arXiv:2603.17889v4 Announce Type: replace Abstract: Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerg

Implicit Neural Speckle Denoising

ResearchDGX agent

arXiv:2608.06574v1 Announce Type: cross Abstract: Speckle fundamentally limits coherent imaging by introducing multiplicative, spatially correlated noise that obscures scene structure. Removing speckl

Improving Low-Resolution Face Recognition under Limited Data: How Synthetic Data Generation Can Close the Domain Gap

ApplicationsDGX agent

arXiv:2608.06580v1 Announce Type: new Abstract: Face Recognition (FR) systems in surveillance settings often encounter Low Resolution (LR) faces, those whose face region falls below the standard 112 i

InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion

Model ReleasesDGX agent

arXiv:2608.06490v1 Announce Type: new Abstract: We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise

InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding

TutorialsDGX agent

arXiv:2608.07144v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene underst

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

SafetyDGX agent

arXiv:2608.06799v1 Announce Type: cross Abstract: Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-

KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement

SafetyDGX agent

arXiv:2608.07057v1 Announce Type: new Abstract: Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identific

Learning Ordinal Degradation Representations with Textual Priors for Diffusion-Based Blind Image Super-Resolution

ApplicationsDGX agent

arXiv:2512.10340v2 Announce Type: replace Abstract: Blind image super-resolution (Blind SR) has achieved remarkable perceptual quality via generative priors. However, lacking clear degradation represe

Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration

ResearchDGX agent

arXiv:2608.06981v1 Announce Type: new Abstract: Diffusion models have demonstrated remarkable effectiveness in image restoration tasks. However, when guiding image reconstruction, existing Diffusion M

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Model ReleasesDGX agent

arXiv:2608.07463v1 Announce Type: new Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging

MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding

ResearchDGX agent

arXiv:2506.07136v2 Announce Type: replace Abstract: First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the app

Multiple Hypothesis Flow Estimation for Video Frame Interpolation under Matching Ambiguity

Model ReleasesDGX agent

arXiv:2608.07120v1 Announce Type: new Abstract: Many flow-based video frame interpolation (VFI) methods synthesize an intermediate frame by estimating optical flow fields, warping the two input frames

MuST-VAD: Mutual Structured Learning for Video Anomaly Detection

TutorialsDGX agent

arXiv:2608.06913v1 Announce Type: new Abstract: In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detec

oldsymbol{lambda}-Orthogonality Regularization for Compatible Representation Learning

ResearchDGX agent

arXiv:2509.16664v2 Announce Type: cross Abstract: Retrieval systems rely on representations learned by increasingly powerful models. However, due to the high training cost and inconsistencies in learn

Panoramic Multimodal Semantic Occupancy Prediction for Quadruped Robots

Model ReleasesDGX agent

arXiv:2603.13108v2 Announce Type: replace-cross Abstract: Panoramic imagery provides holistic 360{eg} visual coverage for environmental perception in quadruped robots. However, existing occupancy pred

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

SafetyDGX agent

arXiv:2608.06794v1 Announce Type: new Abstract: While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objecti

Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings

ResearchDGX agent

arXiv:2511.23204v2 Announce Type: replace Abstract: Pathology foundation models (FMs) have driven significant progress in computational pathology. However, these high-performing models can easily exce

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

ResearchDGX agent

arXiv:2608.06901v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidl

R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim

SafetyDGX agent

arXiv:2608.06827v1 Announce Type: cross Abstract: Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-envi

RegionDet: A Benchmark for Region Detection Beyond Object Instances

Model ReleasesDGX agent

arXiv:2608.06850v1 Announce Type: new Abstract: Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bo

Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT

ResearchDGX agent

arXiv:2512.12236v2 Announce Type: replace-cross Abstract: Sparse-view Computed Tomography (CT) reconstructs images from a limited number of X-ray projections to reduce radiation and scanning time, whi

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

SafetyDGX agent

arXiv:2608.06914v1 Announce Type: new Abstract: Rib fractures are common, clinically significant, and time-consuming to localize on computed tomography (CT). We ask whether fractures detected in two o

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

ResearchDGX agent

arXiv:2608.07012v1 Announce Type: new Abstract: Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2608.07468v1 Announce Type: new Abstract: World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods requir

SkySeaLand: A Wide-Format Satellite Transportation Benchmark with an Ultra-Lightweight Detection Baseline

Model ReleasesDGX agent

arXiv:2608.07382v1 Announce Type: new Abstract: Satellite object detection is challenged by small targets and wide-format scenes that lose detail under standard square-input resizing. We introduce Sky

SparseVoxelDet: Fully Sparse Voxel Networks for Efficient Event-Based Drone Detection

Model ReleasesDGX agent

arXiv:2603.21638v2 Announce Type: replace Abstract: Event cameras excel at detecting small, fast drones, but today's detectors give away their key advantage: they convert the sparse event stream into

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

Model ReleasesDGX agent

arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing

SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching

ResearchDGX agent

arXiv:2608.06408v1 Announce Type: cross Abstract: Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchroniza

Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

HardwareDGX agent

arXiv:2608.06959v1 Announce Type: new Abstract: Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downl

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

Model ReleasesDGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

Symbolic Graphics Programming with Large Language Models

Model ReleasesDGX agent

arXiv:2509.05208v2 Announce Type: replace Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise vis

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07314v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement lea

Test-Time Adaptation with Online Personalized Energy-Based Cache for Fine-Grained Video Expression Recognition

Model ReleasesDGX agent

arXiv:2608.06467v1 Announce Type: new Abstract: Facial expression recognition (FER) in videos is challenging because models must identify subtle, temporally evolving affective states that vary across

Toward surface-based registration of a virtual preoperative cutting guide onto the mandible for reconstruction surgery

SafetyDGX agent

arXiv:2608.06599v1 Announce Type: new Abstract: Mandibular reconstruction restores facial continuity and oral function after segmental resection. Patient-specific cutting guides transfer a computed to

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

Model ReleasesDGX agent

arXiv:2608.06404v1 Announce Type: new Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and manage

Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection

ResearchDGX agent

arXiv:2608.07015v1 Announce Type: new Abstract: Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains an

UniCycleFlow: Bidirectional Unpaired Image Translation with a Shared Rectified Flow

ResearchDGX agent

arXiv:2608.06784v1 Announce Type: new Abstract: Bidirectional unpaired image translation must preserve source-specific structure while learning coherent transformations in both directions without pair

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

TutorialsDGX agent

arXiv:2608.07409v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent s

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

Model ReleasesDGX agent

arXiv:2511.01295v3 Announce Type: replace Abstract: Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still str

Vernata: Self-Supervised Learning of LiDAR Point Representations

ResearchDGX agent

arXiv:2608.06919v1 Announce Type: new Abstract: LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain

Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning

ResearchDGX agent

arXiv:2608.06934v1 Announce Type: cross Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and prefere

WaveFreqAnchor: Wave-Structural Anchoring and Frequency Correction Diffusion for Training-Free Face Restoration

ApplicationsDGX agent

arXiv:2608.06717v1 Announce Type: new Abstract: Diffusion-based face restoration that adjusts the sampling trajectory of pre-trained diffusion models has achieved remarkable progress. However, existin

When One Modality Is Not Enough: Multimodal Sex and Life-Stage Classification of Red Deer from Aerial RGB-Thermal Video

ResearchDGX agent

arXiv:2608.06973v1 Announce Type: new Abstract: Aerial drone surveys increasingly support wildlife population estimation, yet a useful census is more than a count: population dynamics are defined by s

When Semantics Saturate or Emerge: Adaptation-Conditional Semantic Utility in Source-Free Cross-Domain Few-Shot Learning

ResearchDGX agent

arXiv:2608.06673v1 Announce Type: new Abstract: Language descriptions in source-free cross-domain few-shot learning (SF-CDFSL) are often selected according to zero-shot accuracy obtained with a frozen

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Model ReleasesDGX agent

arXiv:2608.07051v1 Announce Type: new Abstract: Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous op

7 Aug 2026

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

SafetyDGX agent

arXiv:2510.06612v2 Announce Type: replace Abstract: Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in En

A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets

ResearchDGX agent

arXiv:2608.05471v1 Announce Type: cross Abstract: Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy-restricted, and hard-to-annotate datasets

A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence

ResearchDGX agent

arXiv:2608.05570v1 Announce Type: cross Abstract: 360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardw

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

Model ReleasesDGX agent

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization

ResearchDGX agent

arXiv:2608.05217v1 Announce Type: cross Abstract: Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets.

Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery

ApplicationsDGX agent

arXiv:2608.05840v1 Announce Type: new Abstract: Accurate vehicle localization from monocular roadside surveillance cameras is important for intelligent transportation systems, traffic monitoring, and

Active View Selection for Scene-level Multi-view Crowd Counting and Localization with Limited Labeling Budget

ResearchDGX agent

arXiv:2509.16684v2 Announce Type: replace Abstract: Multi-view crowd counting and localization fuse the input multi-views for estimating the crowd number or locations on the ground. Existing methods m

Adapting Vision Foundation Models with Cascaded Semantics

Model ReleasesDGX agent

arXiv:2608.05393v1 Announce Type: new Abstract: Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapt

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

Model ReleasesDGX agent

arXiv:2503.03556v3 Announce Type: replace Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and

ALTER: Modeling Longitudinal Changes via Regional Differencing for 3D CT Report Generation

TutorialsDGX agent

arXiv:2608.05615v1 Announce Type: new Abstract: Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow-up, yet automatically generating accurate and complete radiology

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture

TutorialsDGX agent

arXiv:2608.06062v1 Announce Type: new Abstract: Bar charts are commonly used in data visualization, and while they are easily understood by humans, it is non-trivial to extract the underlying data com

← Previous
1…678910…207
Next →