AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Rethinking Pixel Mean Flows via Interval Denoiser

DGX agent

arXiv:2608.04818v1 Announce Type: new Abstract: Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-ste

researcharxiv-cs-cv
6 Aug 2026
Applications

Revisiting Pose Sensitivity in Splat-based Computed Tomography under Sparse-view Reconstruction

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2608.04752v1 Announce Type: new Abstract: X-ray computed tomography (CT) reconstructs volumetric representations of objects from projection images obtained by transmitting X-rays through a targe

applicationsarxiv-cs-cv
6 Aug 2026
Research

REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

DGX agent

arXiv:2608.04480v1 Announce Type: new Abstract: Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A comm

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Robustness Emerges Early in Training Dynamics, but Is Not Preserved

DGX agent

arXiv:2608.04442v1 Announce Type: cross Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon

model-releasesarxiv-cs-cv
6 Aug 2026
Research

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

DGX agent

arXiv:2608.04132v1 Announce Type: new Abstract: High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the re

researcharxiv-cs-cv
6 Aug 2026
Applications

SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

DGX agent

arXiv:2608.04246v1 Announce Type: cross Abstract: Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel obje

applicationsarxiv-cs-cv
6 Aug 2026
Model Releases

SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for Unpaired RGB+Thermal 3D Reconstruction

DGX agent

arXiv:2603.18774v2 Announce Type: replace Abstract: Foundational feed-forward visual geometry models enable accurate and efficient camera pose estimation and scene reconstruction by learning strong sc

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Season: Spectrum-Aware Orthogonal Gradient Refinement for Transfer-Based Adversarial Attacks

DGX agent

arXiv:2608.04441v1 Announce Type: new Abstract: Transfer-based adversarial attacks often transfer poorly across heterogeneous architectures because CNNs favor local textures while Vision Transformers

researcharxiv-cs-cv
6 Aug 2026
Research

Segmentation Pre-training for Label-Efficient Lumbar Spine Degeneration Grading

DGX agent

arXiv:2608.04810v1 Announce Type: new Abstract: Automated assessment of degenerative pathology in the lumbar spine on magnetic resonance imaging (MRI) requires access to large-scale datasets of expert

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Semantic Frame Interpolation

DGX agent

arXiv:2507.05173v2 Announce Type: replace Abstract: Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significan

model-releasesarxiv-cs-cv
6 Aug 2026
Research

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation

DGX agent

arXiv:2608.04196v1 Announce Type: cross Abstract: Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually

researcharxiv-cs-cv
6 Aug 2026
Model Releases

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

DGX agent

arXiv:2608.05137v1 Announce Type: new Abstract: Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, incl

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modeling

DGX agent

arXiv:2608.04764v1 Announce Type: new Abstract: X-ray computed tomography (CT) suffers from severe metal artifacts when high-attenuation objects such as dental fillings or orthopedic implants are pres

safetyarxiv-cs-cv
6 Aug 2026
Hardware

StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning

DGX agent

arXiv:2608.04811v1 Announce Type: new Abstract: Structured pruning enhances the efficiency of deep neural networks (DNNs) by eliminating groups of parameters during inference. Previous methods mostly

hardwarearxiv-cs-cv
6 Aug 2026
Safety

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

DGX agent

arXiv:2608.04887v1 Announce Type: new Abstract: On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student.

safetyarxiv-cs-cv
6 Aug 2026
Tutorials

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding

DGX agent

arXiv:2608.04676v1 Announce Type: new Abstract: Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for

tutorialsarxiv-cs-cv
6 Aug 2026
Model Releases

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

DGX agent

arXiv:2608.04568v1 Announce Type: new Abstract: As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inpu

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

DGX agent

arXiv:2608.04127v1 Announce Type: new Abstract: Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and con

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Thinking with Anchors: Grounded and Efficient Document Reasoning

DGX agent

arXiv:2608.04424v1 Announce Type: new Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reaso

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

DGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

safetyarxiv-cs-cv
6 Aug 2026
Safety

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

DGX agent

arXiv:2608.05000v1 Announce Type: new Abstract: Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, t

safetyarxiv-cs-cv
6 Aug 2026
Research

Towards Valid B-Rep Generation: Training-Free Wireframe Anomaly Detection and Repair

DGX agent

arXiv:2608.04955v1 Announce Type: new Abstract: Multi-stage boundary representation (B-Rep) generation leverages intermediate wireframes to synthesize CAD models. However, geometric and topological ri

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

DGX agent

arXiv:2608.04879v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is indepe

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Transferable Dual-Stream Representations for Mesoscale-Preserving Sea Surface Temperature Downscaling

DGX agent

arXiv:2608.04230v1 Announce Type: cross Abstract: Deep learning models for scientific spatio-temporal downscaling often minimize reconstruction error while failing to preserve physically meaningful mu

researcharxiv-cs-cv
6 Aug 2026
Research

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

DGX agent

arXiv:2608.04606v1 Announce Type: new Abstract: Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintainin

researcharxiv-cs-cv
6 Aug 2026
Safety

TriCLE: Tri-Modal Vision-Language Reasoning for Edge-Deployed Fine-Grained Clustering

DGX agent

arXiv:2608.04175v1 Announce Type: new Abstract: Edge platforms used for aerial observation must interpret aircraft imagery under limited memory, limited compute, and intermittent connectivity. This se

safetyarxiv-cs-cv
6 Aug 2026
Applications

UBLLIE: Unified Backlight and Low-Light Image Enhancement

DGX agent

arXiv:2608.04429v1 Announce Type: new Abstract: Backlit and low-light images often suffer from severe exposure imbalance or global underexposure, presenting significant challenges for both visual perc

applicationsarxiv-cs-cv
6 Aug 2026
Model Releases

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

DGX agent

arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generati

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

DGX agent

arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI

researcharxiv-cs-cv
6 Aug 2026
Research

Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

DGX agent

arXiv:2608.04623v1 Announce Type: new Abstract: Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

DGX agent

arXiv:2608.04902v1 Announce Type: new Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

DGX agent

arXiv:2608.04557v1 Announce Type: new Abstract: High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric g

safetyarxiv-cs-cv
6 Aug 2026
Tutorials

When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions

DGX agent

arXiv:2608.04820v1 Announce Type: new Abstract: Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and

tutorialsarxiv-cs-cv
6 Aug 2026
Research

When Modalities Fail to Tango: Conformal Backdoor Detection in Multimodal Contrastive Learning

DGX agent

arXiv:2608.04052v1 Announce Type: cross Abstract: Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend

researcharxiv-cs-cv
6 Aug 2026
Research

YOLO-PVC: 2D-to-3D Consolidation of Slice-wise Detections for Volumetric Liver Tumor Localization in MRI

DGX agent

arXiv:2608.04642v1 Announce Type: new Abstract: Slice-wise 2D object detectors are increasingly applied to volumetric data due to their computational efficiency and scalability, yet they often yield f

researcharxiv-cs-cv
6 Aug 2026
Safety

YOLOv14:Unified Cross-Domain Real-Time Object Detectionwith Adaptive Multi-View Representation

DGX agent

arXiv:2608.04720v1 Announce Type: new Abstract: Real-time object detectors achieve remarkable accuracy under controlled conditions, yet degrade sharply on non-ideal inputs: fisheye distortion, game-re

safetyarxiv-cs-cv
6 Aug 2026
Safety

YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos

DGX agent

arXiv:2506.18266v2 Announce Type: replace Abstract: 3D semantic occupancy prediction is crucial for fine-grained scene understanding, yet its advancement in privacy-sensitive indoor environments is fu

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment

DGX agent

arXiv:2608.03279v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become a dominant representation for real-time novel view synthesis (NVS), yet its storage footprint makes compression

model-releasesarxiv-cs-cv
5 Aug 2026
Research

A Human-in-the-Loop Deep Learning Framework for Color Reconstruction of Lenticular Films

DGX agent

arXiv:2608.02835v1 Announce Type: new Abstract: Historical lenticular films, such as those created with the Kodacolor process, encode color information in a distinctive spatial format. This structure

researcharxiv-cs-cv
5 Aug 2026
Research

A Unified Resolution-Conditioned Framework for Orthogonal Line-Scanning Image Fusion

DGX agent

arXiv:2608.03107v1 Announce Type: new Abstract: Laser line-scanning microscopy enables fast volumetric imaging but produces anisotropic lateral resolution. Orthogonal line scans provide complementary

researcharxiv-cs-cv
5 Aug 2026
Local Ai

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

DGX agent

arXiv:2608.03779v1 Announce Type: new Abstract: Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences,

local-aiarxiv-cs-cv
5 Aug 2026
Research

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

DGX agent

arXiv:2608.03047v1 Announce Type: new Abstract: Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-b

researcharxiv-cs-cv
5 Aug 2026
Local Ai

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

DGX agent

arXiv:2608.03316v1 Announce Type: cross Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language

local-aiarxiv-cs-cv
5 Aug 2026
Research

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

DGX agent

arXiv:2608.02791v1 Announce Type: new Abstract: MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-pre

researcharxiv-cs-cv
5 Aug 2026
Agents

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

DGX agent

arXiv:2608.03571v1 Announce Type: new Abstract: Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal env

agentsarxiv-cs-cv
5 Aug 2026
Model Releases

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

DGX agent

arXiv:2607.11844v2 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos inv

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Bridging Online and Offline Handwriting via Differentiable Physical Rendering

DGX agent

arXiv:2608.03198v1 Announce Type: new Abstract: Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calli

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

Can Text-to-Image Models Draw from the Right Frame of Reference?

DGX agent

arXiv:2608.03357v1 Announce Type: new Abstract: Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expression

model-releasesarxiv-cs-cv
5 Aug 2026
← Previous
1…1314151617…259
Next →