AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
6 Aug 2026

OutLangSplat: 3D Language Gaussian Splatting for UAV Outdoor Scenes

SafetyDGX agent

arXiv:2608.04560v1 Announce Type: new Abstract: 3D Language Gaussian Splatting embeds open-vocabulary language features into 3D Gaussian Splatting, providing an efficient explicit representation for t

Overcoming Statistical Bias in Action-Controllable World Models

SafetyDGX agent

arXiv:2608.04653v1 Announce Type: new Abstract: Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable f

PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images

Model ReleasesDGX agent

arXiv:2608.04210v1 Announce Type: new Abstract: Pose-agnostic Anomaly Detection (PAD) remains challenging as anomalies can appear under arbitrary viewpoints, requiring methods to handle significant po


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Persistent Object Narratives for Token-Efficient Video Language Models

Model ReleasesDGX agent

arXiv:2608.04866v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have made strong progress in open-ended video understanding. However, their visual interfaces remain token-inte

Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models

SafetyDGX agent

arXiv:2608.04349v1 Announce Type: new Abstract: Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional

Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings

ResearchDGX agent

arXiv:2608.05132v1 Announce Type: new Abstract: Predicting how a subcortical structure's shape will evolve from a few prior scans could support prognosis and clinical-trial enrichment. Existing longit

Privacy-Preserving Action Recognition: Taxonomy, Methods, and Privacy-Utility Trade-offs

Model ReleasesDGX agent

arXiv:2608.04501v1 Announce Type: new Abstract: Video surveillance in public safety, healthcare, and smart environments has made continuous human monitoring routine, raising real risks to personal ide

Promptable Animal Pose Tracking Across Species

Local AiDGX agent

arXiv:2608.04995v1 Announce Type: new Abstract: Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated

Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

Model ReleasesDGX agent

arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modali

RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection

ResearchDGX agent

arXiv:2608.04833v1 Announce Type: new Abstract: RGB-infrared (RGB-IR) object detection benefits from complementary visible and thermal cues, but effective fusion remains challenging under illumination

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Model ReleasesDGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans

Model ReleasesDGX agent

arXiv:2508.14006v2 Announce Type: replace Abstract: We introduce ResPlan, a dataset of 17,000 residential floor plans with vector geometry, room-connectivity graphs, and metric-scale coordinates. Each

Rethinking Pixel Mean Flows via Interval Denoiser

ResearchDGX agent

arXiv:2608.04818v1 Announce Type: new Abstract: Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-ste

Revisiting Pose Sensitivity in Splat-based Computed Tomography under Sparse-view Reconstruction

ApplicationsDGX agent

arXiv:2608.04752v1 Announce Type: new Abstract: X-ray computed tomography (CT) reconstructs volumetric representations of objects from projection images obtained by transmitting X-rays through a targe

REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

ResearchDGX agent

arXiv:2608.04480v1 Announce Type: new Abstract: Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A comm

Robustness Emerges Early in Training Dynamics, but Is Not Preserved

Model ReleasesDGX agent

arXiv:2608.04442v1 Announce Type: cross Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

ResearchDGX agent

arXiv:2608.04132v1 Announce Type: new Abstract: High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the re

SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

ApplicationsDGX agent

arXiv:2608.04246v1 Announce Type: cross Abstract: Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel obje

SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for Unpaired RGB+Thermal 3D Reconstruction

Model ReleasesDGX agent

arXiv:2603.18774v2 Announce Type: replace Abstract: Foundational feed-forward visual geometry models enable accurate and efficient camera pose estimation and scene reconstruction by learning strong sc

Season: Spectrum-Aware Orthogonal Gradient Refinement for Transfer-Based Adversarial Attacks

ResearchDGX agent

arXiv:2608.04441v1 Announce Type: new Abstract: Transfer-based adversarial attacks often transfer poorly across heterogeneous architectures because CNNs favor local textures while Vision Transformers

Segmentation Pre-training for Label-Efficient Lumbar Spine Degeneration Grading

ResearchDGX agent

arXiv:2608.04810v1 Announce Type: new Abstract: Automated assessment of degenerative pathology in the lumbar spine on magnetic resonance imaging (MRI) requires access to large-scale datasets of expert

Semantic Frame Interpolation

Model ReleasesDGX agent

arXiv:2507.05173v2 Announce Type: replace Abstract: Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significan

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation

ResearchDGX agent

arXiv:2608.04196v1 Announce Type: cross Abstract: Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Model ReleasesDGX agent

arXiv:2608.05137v1 Announce Type: new Abstract: Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, incl

Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modeling

SafetyDGX agent

arXiv:2608.04764v1 Announce Type: new Abstract: X-ray computed tomography (CT) suffers from severe metal artifacts when high-attenuation objects such as dental fillings or orthopedic implants are pres

StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning

HardwareDGX agent

arXiv:2608.04811v1 Announce Type: new Abstract: Structured pruning enhances the efficiency of deep neural networks (DNNs) by eliminating groups of parameters during inference. Previous methods mostly

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

SafetyDGX agent

arXiv:2608.04887v1 Announce Type: new Abstract: On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student.

SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding

TutorialsDGX agent

arXiv:2608.04676v1 Announce Type: new Abstract: Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

Model ReleasesDGX agent

arXiv:2608.04568v1 Announce Type: new Abstract: As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inpu

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

Model ReleasesDGX agent

arXiv:2608.04127v1 Announce Type: new Abstract: Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and con

Thinking with Anchors: Grounded and Efficient Document Reasoning

Model ReleasesDGX agent

arXiv:2608.04424v1 Announce Type: new Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reaso

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

SafetyDGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

SafetyDGX agent

arXiv:2608.05000v1 Announce Type: new Abstract: Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, t

Towards Valid B-Rep Generation: Training-Free Wireframe Anomaly Detection and Repair

ResearchDGX agent

arXiv:2608.04955v1 Announce Type: new Abstract: Multi-stage boundary representation (B-Rep) generation leverages intermediate wireframes to synthesize CAD models. However, geometric and topological ri

Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

Model ReleasesDGX agent

arXiv:2608.04879v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is indepe

Transferable Dual-Stream Representations for Mesoscale-Preserving Sea Surface Temperature Downscaling

ResearchDGX agent

arXiv:2608.04230v1 Announce Type: cross Abstract: Deep learning models for scientific spatio-temporal downscaling often minimize reconstruction error while failing to preserve physically meaningful mu

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

ResearchDGX agent

arXiv:2608.04606v1 Announce Type: new Abstract: Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintainin

TriCLE: Tri-Modal Vision-Language Reasoning for Edge-Deployed Fine-Grained Clustering

SafetyDGX agent

arXiv:2608.04175v1 Announce Type: new Abstract: Edge platforms used for aerial observation must interpret aircraft imagery under limited memory, limited compute, and intermittent connectivity. This se

UBLLIE: Unified Backlight and Low-Light Image Enhancement

ApplicationsDGX agent

arXiv:2608.04429v1 Announce Type: new Abstract: Backlit and low-light images often suffer from severe exposure imbalance or global underexposure, presenting significant challenges for both visual perc

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Model ReleasesDGX agent

arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generati

Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

ResearchDGX agent

arXiv:2608.04935v1 Announce Type: new Abstract: Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI

Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

ResearchDGX agent

arXiv:2608.04623v1 Announce Type: new Abstract: Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

Model ReleasesDGX agent

arXiv:2608.04902v1 Announce Type: new Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

SafetyDGX agent

arXiv:2608.04557v1 Announce Type: new Abstract: High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric g

When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions

TutorialsDGX agent

arXiv:2608.04820v1 Announce Type: new Abstract: Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and

When Modalities Fail to Tango: Conformal Backdoor Detection in Multimodal Contrastive Learning

ResearchDGX agent

arXiv:2608.04052v1 Announce Type: cross Abstract: Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend

YOLO-PVC: 2D-to-3D Consolidation of Slice-wise Detections for Volumetric Liver Tumor Localization in MRI

ResearchDGX agent

arXiv:2608.04642v1 Announce Type: new Abstract: Slice-wise 2D object detectors are increasingly applied to volumetric data due to their computational efficiency and scalability, yet they often yield f

YOLOv14:Unified Cross-Domain Real-Time Object Detectionwith Adaptive Multi-View Representation

SafetyDGX agent

arXiv:2608.04720v1 Announce Type: new Abstract: Real-time object detectors achieve remarkable accuracy under controlled conditions, yet degrade sharply on non-ideal inputs: fisheye distortion, game-re

YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos

SafetyDGX agent

arXiv:2506.18266v2 Announce Type: replace Abstract: 3D semantic occupancy prediction is crucial for fine-grained scene understanding, yet its advancement in privacy-sensitive indoor environments is fu

5 Aug 2026

3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment

Model ReleasesDGX agent

arXiv:2608.03279v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become a dominant representation for real-time novel view synthesis (NVS), yet its storage footprint makes compression

A Human-in-the-Loop Deep Learning Framework for Color Reconstruction of Lenticular Films

ResearchDGX agent

arXiv:2608.02835v1 Announce Type: new Abstract: Historical lenticular films, such as those created with the Kodacolor process, encode color information in a distinctive spatial format. This structure

A Unified Resolution-Conditioned Framework for Orthogonal Line-Scanning Image Fusion

ResearchDGX agent

arXiv:2608.03107v1 Announce Type: new Abstract: Laser line-scanning microscopy enables fast volumetric imaging but produces anisotropic lateral resolution. Orthogonal line scans provide complementary

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

Local AiDGX agent

arXiv:2608.03779v1 Announce Type: new Abstract: Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences,

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

ResearchDGX agent

arXiv:2608.03047v1 Announce Type: new Abstract: Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-b

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

Local AiDGX agent

arXiv:2608.03316v1 Announce Type: cross Abstract: On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

ResearchDGX agent

arXiv:2608.02791v1 Announce Type: new Abstract: MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-pre

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

AgentsDGX agent

arXiv:2608.03571v1 Announce Type: new Abstract: Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal env

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

Model ReleasesDGX agent

arXiv:2607.11844v2 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos inv

Bridging Online and Offline Handwriting via Differentiable Physical Rendering

Model ReleasesDGX agent

arXiv:2608.03198v1 Announce Type: new Abstract: Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calli

Can Text-to-Image Models Draw from the Right Frame of Reference?

Model ReleasesDGX agent

arXiv:2608.03357v1 Announce Type: new Abstract: Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expression

← Previous
1…1011121314…207
Next →