AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Jun 2026

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration

Model ReleasesDGX agent

arXiv:2606.28049v1 Announce Type: new Abstract: In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometric

An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models

ResearchDGX agent

arXiv:2604.00784v2 Announce Type: replace Abstract: Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently be

An Embedded Real-Time License Plate Recognition System for Complex Traffic Scenes

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.27772v1 Announce Type: new Abstract: Vehicle license plate recognition is an integral component of intelligent transportation systems. In this work, we present an embedded real-time license

An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations

ResearchDGX agent

arXiv:2407.01014v2 Announce Type: replace Abstract: Diffusion models excel in solving imaging inverse problems due to their ability to model complex image priors. However, their reliance on large, cle

Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling

ResearchDGX agent

arXiv:2606.27547v1 Announce Type: new Abstract: Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, re

Beyond Points: Spherical Distributional Part Prototypes for Interpretable Classification

ResearchDGX agent

arXiv:2606.27582v1 Announce Type: new Abstract: Prototype-based neural networks aim to provide intrinsic interpretability by grounding predictions in a small set of part prototypes. However, modern vi

Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding

Local AiDGX agent

arXiv:2603.10863v2 Announce Type: replace Abstract: Despite the remarkable capabilities of Multimodal Large Language Models (MLLMs), they still suffer from visual fading in long-context scenarios. Spe

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

AgentsDGX agent

arXiv:2606.27644v1 Announce Type: new Abstract: This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for a

Complex-Valued 2D Gaussian Representation for Computer-Generated Holography

Model ReleasesDGX agent

arXiv:2511.15022v2 Announce Type: replace Abstract: Complex-valued Gaussian primitives have recently been explored for representing holographic radiance fields in 3D novel view synthesis. In this work

Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding

SafetyDGX agent

arXiv:2604.02546v2 Announce Type: replace Abstract: Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable r

Controllable Histopathology Image Synthesis with Training-free Structural Initialization and Textural Modulation

ResearchDGX agent

arXiv:2606.27935v1 Announce Type: new Abstract: Deep learning has demonstrated remarkable success in high-throughput histopathology image analysis. However, the performance of learning-based models cr

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

ResearchDGX agent

arXiv:2606.28104v1 Announce Type: new Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action

CSD: Content-aware Speculative Decoding for Efficient Image Generation

SafetyDGX agent

arXiv:2606.27829v1 Announce Type: new Abstract: Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation,

Curriculum-guided Change Detection Training: Toward Accurate Serac Fall Monitoring

Model ReleasesDGX agent

arXiv:2606.28012v1 Announce Type: new Abstract: Change Detection (CD) aims to identify semantic or structural changes from nearly registered multi-temporal images. While recent advances in training me

DeLux: Cross-Modal Local Artifact Restoration in Video Using Neuromorphic Data

TutorialsDGX agent

arXiv:2606.27576v1 Announce Type: new Abstract: Conventional RGB cameras suffer from lighting artifacts such as flare, glare, flicker, and overexposure, leading to irrecoverable information loss that

Denoising ICF Images with Multiplicative Uniform Noise: A Self-Supervised Study Based on the Log-Domain Noisier2Inverse Framework

ResearchDGX agent

arXiv:2606.27635v1 Announce Type: new Abstract: This paper documents the implementation and evaluation of a self-supervised denoising framework on Inertial Confinement Fusion (ICF) images corrupted by

Differentiable design of the PIAA-ZWFS: a flexible wavefront sensor that approaches the fundamental limit

ResearchDGX agent

arXiv:2606.28136v1 Announce Type: cross Abstract: Extreme adaptive optics (AO) is necessary for high contrast astronomy at scales of the habitable zone of nearby systems. We seek to evaluate wavefront

Diffusion Model Attribution via Spectral Coupling of Denoiser Responses

ResearchDGX agent

arXiv:2606.28092v1 Announce Type: new Abstract: Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection.

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

Local AiDGX agent

arXiv:2606.27677v1 Announce Type: cross Abstract: World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing meth

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

SafetyDGX agent

arXiv:2606.27964v1 Announce Type: new Abstract: Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video

EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography

SafetyDGX agent

arXiv:2606.28164v1 Announce Type: new Abstract: Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreti

EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

ResearchDGX agent

arXiv:2606.28026v1 Announce Type: new Abstract: High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods f

Enhanced Neural Video Representation Compression across Extreme Complexity and Quality Scales

ApplicationsDGX agent

arXiv:2606.28163v1 Announce Type: cross Abstract: Implicit neural representations (INRs) have recently emerged as a promising approach to video compression, delivering competitive rate-distortion perf

Enhancing Co-packaging Optics Enabled Silicon Photonics Security Assurance Hardware Fingerprinting

HardwareDGX agent

arXiv:2606.27612v1 Announce Type: cross Abstract: Silicon photonics enables integration of optical components using standard semiconductor processes, greatly improving data communication bandwidth and

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

ResearchDGX agent

arXiv:2603.25168v2 Announce Type: replace Abstract: Previous works based on Segment Anything Model (SAM) have achieved promising performance in unified scene text detection and layout analysis. Howeve

Fine-Grained Behavior and Lane Constraints Guided Trajectory Prediction Method

AgentsDGX agent

arXiv:2503.21477v3 Announce Type: replace Abstract: Trajectory prediction, as a critical component of autonomous driving systems, has attracted the attention of many researchers. Existing prediction a

Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

Model ReleasesDGX agent

arXiv:2606.27484v1 Announce Type: new Abstract: Autism spectrum disorder (ASD) affects 1 in 31 US children, yet median age at diagnosis exceeds four years. Artificial intelligence pipelines that provi

GeoFace: Consistent Multi-View Face Generation with Geometry-Constrained Diffusion

SafetyDGX agent

arXiv:2606.27659v1 Announce Type: new Abstract: We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view

Graph Unfolding and Sampling for Transitory Video Keyframe Selection via Gershgorin Disc Alignment

SafetyDGX agent

arXiv:2408.01859v2 Announce Type: replace Abstract: User-generated videos (UGVs) uploaded from mobile phones to social media sites like YouTube and TikTok are short and non-repetitive. We summarize a

GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving

AgentsDGX agent

arXiv:2511.11266v4 Announce Type: replace Abstract: Vision-language models have recently emerged as promising planners for autonomous driving, where success hinges on topology-aware reasoning over spa

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

Model ReleasesDGX agent

arXiv:2606.27999v1 Announce Type: new Abstract: Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models

HunyuanImage 3.0 Technical Report

SafetyDGX agent

arXiv:2509.23951v3 Announce Type: replace Abstract: We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with

Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?

Local AiDGX agent

arXiv:2501.16947v2 Announce Type: replace Abstract: The advances in Vision-Language models (VLMs) offer exciting opportunities for robotic applications involving image geo-localization - the problem o

Instant Expressive Gaussian Head Avatars at Over 100 FPS

ApplicationsDGX agent

arXiv:2512.16893v2 Announce Type: replace Abstract: Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods ofte

Latent Visual Diffusion Reasoning with Monte Carlo Tree Search

TutorialsDGX agent

arXiv:2606.27988v1 Announce Type: new Abstract: Analyzing fine-grained skill activities (e.g., sports, surgery) requires not only recognizing visual patterns but also performing step-by-step visual re

Learning 1-Bit LiDAR-based Localization with Auxiliary Objective

AgentsDGX agent

arXiv:2606.27729v1 Announce Type: new Abstract: 6-DoF LiDAR-based localization is a fundamental capability for autonomous systems operating in large-scale outdoor environments. Many deep-learning-base

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

ResearchDGX agent

arXiv:2601.12066v4 Announce Type: replace Abstract: Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninfo

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

SafetyDGX agent

arXiv:2511.14271v2 Announce Type: replace Abstract: Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still

Long-Term Prediction of Local and Global Human Motion with Occlusion Recovery

Local AiDGX agent

arXiv:2606.27900v1 Announce Type: new Abstract: Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of

MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation

TutorialsDGX agent

arXiv:2606.27718v1 Announce Type: new Abstract: Video frame interpolation (VFI) remains a challenging task, particularly when dealing with large, non-linear motions and complex occlusions. While flow-

MeDUET: Disentangled Unified Pretraining for 3D Medical Image Synthesis and Analysis

TutorialsDGX agent

arXiv:2602.17901v3 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) and diffusion models have respectively advanced representation learning and generative modeling for high-dimens

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

Model ReleasesDGX agent

arXiv:2606.27537v1 Announce Type: new Abstract: Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most ass

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations

ResearchDGX agent

arXiv:2606.27779v1 Announce Type: new Abstract: Generating lifelike facial animation for dyadic conversations requires reconciling high-level cognitive intent with precise low-level motor reflexes, ye

ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On

SafetyDGX agent

arXiv:2606.27773v1 Announce Type: new Abstract: Image-based virtual try-on has emerged as a compelling task in e-commerce and augmented reality, yet existing methods struggle to simultaneously preserv

Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading

Model ReleasesDGX agent

arXiv:2606.28144v1 Announce Type: new Abstract: Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarc

Multi-Modal Conditioned High-Resolution Transformer for Urban Electromagnetic Field Map Prediction Download PDF

ResearchDGX agent

arXiv:2606.27671v1 Announce Type: new Abstract: Predicting electromagnetic field (EMF) strength in urban environments is essential for cellular network planning but computationally expensive with phys

MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis

ResearchDGX agent

arXiv:2410.02103v4 Announce Type: replace Abstract: Recent works in volume rendering, extit{e.g.} NeRF and 3D Gaussian Splatting (3DGS), significantly advance the rendering quality and efficiency with

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.27660v1 Announce Type: new Abstract: Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual tok

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

Local AiDGX agent

arXiv:2606.27771v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that a

On the stability of scale-space metrics

ResearchDGX agent

arXiv:2606.27605v1 Announce Type: cross Abstract: We study the stability of a classical family of metrics defined over functions' Gaussian scale-space representations, focusing on the comparison of im

OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation

Model ReleasesDGX agent

arXiv:2606.27880v1 Announce Type: new Abstract: Unified fashion generation integrates tasks like virtual try-on and garment reconstruction into a single model to reduce task-specific adaptation costs.

Panoramic Scene Analysis: A Survey from Distortion-Aware Engineering to Sphere-Native Foundation Modeling

ResearchDGX agent

arXiv:2606.27745v1 Announce Type: new Abstract: Panoramic images capture the complete visual sphere in a single frame, providing spatial context unattainable by conventional cameras. Yet this complete

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Model ReleasesDGX agent

arXiv:2606.28322v1 Announce Type: new Abstract: We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness

Perceptual 3D Simulation With Physical World Modeling

Local AiDGX agent

arXiv:2606.27575v1 Announce Type: new Abstract: Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal s

Permutation Learning with Only N Parameters: From SoftSort to Self-Organizing Gaussians

ResearchDGX agent

arXiv:2503.13051v3 Announce Type: replace-cross Abstract: Sorting and permutation learning are key concepts in optimization and machine learning, especially when organizing high-dimensional data into

PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding

TutorialsDGX agent

arXiv:2511.20562v2 Announce Type: replace Abstract: While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability

PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion

ResearchDGX agent

arXiv:2606.27760v1 Announce Type: new Abstract: End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency seman

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2603.19466v2 Announce Type: replace Abstract: Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone

QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception

Model ReleasesDGX agent

arXiv:2509.03704v2 Announce Type: replace Abstract: Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating

Qwen-Image-2.0-RL Technical Report

Model ReleasesDGX agent

arXiv:2606.27608v1 Announce Type: new Abstract: We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) t

← Previous
1…6566676869…209
Next →