AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Diffusion Model Attribution via Spectral Coupling of Denoiser Responses

DGX agent

arXiv:2606.28092v1 Announce Type: new Abstract: Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection.

researcharxiv-cs-cv
29 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

DGX agent

arXiv:2606.27677v1 Announce Type: cross Abstract: World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing meth

local-aiarxiv-cs-cv
29 Jun 2026
Safety

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

DGX agent

arXiv:2606.27964v1 Announce Type: new Abstract: Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video

safetyarxiv-cs-cv
29 Jun 2026
Safety

EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography

DGX agent

arXiv:2606.28164v1 Announce Type: new Abstract: Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreti

safetyarxiv-cs-cv
29 Jun 2026
Research

EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

DGX agent

arXiv:2606.28026v1 Announce Type: new Abstract: High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods f

researcharxiv-cs-cv
29 Jun 2026
Applications

Enhanced Neural Video Representation Compression across Extreme Complexity and Quality Scales

DGX agent

arXiv:2606.28163v1 Announce Type: cross Abstract: Implicit neural representations (INRs) have recently emerged as a promising approach to video compression, delivering competitive rate-distortion perf

applicationsarxiv-cs-cv
29 Jun 2026
Hardware

Enhancing Co-packaging Optics Enabled Silicon Photonics Security Assurance Hardware Fingerprinting

DGX agent

arXiv:2606.27612v1 Announce Type: cross Abstract: Silicon photonics enables integration of optical components using standard semiconductor processes, greatly improving data communication bandwidth and

hardwarearxiv-cs-cv
29 Jun 2026
Research

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

DGX agent

arXiv:2603.25168v2 Announce Type: replace Abstract: Previous works based on Segment Anything Model (SAM) have achieved promising performance in unified scene text detection and layout analysis. Howeve

researcharxiv-cs-cv
29 Jun 2026
Agents

Fine-Grained Behavior and Lane Constraints Guided Trajectory Prediction Method

DGX agent

arXiv:2503.21477v3 Announce Type: replace Abstract: Trajectory prediction, as a critical component of autonomous driving systems, has attracted the attention of many researchers. Existing prediction a

agentsarxiv-cs-cv
29 Jun 2026
Model Releases

Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

DGX agent

arXiv:2606.27484v1 Announce Type: new Abstract: Autism spectrum disorder (ASD) affects 1 in 31 US children, yet median age at diagnosis exceeds four years. Artificial intelligence pipelines that provi

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

GeoFace: Consistent Multi-View Face Generation with Geometry-Constrained Diffusion

DGX agent

arXiv:2606.27659v1 Announce Type: new Abstract: We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view

safetyarxiv-cs-cv
29 Jun 2026
Safety

Graph Unfolding and Sampling for Transitory Video Keyframe Selection via Gershgorin Disc Alignment

DGX agent

arXiv:2408.01859v2 Announce Type: replace Abstract: User-generated videos (UGVs) uploaded from mobile phones to social media sites like YouTube and TikTok are short and non-repetitive. We summarize a

safetyarxiv-cs-cv
29 Jun 2026
Agents

GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving

DGX agent

arXiv:2511.11266v4 Announce Type: replace Abstract: Vision-language models have recently emerged as promising planners for autonomous driving, where success hinges on topology-aware reasoning over spa

agentsarxiv-cs-cv
29 Jun 2026
Model Releases

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

DGX agent

arXiv:2606.27999v1 Announce Type: new Abstract: Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

HunyuanImage 3.0 Technical Report

DGX agent

arXiv:2509.23951v3 Announce Type: replace Abstract: We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with

safetyarxiv-cs-cv
29 Jun 2026
Local Ai

Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?

DGX agent

arXiv:2501.16947v2 Announce Type: replace Abstract: The advances in Vision-Language models (VLMs) offer exciting opportunities for robotic applications involving image geo-localization - the problem o

local-aiarxiv-cs-cv
29 Jun 2026
Applications

Instant Expressive Gaussian Head Avatars at Over 100 FPS

DGX agent

arXiv:2512.16893v2 Announce Type: replace Abstract: Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods ofte

applicationsarxiv-cs-cv
29 Jun 2026
Tutorials

Latent Visual Diffusion Reasoning with Monte Carlo Tree Search

DGX agent

arXiv:2606.27988v1 Announce Type: new Abstract: Analyzing fine-grained skill activities (e.g., sports, surgery) requires not only recognizing visual patterns but also performing step-by-step visual re

tutorialsarxiv-cs-cv
29 Jun 2026
Agents

Learning 1-Bit LiDAR-based Localization with Auxiliary Objective

DGX agent

arXiv:2606.27729v1 Announce Type: new Abstract: 6-DoF LiDAR-based localization is a fundamental capability for autonomous systems operating in large-scale outdoor environments. Many deep-learning-base

agentsarxiv-cs-cv
29 Jun 2026
Research

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

DGX agent

arXiv:2601.12066v4 Announce Type: replace Abstract: Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninfo

researcharxiv-cs-cv
29 Jun 2026
Safety

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

DGX agent

arXiv:2511.14271v2 Announce Type: replace Abstract: Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still

safetyarxiv-cs-cv
29 Jun 2026
Local Ai

Long-Term Prediction of Local and Global Human Motion with Occlusion Recovery

DGX agent

arXiv:2606.27900v1 Announce Type: new Abstract: Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of

local-aiarxiv-cs-cv
29 Jun 2026
Tutorials

MASS: Motion-Aligned Selective Scan for Refinement in Flow-Based Video Frame Interpolation

DGX agent

arXiv:2606.27718v1 Announce Type: new Abstract: Video frame interpolation (VFI) remains a challenging task, particularly when dealing with large, non-linear motions and complex occlusions. While flow-

tutorialsarxiv-cs-cv
29 Jun 2026
Tutorials

MeDUET: Disentangled Unified Pretraining for 3D Medical Image Synthesis and Analysis

DGX agent

arXiv:2602.17901v3 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) and diffusion models have respectively advanced representation learning and generative modeling for high-dimens

tutorialsarxiv-cs-cv
29 Jun 2026
Model Releases

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

DGX agent

arXiv:2606.27537v1 Announce Type: new Abstract: Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most ass

model-releasesarxiv-cs-cv
29 Jun 2026
Research

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations

DGX agent

arXiv:2606.27779v1 Announce Type: new Abstract: Generating lifelike facial animation for dyadic conversations requires reconciling high-level cognitive intent with precise low-level motor reflexes, ye

researcharxiv-cs-cv
29 Jun 2026
Safety

ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On

DGX agent

arXiv:2606.27773v1 Announce Type: new Abstract: Image-based virtual try-on has emerged as a compelling task in e-commerce and augmented reality, yet existing methods struggle to simultaneously preserv

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading

DGX agent

arXiv:2606.28144v1 Announce Type: new Abstract: Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarc

model-releasesarxiv-cs-cv
29 Jun 2026
Research

Multi-Modal Conditioned High-Resolution Transformer for Urban Electromagnetic Field Map Prediction Download PDF

DGX agent

arXiv:2606.27671v1 Announce Type: new Abstract: Predicting electromagnetic field (EMF) strength in urban environments is essential for cellular network planning but computationally expensive with phys

researcharxiv-cs-cv
29 Jun 2026
Research

MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis

DGX agent

arXiv:2410.02103v4 Announce Type: replace Abstract: Recent works in volume rendering, extit{e.g.} NeRF and 3D Gaussian Splatting (3DGS), significantly advance the rendering quality and efficiency with

researcharxiv-cs-cv
29 Jun 2026
Model Releases

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving

DGX agent

arXiv:2606.27660v1 Announce Type: new Abstract: Vision-Language Models (VLMs) improve generalization and interpretability in autonomous driving but suffer from efficiency issues due to long visual tok

model-releasesarxiv-cs-cv
29 Jun 2026
Local Ai

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

DGX agent

arXiv:2606.27771v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that a

local-aiarxiv-cs-cv
29 Jun 2026
Research

On the stability of scale-space metrics

DGX agent

arXiv:2606.27605v1 Announce Type: cross Abstract: We study the stability of a classical family of metrics defined over functions' Gaussian scale-space representations, focusing on the comparison of im

researcharxiv-cs-cv
29 Jun 2026
Model Releases

OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation

DGX agent

arXiv:2606.27880v1 Announce Type: new Abstract: Unified fashion generation integrates tasks like virtual try-on and garment reconstruction into a single model to reduce task-specific adaptation costs.

model-releasesarxiv-cs-cv
29 Jun 2026
Research

Panoramic Scene Analysis: A Survey from Distortion-Aware Engineering to Sphere-Native Foundation Modeling

DGX agent

arXiv:2606.27745v1 Announce Type: new Abstract: Panoramic images capture the complete visual sphere in a single frame, providing spatial context unattainable by conventional cameras. Yet this complete

researcharxiv-cs-cv
29 Jun 2026
Model Releases

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

DGX agent

arXiv:2606.28322v1 Announce Type: new Abstract: We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness

model-releasesarxiv-cs-cv
29 Jun 2026
Local Ai

Perceptual 3D Simulation With Physical World Modeling

DGX agent

arXiv:2606.27575v1 Announce Type: new Abstract: Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal s

local-aiarxiv-cs-cv
29 Jun 2026
Research

Permutation Learning with Only N Parameters: From SoftSort to Self-Organizing Gaussians

DGX agent

arXiv:2503.13051v3 Announce Type: replace-cross Abstract: Sorting and permutation learning are key concepts in optimization and machine learning, especially when organizing high-dimensional data into

researcharxiv-cs-cv
29 Jun 2026
Tutorials

PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding

DGX agent

arXiv:2511.20562v2 Announce Type: replace Abstract: While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability

tutorialsarxiv-cs-cv
29 Jun 2026
Research

PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion

DGX agent

arXiv:2606.27760v1 Announce Type: new Abstract: End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency seman

researcharxiv-cs-cv
29 Jun 2026
Model Releases

ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models

DGX agent

arXiv:2603.19466v2 Announce Type: replace Abstract: Effective collaboration begins with knowing when to ask for help. For example, when trying to identify an occluded object, a human would ask someone

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception

DGX agent

arXiv:2509.03704v2 Announce Type: replace Abstract: Cooperative perception through Vehicle-to-Everything (V2X) communication offers significant potential for enhancing vehicle perception by mitigating

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

Qwen-Image-2.0-RL Technical Report

DGX agent

arXiv:2606.27608v1 Announce Type: new Abstract: We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) t

model-releasesarxiv-cs-cv
29 Jun 2026
Local Ai

Radar Guided Camera Verification for Automatic Emergency Braking Rethinking Object Detection in Radar Camera Fusion

DGX agent

arXiv:2606.27556v1 Announce Type: new Abstract: Radar camera fusion is widely used in Automatic Emergency Braking AEB systems because radar provides reliable range and velocity measurements while came

local-aiarxiv-cs-cv
29 Jun 2026
Tutorials

RAE-NWM: Navigation World Model in Dense Visual Representation Space

DGX agent

arXiv:2603.09241v2 Announce Type: replace Abstract: Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulati

tutorialsarxiv-cs-cv
29 Jun 2026
Model Releases

RANSAC Scoring Done Right

DGX agent

arXiv:2606.27385v1 Announce Type: cross Abstract: The most widely used RANSAC variants score candidate models by counting inliers or summing per-point scores that saturate beyond a residual threshold.

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures

DGX agent

arXiv:2606.28060v1 Announce Type: new Abstract: Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require o

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

ReWorld: Learning Better Representations for World Action Models

DGX agent

arXiv:2606.27504v1 Announce Type: new Abstract: World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, e

safetyarxiv-cs-cv
29 Jun 2026
← Previous
1…8485868788…263
Next →