AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models

SafetyDGX agent

arXiv:2604.10385v1 Announce Type: new Abstract: Generating complex multi-actor scenario videos remains difficult even for state-of-the-art neural generators, while evaluating them is hard due to the l

H-SPAM: Hierarchical Superpixel Anything Model

ResearchDGX agent

arXiv:2604.11218v1 Announce Type: new Abstract: Superpixels offer a compact image representation by grouping pixels into coherent regions. Recent methods have reached a plateau in terms of segmentatio

HDR 3D Gaussian Splatting via Luminance-Chromaticity Decomposition

Model ReleasesDGX agent

arXiv:2511.12895v2 Announce Type: replace Abstract: High Dynamic Range (HDR) 3D reconstruction is pivotal for professional content creation in filmmaking and virtual production. Existing methods typic


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

HDR Video Generation via Latent Alignment with Logarithmic Encoding

SafetyDGX agent

arXiv:2604.11788v1 Announce Type: new Abstract: High dynamic range (HDR) imagery offers a rich and faithful representation of scene radiance, but remains challenging for generative models due to its m

HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images

ResearchDGX agent

arXiv:2412.20704v2 Announce Type: replace Abstract: Dramatic advances in the quality of the latent diffusion models (LDMs) also led to the malicious use of AI-generated images. While current AI-genera

HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

Model ReleasesDGX agent

arXiv:2603.10128v2 Announce Type: replace Abstract: Lane detection is a crucial task in autonomous driving, as it helps ensure the safe operation of vehicles. However, existing datasets such as CULane

HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement

TutorialsDGX agent

arXiv:2604.10675v1 Announce Type: new Abstract: We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement know

Hide-and-Seek Attribution: Weakly Supervised Segmentation of Vertebral Metastases in CT

ResearchDGX agent

arXiv:2512.06849v2 Announce Type: replace Abstract: Accurate segmentation of vertebral metastasis in CT is clinically important yet difficult to scale, as voxel-level annotations are scarce and both l

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

TutorialsDGX agent

arXiv:2604.10836v1 Announce Type: new Abstract: Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and

HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models

ResearchDGX agent

arXiv:2604.10772v1 Announce Type: new Abstract: 3D layout generation and editing play a crucial role in Embodied AI and immersive VR interaction. However, manual creation requires tedious labor, while

How to Design a Compact High-Throughput Video Camera?

TutorialsDGX agent

arXiv:2604.10619v1 Announce Type: new Abstract: High throughput video acquisition is a challenging problem and has been drawing increasing attention. Existing high throughput imaging systems splice hu

How to Spin an Object: First, Get the Shape Right

TutorialsDGX agent

arXiv:2412.10273v3 Announce Type: replace Abstract: Image-to-3D models increasingly rely on hierarchical generation to disentangle geometry and texture. However, the design choices underlying these tw

HuiYanEarth-SAR: A Foundation Model for High-Fidelity and Low-Cost Global Remote Sensing Imagery Generation

ResearchDGX agent

arXiv:2604.11444v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) imagery generation is essential for deepening the study of scattering mechanisms, establishing trustworthy electromagneti

Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation

ResearchDGX agent

arXiv:2604.10837v1 Announce Type: new Abstract: Image-to-video (I2V) generation has the potential for societal harm because it enables the unauthorized animation of static images to create realistic d

Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps

Model ReleasesDGX agent

arXiv:2604.09688v1 Announce Type: new Abstract: Recent large-scale generative models enable high-quality 3D synthesis. However, the public accessibility of pre-trained weights introduces a critical vu

IMPLICITSTAINER: Resolution Agnostic Data-Efficient Virtual Staining Using Neural Implicit Functions

Local AiDGX agent

arXiv:2505.09831v2 Announce Type: replace-cross Abstract: Hematoxylin and eosin (H&E)-stained slides are central to cancer diagnosis and monitoring, visualizing tissue architecture and cellular morpho

Improving Deep Learning-Based Target Volume Auto-Delineation for Adaptive MR-Guided Radiotherapy in Head and Neck Cancer: Impact of a Volume-Aware Dice Loss

ResearchDGX agent

arXiv:2604.10130v1 Announce Type: new Abstract: Background: Manual delineation of target volumes in head and neck cancer (HNC) remains a significant bottleneck in radiotherapy planning, characterized

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization

AgentsDGX agent

arXiv:2604.11042v1 Announce Type: new Abstract: Fine-tuning object detection (OD) models on combined datasets assumes annotation compatibility, yet datasets often encode conflicting spatial definition

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

TutorialsDGX agent

arXiv:2507.00748v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image re

Inferring Dynamic Physical Properties from Video Foundation Models

ApplicationsDGX agent

arXiv:2510.02311v2 Announce Type: replace Abstract: We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal in

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

Model ReleasesDGX agent

arXiv:2604.07209v2 Announce Type: replace Abstract: Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generat

Intelligent bear deterrence system based on computer vision: Reducing human bear conflicts in remote areas

Local AiDGX agent

arXiv:2503.23178v2 Announce Type: replace Abstract: Conflicts between humans and bears on the Tibetan Plateau present substantial threats to local communities and hinder wildlife preservation initiati

Interactive Interface For Semantic Segmentation Dataset Synthesis

ApplicationsDGX agent

arXiv:2506.23470v2 Announce Type: replace Abstract: The rapid advancement of AI and computer vision has significantly increased the demand for high-quality annotated datasets, particularly for semanti

Intra-finger Variability of Diffusion-based Latent Fingerprint Generation

ResearchDGX agent

arXiv:2604.10040v1 Announce Type: new Abstract: The primary goal of this work is to systematically evaluate the intra-finger variability of synthetic fingerprints (particularly latent prints) generate

Investigating Bias and Fairness in Appearance-based Gaze Estimation

Model ReleasesDGX agent

arXiv:2604.10707v1 Announce Type: new Abstract: While appearance-based gaze estimation has achieved significant improvements in accuracy and domain adaptation, the fairness of these systems across dif

Iterative Inference-time Scaling with Adaptive Frequency Steering for Image Super-Resolution

ResearchDGX agent

arXiv:2512.23532v2 Announce Type: replace Abstract: Diffusion models have become a leading paradigm for image super-resolution (SR), but existing methods struggle to guarantee both the high-frequency

ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs

Model ReleasesDGX agent

arXiv:2604.03765v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing

K-STEMIT: Knowledge-Informed Spatio-Temporal Efficient Multi-Branch Graph Neural Network for Subsurface Stratigraphy Thickness Estimation from Radar Data

ResearchDGX agent

arXiv:2604.09922v1 Announce Type: cross Abstract: Subsurface stratigraphy contains important spatio-temporal information about accumulation, deformation, and layer formation in polar ice sheets. In pa

KiseKloset for Fashion Retrieval and Recommendation

ResearchDGX agent

arXiv:2506.23471v2 Announce Type: replace-cross Abstract: The global fashion e-commerce industry has become integral to people's daily lives, leveraging technological advancements to offer personalize

Language Prompt vs. Image Enhancement: Boosting Object Detection With CLIP in Hazy Environments

Model ReleasesDGX agent

arXiv:2604.10637v1 Announce Type: new Abstract: Object detection in hazy environments is challenging because degraded objects are nearly invisible and their semantics are weakened by environmental noi

LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment

Model ReleasesDGX agent

arXiv:2604.11689v1 Announce Type: new Abstract: While the shortage of explicit action data limits Vision-Language-Action (VLA) models, human action videos offer a scalable yet unlabeled data source. A

LDEPrompt: Layer-importance guided Dual Expandable Prompt Pool for Pre-trained Model-based Class-Incremental Learning

ResearchDGX agent

arXiv:2604.11091v1 Announce Type: new Abstract: Prompt-based class-incremental learning methods typically construct a prompt pool consisting of multiple trainable key-prompts and perform instance-leve

LEADER: Learning Reliable Local-to-Global Correspondences for LiDAR Relocalization

Model ReleasesDGX agent

arXiv:2604.11355v1 Announce Type: new Abstract: LiDAR relocalization has attracted increasing attention as it can deliver accurate 6-DoF pose estimation in complex 3D environments. Recent learning-bas

Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation

ApplicationsDGX agent

arXiv:2604.09955v1 Announce Type: new Abstract: Video Unsupervised Domain Adaptation (VUDA) poses a significant challenge in action recognition, requiring the adaptation of a model from a labeled sour

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

ResearchDGX agent

arXiv:2604.10573v1 Announce Type: new Abstract: Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied

Learning Long-term Motion Embeddings for Efficient Kinematics Generation

TutorialsDGX agent

arXiv:2604.11737v1 Announce Type: new Abstract: Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scen

Learning Robustness at Test-Time from a Non-Robust Teacher

Model ReleasesDGX agent

arXiv:2604.11590v1 Announce Type: new Abstract: Nowadays, pretrained models are increasingly used as general-purpose backbones and adapted at test-time to downstream environments where target data are

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2603.11346v2 Announce Type: replace Abstract: Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking w

Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video

ResearchDGX agent

arXiv:2511.18322v3 Announce Type: replace-cross Abstract: Learning soft continuum robot (SCR) dynamics from video offers flexibility but existing methods lack interpretability or rely on prior assumpt

LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning

Model ReleasesDGX agent

arXiv:2604.10780v1 Announce Type: new Abstract: Three-dimensional (3D) point cloud analysis has become central to applications ranging from autonomous driving and robotics to forestry and ecological m

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment

SafetyDGX agent

arXiv:2604.10677v1 Announce Type: cross Abstract: Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction dat

LiveGesture Streamable Co-Speech Gesture Generation Model

ResearchDGX agent

arXiv:2604.10927v1 Announce Type: new Abstract: We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

ResearchDGX agent

arXiv:2604.11789v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks req

LogitDynamics: Reliable ViT Error Detection from Layerwise Logit Trajectories

ResearchDGX agent

arXiv:2604.10643v1 Announce Type: new Abstract: Reliable confidence estimation is critical when deploying vision models. We study error prediction: determining whether an image classifier's output is

LoGo-MR: Screening Breast MRI for Cancer Risk Prediction by Efficient Omni-Slice Modeling

Model ReleasesDGX agent

arXiv:2604.11348v1 Announce Type: new Abstract: Efficient and explainable breast cancer (BC) risk prediction is critical for large-scale population-based screening. Breast MRI provides functional info

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation

Local AiDGX agent

arXiv:2604.10103v1 Announce Type: new Abstract: Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window att

LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval

Model ReleasesDGX agent

arXiv:2601.14706v3 Announce Type: replace Abstract: In this paper, we present LookBench (We use the term 'look' to reflect retrieval that mirrors how people shop -- finding the exact item, a close sub

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

Model ReleasesDGX agent

arXiv:2604.11792v1 Announce Type: new Abstract: Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multim

LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment: Methods and Results

Model ReleasesDGX agent

arXiv:2604.11207v1 Announce Type: new Abstract: This paper reviews the LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment. This challenge aims to raise a new direction, i.e., how

LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection

Model ReleasesDGX agent

arXiv:2604.10862v1 Announce Type: new Abstract: The rapid advancement of diffusion-based generative models has made face forgery detection a critical challenge in digital forensics. Current detection

LumiMotion: Improving Gaussian Relighting with Scene Dynamics

Model ReleasesDGX agent

arXiv:2604.10994v1 Announce Type: new Abstract: In 3D reconstruction, the problem of inverse rendering, namely recovering the illumination of the scene and the material properties, is fundamental. Exi

M^{2}SNet: Multi-scale in Multi-scale Subtraction Network for Medical Image Segmentation

ResearchDGX agent

arXiv:2303.10894v3 Announce Type: replace Abstract: Accurate medical image segmentation is critical for early medical diagnosis. Most existing methods are based on U-shape structure and use element-wi

MapATM: Enhancing HD Map Construction through Actor Trajectory Modeling

AgentsDGX agent

arXiv:2604.11081v1 Announce Type: new Abstract: High-definition (HD) mapping tasks, which perform lane detections and predictions, are extremely challenging due to non-ideal conditions such as view oc

Masked Training for Robust Arrhythmia Detection from Digitalized Multiple Layout ECG Images

ApplicationsDGX agent

arXiv:2508.09165v3 Announce Type: replace-cross Abstract: Background: Electrocardiograms are indispensable for diagnosing cardiovascular diseases, yet in many settings they exist only as paper printou

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

Model ReleasesDGX agent

arXiv:2511.18373v2 Announce Type: replace Abstract: Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial

MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration

Local AiDGX agent

arXiv:2604.11197v1 Announce Type: new Abstract: Contrastive Language-Image Pre-training (CLIP) has demonstrated outstanding performance in global image understanding and zero-shot transfer through lar

MedVeriSeg: Teaching MLLM-Based Medical Segmentation Models to Verify Query Validity Without Extra Training

Model ReleasesDGX agent

arXiv:2604.10242v1 Announce Type: new Abstract: Despite recent advances in MLLM-based medical image segmentation, existing LISA-like methods cannot reliably reject false queries and often produce hall

MetroGS: Efficient and Stable Reconstruction of Geometrically Accurate High-Fidelity Large-Scale Scenes

TutorialsDGX agent

arXiv:2511.19172v4 Announce Type: replace Abstract: Recently, 3D Gaussian Splatting and its derivatives have achieved significant breakthroughs in large-scale scene reconstruction. However, how to eff

Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models

ResearchDGX agent

arXiv:2604.10095v1 Announce Type: new Abstract: With the emergence of 3D foundation models, there is growing interest in fine-tuning them for downstream tasks, where LoRA is the dominant fine-tuning p

Mirai: Autoregressive Visual Generation Needs Foresight

Model ReleasesDGX agent

arXiv:2601.14671v2 Announce Type: replace Abstract: Autoregressive (AR) visual generators model images as sequences of discrete tokens and are trained with a next-token likelihood objective. This stri

← Previous
1…196197198199200…207
Next →