AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
21 May 2026

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

SafetyDGX agent

arXiv:2605.20780v1 Announce Type: cross Abstract: Physics-informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and pron

LER-YOLO: Reliability-Aware Expert Routing for Misaligned RGB-Infrared UAV Detection

Model ReleasesDGX agent

arXiv:2605.20667v1 Announce Type: new Abstract: Detecting small unmanned aerial vehicles from RGB-infrared remote-sensing pairs remains challenging due to tiny target scale, cluttered backgrounds, and

Let EEG Models Learn EEG

TutorialsDGX agent

arXiv:2605.21280v1 Announce Type: new Abstract: High-fidelity EEG generation is critical for alleviating data scarcity and addressing privacy constraints in large-scale neural modeling. Despite recent


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching

SafetyDGX agent

arXiv:2510.09060v2 Announce Type: replace-cross Abstract: Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. E

Leveraging Vision-Language Models to Detect Attention in Educational Videos

Model ReleasesDGX agent

arXiv:2605.20211v1 Announce Type: new Abstract: Educational videos are a cornerstone of remote and blended learning. However, learners' fluctuating attention remains a significant barrier to effective

Lighting-aware Unified Model for Instance Segmentation

ApplicationsDGX agent

arXiv:2605.20436v1 Announce Type: new Abstract: Foundation models like the Segment Anything Model (SAM) demonstrate impressive zero-shot generalization but frequently degrade under diverse real-world

Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

SafetyDGX agent

arXiv:2605.21123v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are co

LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation

AgentsDGX agent

arXiv:2605.21007v1 Announce Type: new Abstract: Road segmentation is a fundamental perception task for autonomous driving and intelligent robotic systems, requiring both high accuracy and real-time in

Local-sensitive connectivity filter (ls-cf): A post-processing unsupervised improvement of the frangi, hessian and vesselness filters for multimodal vessel segmentation

ResearchDGX agent

arXiv:2605.21251v1 Announce Type: cross Abstract: A retinal vessel analysis is a procedure that can be used as an assessment of risks to the eye. This work proposes an unsupervised multimodal approach

Lowering the Barrier to IREX Participation: Open-Source Algorithms, Toolkit, and Benchmarking for Iris Recognition

ResearchDGX agent

arXiv:2605.20735v1 Announce Type: new Abstract: This paper proposes two new open-source iris recognition algorithms, providing both Python and IREX-compliant C++ implementations to be submitted to the

Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video

ResearchDGX agent

arXiv:2605.20889v1 Announce Type: new Abstract: Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within t

MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space

ApplicationsDGX agent

arXiv:2605.20549v1 Announce Type: new Abstract: Modern vision models achieve strong performance on standard benchmarks, yet their aggregate accuracy reveals little about which scene properties drive t

Mechanistic Interpretability for Learning Assurance of a Vision-Based Landing System

SafetyDGX agent

arXiv:2605.20607v1 Announce Type: cross Abstract: EASA's learning-assurance guidance requires data-driven aviation systems to build and monitor their own situation representation, yet for neural netwo

MedCRP-CL: Continual Medical Image Segmentation via Bayesian Nonparametric Semantic Modality Discovery

Model ReleasesDGX agent

arXiv:2605.20297v1 Announce Type: new Abstract: Medical image segmentation faces a fundamental challenge in continual learning: data arrives sequentially from heterogeneous sources, yet effective cont

MeshTailor: Cutting Seams via Generative Mesh Traversal

ResearchDGX agent

arXiv:2603.27309v2 Announce Type: replace-cross Abstract: We present MeshTailor, the first mesh-native generative framework for synthesizing edge-aligned seams on 3D surfaces. Unlike prior optimizatio

Mind Your Margin and Boundary: Are Your Distilled Datasets Truly Robust?

ResearchDGX agent

arXiv:2605.20606v1 Announce Type: new Abstract: Dataset distillation (DD) compresses a large training set into a small synthetic set for efficient training, but most DD methods optimize only clean acc

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

Model ReleasesDGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

Multi-needle Localization for Pelvic Seed Implant Brachytherapy based on Tip-handle Detection and Matching

Local AiDGX agent

arXiv:2509.17931v2 Announce Type: replace Abstract: Accurate multi-needle localization in intraoperative CT images is crucial for optimizing seed placement in pelvic seed implant brachytherapy. Howeve

Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning

ApplicationsDGX agent

arXiv:2507.09180v4 Announce Type: replace Abstract: Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision tr

Multimodal LLMs under Pairwise Modalities

SafetyDGX agent

arXiv:2605.21059v1 Announce Type: new Abstract: Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data,

Multimodal Optimal Transport for Training-free Temporal Segmentation in Surgical Robotics

Model ReleasesDGX agent

arXiv:2602.24138v2 Announce Type: replace Abstract: Automated recognition of surgical phases and steps is a fundamental capability for intraoperative decision support, workflow automation, and skill a

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

SafetyDGX agent

arXiv:2605.20302v1 Announce Type: cross Abstract: Supervised classification has a theoretical optimum, Neural Collapse (NC), yet neither of its two dominant paradigms reaches it in practice. Cross ent

OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation

SafetyDGX agent

arXiv:2605.21343v1 Announce Type: new Abstract: Recent layout-to-image models have achieved remarkable progress in spatial controllability. However, they still struggle with inter-object occlusion. Wh

OlmoEarth v1.1: A more efficient family of OlmoEarth models

HardwareDGX agent

arXiv:2605.20804v1 Announce Type: new Abstract: We present a set of improvements to the OlmoEarth family. These improvements allow us to cut compute costs during training (1.7 imes reduction in GPU ho

One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration

Local AiDGX agent

arXiv:2605.21484v1 Announce Type: new Abstract: Discrete diffusion models excel at visual synthesis but rely on slow, iterative decoding. Existing single-step distillation methods attempt to bypass th

Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation

Local AiDGX agent

arXiv:2509.09946v2 Announce Type: replace Abstract: Multi-Target Multi-Camera Tracking (MTMC) is an essential computer vision task for automating large-scale surveillance. With camera calibration and

Oracle Supervision Transfers for Hyperparameter Prediction in Model-Based Image Denoising

ResearchDGX agent

arXiv:2605.20479v1 Announce Type: new Abstract: Hyperparameter prediction is a critical practical bottleneck for model-based image denoisers, ranging from classical TV/TGV variational solvers to moder

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

ResearchDGX agent

arXiv:2605.21417v1 Announce Type: new Abstract: Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a sing

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

Local AiDGX agent

arXiv:2605.20818v1 Announce Type: new Abstract: In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 20

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation

Local AiDGX agent

arXiv:2605.20941v1 Announce Type: new Abstract: We present PaintCopilot, a co-creative neural painting assistant that models painting as an open-ended autoregressive artistic behavior conditioned on e

Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing

Model ReleasesDGX agent

arXiv:2602.06862v2 Announce Type: replace Abstract: Adapting pre-trained vision models using parameter-efficient fine-tuning (PEFT) remains challenging, as it aims to achieve performance comparable to

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

AgentsDGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

SafetyDGX agent

arXiv:2605.20640v1 Announce Type: new Abstract: Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthe

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

ResearchDGX agent

arXiv:2602.04876v2 Announce Type: replace Abstract: We introduce PerpetualWonder, a hybrid generative simulator that enables long-horizon, action-conditioned 4D scene generation from a single image. C

PGC: Peak-Guided Calibration for Generalizable AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2605.21207v1 Announce Type: new Abstract: The rapid evolution of generative AI, from GANs to modern diffusion models, has resulted in increasingly subtle discriminative clues. These fine-grained

Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated Image Segmentation Architectures

ResearchDGX agent

arXiv:2605.20459v1 Announce Type: new Abstract: In recent years, there has been a notable increase in the level of attention that is given to algorithms based on deep learning in the context of medica

Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry

ResearchDGX agent

arXiv:2605.20496v1 Announce Type: cross Abstract: The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

SafetyDGX agent

arXiv:2605.21414v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-languag

PREF: Phasorial Embedding Fields for Compact Neural Representations

ResearchDGX agent

arXiv:2205.13524v4 Announce Type: replace Abstract: We present an efficient frequency-based neural representation termed PREF: a shallow MLP augmented with a phasor volume that covers significant bord

Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

Model ReleasesDGX agent

arXiv:2605.20961v1 Announce Type: new Abstract: Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions whi

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

SafetyDGX agent

arXiv:2510.21583v2 Announce Type: replace Abstract: Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated st

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection

AgentsDGX agent

arXiv:2605.20867v1 Announce Type: cross Abstract: Multimodal sarcasm detection requires reasoning over cross-modal incongruities between literal expression and intended meaning, yet the specific analy

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

ResearchDGX agent

arXiv:2602.06886v3 Announce Type: replace Abstract: Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information fl

ProtoPathway: Biologically Structured Prototype-Pathway Fusion for Multimodal Cancer Survival Prediction

ResearchDGX agent

arXiv:2605.21454v1 Announce Type: new Abstract: We introduce ProtoPathway, an interpretable-by-design multimodal framework for cancer survival prediction that unifies whole slide imaging and transcrip

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

ResearchDGX agent

arXiv:2605.21072v1 Announce Type: new Abstract: Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time inte

Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution

Model ReleasesDGX agent

arXiv:2602.01273v4 Announce Type: replace Abstract: Recently, Diffusion Transformers (DiTs) have emerged in Real-World Image Super-Resolution (Real-ISR) to generate high-quality textures, yet their he

Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments

Model ReleasesDGX agent

arXiv:2512.09447v2 Announce Type: replace-cross Abstract: Structurally repetitive environments produce visually plausible but aliased LiDAR loop candidates that can destabilize pose-graph optimization

QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs

Model ReleasesDGX agent

arXiv:2605.20584v1 Announce Type: new Abstract: Mobile app marketplaces require developers to disclose standardized content rating descriptors (CRDs) to inform users about potentially sensitive or res

R2AoP: Reliable and Robust Angle of Progression Estimation from Intrapartum Ultrasound

Local AiDGX agent

arXiv:2605.21099v1 Announce Type: new Abstract: Accurate estimation of the Angle of Progression (AoP) from intrapartum transperineal ultrasound is critical for objective assessment of labor progressio

RadProPoser: Probabilistic Radar Tensor Human Pose Estimation That Knows Its Limits

Model ReleasesDGX agent

arXiv:2508.03578v2 Announce Type: replace Abstract: Radar-based human pose estimation enables privacy-preserving motion tracking for ambient intelligence, yet the noisy nature of radar sensing makes u

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

Model ReleasesDGX agent

arXiv:2605.21195v1 Announce Type: new Abstract: Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the pol

RCGDet3D: Rethinking 4D Radar-Camera Fusion-based 3D Object Detection with Enhanced Radar Feature Encoding

Model ReleasesDGX agent

arXiv:2605.21112v1 Announce Type: new Abstract: 4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detect

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens

TutorialsDGX agent

arXiv:2605.21300v1 Announce Type: new Abstract: Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that on

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

SafetyDGX agent

arXiv:2605.20277v1 Announce Type: new Abstract: Medical vision-language models (VLMs) have rapidly advanced as general-purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT

RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses

ResearchDGX agent

arXiv:2605.20823v1 Announce Type: new Abstract: Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central

ReMATF: Recurrent Motion-Adaptive Multi-scale Turbulence Mitigation for Dynamic Scenes

ResearchDGX agent

arXiv:2605.21440v1 Announce Type: new Abstract: Atmospheric turbulence severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing signi

RePCM: Region-Specific and Phenotype-Adaptive Bi-Ventricular Cardiac Motion Synthesis

Local AiDGX agent

arXiv:2605.21237v1 Announce Type: new Abstract: Cardiac motion over a cardiac cycle is crucial for quantifying regional function and is strongly affected by cardiovascular diseases. Since temporally d

ResNet-50 with Class Reweighting and Anatomy-Guided Temporal Decoding for Gastrointestinal Video Analysis

ResearchDGX agent

arXiv:2603.17784v2 Announce Type: replace Abstract: We developed a multi-label gastrointestinal video analysis pipeline based on a ResNet-50 frame classifier followed by anatomy-guided temporal event

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors

Model ReleasesDGX agent

arXiv:2605.20737v1 Announce Type: new Abstract: Existing approaches for unsupervised 3D point cloud segmentation predominantly rely on a purely visual similarity-based learning-by-clustering paradigm,

Rethinking Cross-Layer Information Routing in Diffusion Transformers

SafetyDGX agent

arXiv:2605.20708v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization,

← Previous
1…121122123124125…211
Next →