AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
Human
87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,480 results
23 Jun 2026

Unified Motion-Action Modeling for Heterogeneous Robot Learning

ResearchDGX agent

arXiv:2606.16917v3 Announce Type: replace Abstract: We present Unified Motion-Action (UMA) Model, an approach that uses 3D object motion trajectories as a shared interface to bridge visuomotor control

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models

ResearchDGX agent

arXiv:2606.22794v1 Announce Type: new Abstract: Mainstream Fast-Slow dual system vision-language-action models decouple a high-frequency action expert from a low-frequency vision-language model for ef

UniRank: Unified Rank Allocation for Low-Rank LLM Compression

Model ReleasesDGX agent

arXiv:2606.21847v1 Announce Type: new Abstract: Low-rank decomposition serves as a promising compression paradigm for large language models, however, rank allocation remains challenging: manual rules

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

UniSLAD: A Unified Framework for Structural and Logical Industrial Visual Anomaly Detection

ResearchDGX agent

arXiv:2606.20768v1 Announce Type: new Abstract: Visual anomaly detection is a fundamental task in industrial automation. While existing approaches have achieved notable progress in identifying structu

UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion

SafetyDGX agent

arXiv:2606.20971v1 Announce Type: new Abstract: We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike pri

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Model ReleasesDGX agent

arXiv:2606.21661v1 Announce Type: new Abstract: Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist acros

Universal Encoders for Modular Relational Deep Learning

ResearchDGX agent

arXiv:2606.21434v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs for end-to-end representation learning. While RDL is evol

Universal priors: solving empirical Bayes via Bayesian inference and pretraining

ResearchDGX agent

arXiv:2602.15136v2 Announce Type: replace-cross Abstract: We theoretically justify the recent empirical finding of [Teh et al., 2025] that a transformer pretrained on synthetically generated data achi

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

ResearchDGX agent

arXiv:2606.23503v1 Announce Type: new Abstract: Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling

AgentsDGX agent

arXiv:2606.21501v1 Announce Type: new Abstract: Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference came

Unlimited OCR Works

Model ReleasesDGX agent

arXiv:2606.23050v1 Announce Type: new Abstract: Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a larg

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

ResearchDGX agent

arXiv:2606.23243v1 Announce Type: new Abstract: Clinical audio diagnosis in low-resource settings requires models that identify conditions from minimal examples without large annotated corpora. We pro

Unmasking LAION-5B: Age, Gender, Race, and Emotion Biases in Large-Scale Image Datasets

ResearchDGX agent

arXiv:2606.23204v1 Announce Type: new Abstract: Large-scale image-text datasets, such as LAION-5B, are foundational to modern AI systems, yet their vast scale and uncurated nature raise significant co

UNSEEN: Uncertainty-aware Navigation via Sparse Estimation in Unknown Environments

ApplicationsDGX agent

arXiv:2606.20755v1 Announce Type: new Abstract: Visual navigation in unknown environments remains a core challenge in mobile robotics, especially for resource-constrained platforms. Most existing appr

Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability

ResearchDGX agent

arXiv:2606.21385v1 Announce Type: new Abstract: This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors that influence

Unsupervised Domain Adaptation for Sim-to-Real Object Pose Estimation with Contrastive Alignment and Pseudo-Label Refinement

SafetyDGX agent

arXiv:2606.21287v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) enables robust transfer of knowledge from simulated to real environments while exploiting a subset of unlabeled tar

Unsupervised Susceptibility Distortion Correction of EPI without Calibration Scans via Image Translation-Based Registration

ResearchDGX agent

arXiv:2606.21588v1 Announce Type: cross Abstract: Functional magnetic resonance imaging (fMRI) utilizes echo-planar imaging (EPI) to capture blood-oxygen-level-dependent (BOLD) signals with high tempo

Urban Power Grid Topology and Hierarchy Identification from Open Data

ResearchDGX agent

arXiv:2606.21352v1 Announce Type: new Abstract: Understanding the complex topology and hierarchy of urban power grid is crucial for energy prognosis, power flow management, and system resilience analy

Using predictive multiplicity to measure individual performance within the AI Act

SafetyDGX agent

arXiv:2602.11944v2 Announce Type: replace Abstract: When building AI systems for decision support, one often encounters the phenomenon of predictive multiplicity: a single best model does not exist; i

Variance-Tilted Diffusion Models for Diverse Sampling

ResearchDGX agent

arXiv:2606.22239v1 Announce Type: cross Abstract: Diffusion models are typically sampled independently, even when the downstream objective is to obtain a diverse set of candidates. We introduce a vari

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

AgentsDGX agent

arXiv:2512.11061v2 Announce Type: replace Abstract: Generative video models, a leading approach to world modelling, face fundamental limitations. They often violate physical and logical rules, lack in

VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks

ApplicationsDGX agent

arXiv:2510.01022v3 Announce Type: replace Abstract: We introduce vector diffusion wavelets (VDWs), a novel family of wavelets inspired by the vector diffusion maps algorithm that was introduced to ana

Vec-QMDP: Vectorized POMDP Planning on CPUs for Real-Time Autonomous Driving

HardwareDGX agent

arXiv:2602.08334v2 Announce Type: replace Abstract: Planning under uncertainty for real-world robotics tasks, such as autonomous driving, requires reasoning in enormous high-dimensional belief spaces,

VegSim: A Geospatial World Model for Scenario-Conditioned Vegetation Simulation

ApplicationsDGX agent

arXiv:2606.21961v1 Announce Type: new Abstract: Vegetation monitoring under climate stress requires answering not only how it will evolve given the expected weather, but how it would respond to altern

Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation

ResearchDGX agent

arXiv:2606.22546v1 Announce Type: new Abstract: Modern Referring Image Segmentation (RIS) systems generate multiple candidate masks per expression but rely on a simple heuristic--typically the argmax

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Model ReleasesDGX agent

arXiv:2606.23610v1 Announce Type: new Abstract: Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existin

VeriBound: PAC-Bayesian Generalization Bounds for Process Reward Models Trained with Formal Verification Tools

ResearchDGX agent

arXiv:2606.20740v1 Announce Type: cross Abstract: Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, yet their training data acquisition remains a b

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Model ReleasesDGX agent

arXiv:2606.23543v1 Announce Type: cross Abstract: Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labe

Vesta: A Generalist Embodied Reasoning Model

Local AiDGX agent

arXiv:2606.20905v1 Announce Type: new Abstract: Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While spec

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

ResearchDGX agent

arXiv:2606.20711v1 Announce Type: new Abstract: UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. H

VideoAgent: All-in-One Framework for Video Understanding and Editing

Model ReleasesDGX agent

arXiv:2606.23327v1 Announce Type: new Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-speci

VideoLatent: Video-Language Learning via Latent Self-Forcing

SafetyDGX agent

arXiv:2606.22870v1 Announce Type: new Abstract: Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal lar

Vision-language models for chest radiography do not always need the image

Model ReleasesDGX agent

arXiv:2606.17710v2 Announce Type: replace Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That infe

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

ApplicationsDGX agent

arXiv:2606.22787v1 Announce Type: new Abstract: Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire di

VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2606.21386v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) achieve state-of-the-art performance on many robotic manipulation tasks, yet they can still behave unpredictably

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies

ApplicationsDGX agent

arXiv:2602.21445v2 Announce Type: replace Abstract: Action chunking has recently emerged as a standard practice in flow-based Vision-Language-Action (VLA) models. However, the effect and choice of the

VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes

Model ReleasesDGX agent

arXiv:2606.23062v1 Announce Type: cross Abstract: We introduce VolHuMe, a dataset of high-quality 4D human scans captured with a state-of-the-art volumetric studio using 64 RGB and 32 depth cameras. V

VQActFlow: Vector-Quantized Action Mode Steering for Multi-Task Robot Manipulation

SafetyDGX agent

arXiv:2606.21600v1 Announce Type: new Abstract: Multi-task robot manipulation policies are challenging to learn from demonstration because traditionally a single network must select among qualitativel

VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training

SafetyDGX agent

arXiv:2508.03058v2 Announce Type: replace Abstract: Reinforcement Learning (RL) in real-world environments often suffers from ambiguous or incomplete reward supervision, which undermines policy stabil

VT-DUDA: Visual Token Conditioning for Diffusion-guided Unsupervised Domain Adaptation

TutorialsDGX agent

arXiv:2606.21700v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) aims to learn a target-domain classifier from labeled source data and unlabeled target data under distribution shif

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

AgentsDGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

WarPGNN: A Parametric Thermal Warpage Analysis Framework with Physics-aware Graph Neural Network

ResearchDGX agent

arXiv:2603.18581v2 Announce Type: replace-cross Abstract: With the advent of system-in-package (SiP) chiplet-based design and heterogeneous 2.5D/3D integration, thermal-induced warpage has become a cr

WebCryptoAgent: Agentic Crypto Trading with Web Informatics

AgentsDGX agent

arXiv:2601.04687v2 Announce Type: replace Abstract: Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support shor

Weighted Score-Oriented Losses for Temporally Localized Event Prediction

Model ReleasesDGX agent

arXiv:2606.23145v1 Announce Type: new Abstract: Operational event-detection systems are rarely assessed by pointwise accuracy alone. In anomaly detection, changepoint detection, and warning systems, t

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

SafetyDGX agent

arXiv:2606.22136v1 Announce Type: new Abstract: Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and s

What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility

SafetyDGX agent

arXiv:2606.21126v1 Announce Type: new Abstract: Despite the success of deep learning, training deep networks in biologically plausible and hardware-efficient ways remains an open challenge. Feedback a

What Do Lorentz-Equivariant Jet Taggers Learn?

TutorialsDGX agent

arXiv:2606.21790v1 Announce Type: new Abstract: We study what Lorentz-equivariant jet taggers learn internally, using equivariance tests, linear probes and grade ablations across five models including

What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study

TutorialsDGX agent

arXiv:2606.22020v1 Announce Type: cross Abstract: Neural networks outperform classical GCC-PHAT for Time-Difference-of-Arrival (TDOA) estimation in noise and reverberation, yet their internal strategy

What Does a Chemical Language Model Know About Molecules?

TutorialsDGX agent

arXiv:2606.23443v1 Announce Type: new Abstract: Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here,

What if? Emulative Simulation with World Models for Situated Reasoning

SafetyDGX agent

arXiv:2603.06445v2 Announce Type: replace Abstract: Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints o

What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

Local AiDGX agent

arXiv:2606.20814v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evalu

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.22864v1 Announce Type: new Abstract: Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for

When Calibration Fails the Vulnerable Hospital: Federated Conformal Risk Control via Risk-Curve Shrinkage

ResearchDGX agent

arXiv:2606.20115v2 Announce Type: replace-cross Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-o

When Confidence Lacks Concepts: Interpretable OOD Detection via Representation Perturbations

SafetyDGX agent

arXiv:2606.16196v2 Announce Type: replace-cross Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributio

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR

ResearchDGX agent

arXiv:2606.22043v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is increasingly applied to large vision-language models (LVLMs), yet outcome-only optimization c

When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model

ResearchDGX agent

arXiv:2606.21641v1 Announce Type: new Abstract: Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that 'warm-start' search from prior knowledge, proposing s

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure

Model ReleasesDGX agent

arXiv:2606.23339v1 Announce Type: new Abstract: Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We pro

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

Model ReleasesDGX agent

arXiv:2606.20724v1 Announce Type: cross Abstract: Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, and terminate

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

Model ReleasesDGX agent

arXiv:2606.22079v1 Announce Type: cross Abstract: Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine,

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

Model ReleasesDGX agent

arXiv:2606.23057v1 Announce Type: cross Abstract: Large language models now mediate how buyers discover products and services, making the competitive structure of AI-generated recommendations a strate

← Previous
1…405406407408409…1042
Next →