AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
Human
86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
61,904 results
30 Jun 2026

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control

SafetyDGX agent

arXiv:2606.28760v1 Announce Type: new Abstract: Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues t

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

Local AiDGX agent

arXiv:2606.30288v1 Announce Type: new Abstract: Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images

VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction

SafetyDGX agent

arXiv:2606.29548v1 Announce Type: cross Abstract: Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must decide whether

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

SafetyDGX agent

arXiv:2508.14483v4 Announce Type: replace Abstract: We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to c

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

SafetyDGX agent

arXiv:2606.30645v1 Announce Type: cross Abstract: Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapp

VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

SafetyDGX agent

arXiv:2510.00458v3 Announce Type: replace Abstract: Vision-language object detectors (VLODs) such as YOLO-World and Grounding DINO exhibit strong zero-shot generalization, but their performance degrad

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

Model ReleasesDGX agent

arXiv:2603.11734v2 Announce Type: replace Abstract: As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing spe

W4A4 Quantization for Inference on Wan2.2-I2V-A14B

ResearchDGX agent

arXiv:2606.29337v1 Announce Type: new Abstract: We summarize our submission to Sub-Challenge 1: W4A4 Quantization for Inference (HiF4 / MXFP4) of the ICME 2026 Low-Bit-width Large-Model Quantization C

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Local AiDGX agent

arXiv:2606.30045v1 Announce Type: new Abstract: Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transit

Warm-Starting Iterative Gaussian Processes for Faster Sequential Inference

ResearchDGX agent

arXiv:2511.16340v2 Announce Type: replace Abstract: Efficient Gaussian process (GP) inference is critical for sequential decision-making tasks such as active learning, online prediction, and Bayesian

WARP: Whole-Body Retargeting for Learning from Offline Human Demonstrations

ApplicationsDGX agent

arXiv:2606.29940v1 Announce Type: new Abstract: Direct transfer from human demonstration to learnable robot action is a crucial step towards scalable whole-body mobile manipulation. While human data s

Wasserstein Distributionally Robust Regret Optimization

ResearchDGX agent

arXiv:2504.10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

ResearchDGX agent

arXiv:2606.28857v1 Announce Type: cross Abstract: While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction o

Weak Dominant Balance for Robust Identification of Dynamically Consistent Fluid Flow Structure

ResearchDGX agent

arXiv:2606.29047v1 Announce Type: cross Abstract: Extracting interpretable, localized physical mechanisms from complex spatiotemporal data is a foundational challenge across physics, biology, and engi

Weighted Contrastive Learning for Anomaly-Aware Time-Series Forecasting

ResearchDGX agent

arXiv:2512.07569v2 Announce Type: replace-cross Abstract: Reliable forecasting of multivariate time series under anomalous conditions is crucial in applications such as ATM cash logistics, where sudde

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

AgentsDGX agent

arXiv:2603.02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? Clas

What Color is the Sky (for a non-human) ?

ResearchDGX agent

arXiv:2606.28912v1 Announce Type: new Abstract: The light of the daytime sky contains a mixture of many colors yet is perceived as blue by human observers. This is largely due to the particular respon

What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics

Model ReleasesDGX agent

arXiv:2606.29791v1 Announce Type: cross Abstract: Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particularly challen

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

TutorialsDGX agent

arXiv:2606.28615v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rati

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

Model ReleasesDGX agent

arXiv:2606.28438v1 Announce Type: cross Abstract: Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We st

When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation for Structured Generation

ResearchDGX agent

arXiv:2606.29054v1 Announce Type: new Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and s

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

SafetyDGX agent

arXiv:2606.30445v1 Announce Type: new Abstract: Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline sup

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2606.28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memor

When Does Sparsity Mitigate the Curse of Depth in LLMs

ResearchDGX agent

arXiv:2603.15389v2 Announce Type: replace Abstract: Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation t

When Does Synthetic CT Transfer? A Label-Free Donor/Host Diagnostic for Medical Vision-Language Model Routing on Real Lung CT

ResearchDGX agent

arXiv:2606.29232v1 Announce Type: new Abstract: A synthetic measurement of model competence is useful only if it survives the move to real data, yet the real labels that would verify it are exactly wh

When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate

ApplicationsDGX agent

arXiv:2512.03578v3 Announce Type: replace-cross Abstract: Time series extrinsic regression (TSER) refers to the task of predicting a continuous target variable from an input time series. It appears in

When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding

Local AiDGX agent

arXiv:2606.30265v1 Announce Type: cross Abstract: Speculative decoding accelerates language model inference by using a fast drafter to propose candidate tokens that are then verified by a larger targe

When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning

AgentsDGX agent

arXiv:2606.29354v1 Announce Type: new Abstract: Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language rationales that are

When May I Help You? On The Effect of Proactivity on Group Human-Robot Collaboration

ResearchDGX agent

arXiv:2606.28469v1 Announce Type: new Abstract: Robot initiative is a central challenge in multi-party human-robot collaboration. A robot that contributes without being addressed may provide timely su

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

Model ReleasesDGX agent

arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p

When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning

TutorialsDGX agent

arXiv:2601.07965v2 Announce Type: replace Abstract: When a model knows when it does not know, many possibilities emerge. The first question is how to enable a model to recognize that it does not know.

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

Model ReleasesDGX agent

arXiv:2606.28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sam

When Prices Double in a Week: Forecasting of Agricultural Volatility in Import-Isolated Markets

ResearchDGX agent

arXiv:2606.29248v1 Announce Type: new Abstract: Vegetable prices in Sri Lanka are highly volatile because the market is largely import-isolated, so supply disruptions quickly drive prices up. This stu

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

SafetyDGX agent

arXiv:2606.29115v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoi

When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

AgentsDGX agent

arXiv:2606.29251v1 Announce Type: new Abstract: Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large language models (LL

Where Do Humans Look When Demonstrating to Robots? Human Gaze Behavior in Pick-and-Place Tasks Across Demonstration Devices

ResearchDGX agent

arXiv:2506.05808v2 Announce Type: replace Abstract: Imitation learning for generalizable performance often requires a large volume of demonstration data, making the process significantly costly. One p

Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy

ResearchDGX agent

arXiv:2606.29489v1 Announce Type: new Abstract: When humans translate, not every word depends equally on the surrounding context. Some tokens, particularly function words like pronouns and auxiliaries

Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance Prediction

TutorialsDGX agent

arXiv:2606.28544v1 Announce Type: cross Abstract: Team roles offer an interpretable lens on collaboration, yet computational studies of roles often rely on domain-specific personas or data-driven clus

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

Model ReleasesDGX agent

arXiv:2606.30383v1 Announce Type: new Abstract: A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also

Why Do We Need Warm-up? A Theoretical Perspective

Model ReleasesDGX agent

arXiv:2510.03164v2 Announce Type: replace Abstract: Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

Model ReleasesDGX agent

arXiv:2606.29712v1 Announce Type: new Abstract: Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and

Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare

Model ReleasesDGX agent

arXiv:2606.28666v1 Announce Type: cross Abstract: Agent-based AI has enabled the automation of tasks by exposing application tools and resources to large language models (LLMs). However, to improve sc

Wireless Backdoor Attack and Defense for Semantic Communications over Multiple Access Channel

ResearchDGX agent

arXiv:2606.30595v1 Announce Type: cross Abstract: Semantic communication (SemCom) aims to preserve semantic meaning and task-oriented information beyond conventional message recovery over wireless cha

Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection

SafetyDGX agent

arXiv:2606.30587v1 Announce Type: cross Abstract: Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection. Recent work has shown that LLMs a

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

SafetyDGX agent

arXiv:2602.13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requi

X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving

SafetyDGX agent

arXiv:2606.28758v1 Announce Type: cross Abstract: Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relyi

X-Morph: Human Motion Priors for Scalable Robot Learning Across Morphologies

SafetyDGX agent

arXiv:2606.30290v1 Announce Type: new Abstract: Recent progress in humanoid behavior models has been driven in large part by abundant human motion data, but comparable motion data is scarce for non-hu

X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining

SafetyDGX agent

arXiv:2606.14752v2 Announce Type: replace-cross Abstract: Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing act

xGR: Efficient Generative Recommendation Serving at Scale

ApplicationsDGX agent

arXiv:2512.11529v3 Announce Type: replace Abstract: Recommendation system delivers substantial economic benefits by providing personalized predictions. Generative recommendation (GR) integrates LLMs t

XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2412.15529v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) synergizes the retrieval of pertinent data with the generative capabilities of Large Language Models (LLM

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

Model ReleasesDGX agent

arXiv:2506.00599v3 Announce Type: replace Abstract: While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical

You Only Touch Once: 6-DoF Object Pose Estimation from Single Tactile Contact

Model ReleasesDGX agent

arXiv:2606.28899v1 Announce Type: new Abstract: Accurate 6-DoF object pose estimation is fundamental to robotic manipulation, yet vision-based methods often fail under occlusion, poor lighting, and re

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

Local AiDGX agent

arXiv:2606.30248v1 Announce Type: new Abstract: Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with hu

Zero-Gated Language-conditioned Human Motion Prediction

Model ReleasesDGX agent

arXiv:2606.29208v1 Announce Type: new Abstract: Pose histories provide the core kinematic evidence for 3D human motion prediction, but they lack explicit high-level semantic guidance. This paper intro

Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture

SafetyDGX agent

arXiv:2606.28383v1 Announce Type: new Abstract: Identifying complex and safety-critical driving scenarios in large unlabelled datasets is an important but expensive problem. Existing approaches rely o

Zero-Shot Depth from Defocus

Model ReleasesDGX agent

arXiv:2603.26658v2 Announce Type: replace Abstract: Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain datas

29 Jun 2026

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

Model ReleasesDGX agent

arXiv:2606.27886v1 Announce Type: new Abstract: Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their

A Comprehensive Survey on World Models for Embodied AI

AgentsDGX agent

arXiv:2510.16732v3 Announce Type: replace Abstract: Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators th

A Gaussian Perspective for Distributional Discrepancy in Generative Diffusion Models

ResearchDGX agent

arXiv:2601.13602v3 Announce Type: replace-cross Abstract: This paper introduces an analytical approach to quantifying and optimizing the distributional discrepancy in generative diffusion models. For

A Multi-Attribute Latent Space for Visual Analysis of Watches

Model ReleasesDGX agent

arXiv:2606.27897v1 Announce Type: new Abstract: We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneous

← Previous
1…336337338339340…1032
Next →