AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
Human
88,343Total entries
1Added by human
88,342Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
9 Jun 2026

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

ResearchDGX agent

arXiv:2606.09125v1 Announce Type: cross Abstract: Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. Howe

UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough

ResearchDGX agent

arXiv:2603.29875v3 Announce Type: replace-cross Abstract: One of the key problems in Retrieval-augmented generation (RAG) systems is that chunk-based retrieval pipelines represent the source chunks as

VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands

SafetyDGX agent

arXiv:2606.09286v1 Announce Type: new Abstract: Humanoid robots hold immense potential for real-world assistance, yet agile interaction with objects in unstructured environments demands tightly couple

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Variational Proximal Policy Optimization

Local AiDGX agent

arXiv:2606.08032v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback via Proximal Policy Optimization often suffers from policy mode collapse, brittle exploration loops, and di

Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance

ResearchDGX agent

arXiv:2602.05774v4 Announce Type: replace-cross Abstract: Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single g

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

Model ReleasesDGX agent

arXiv:2606.07992v1 Announce Type: new Abstract: As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-hand

Vector Space of Cycles

ResearchDGX agent

arXiv:2606.08202v1 Announce Type: cross Abstract: Most statistical and machine learning methods for directed interactions focus on pairwise effects among variables. Even existing cyclic models represe

Vessel Traffic Flow Prediction on Sparse Data via Spatio-Temporal Graph Neural Networks with a Learnable Tweedie Head

SafetyDGX agent

arXiv:2606.07694v1 Announce Type: new Abstract: Accurate vessel traffic flow prediction is crucial for smart port operations and navigational safety. However, maritime traffic flow data are often high

vesselFM-CT: Segmenting All Blood Vessels in CT Images for System-Level Cardiovascular Analysis

ResearchDGX agent

arXiv:2606.09400v1 Announce Type: new Abstract: The vascular network in the human body is characterized by blood vessels exhibiting drastic structural variations in radius, length, topological propert

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

SafetyDGX agent

arXiv:2606.08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, a

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

ResearchDGX agent

arXiv:2510.03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores c

VGP-Nav: Metric-Aware Visual Geometric Perception for Robot Navigation

AgentsDGX agent

arXiv:2606.09268v1 Announce Type: new Abstract: Reliable robotic navigation necessitates the seamless integration of accurate global localization and dense, metric-consistent obstacle perception. A co

Video Understanding by Design: How Datasets Shape Video Models

SafetyDGX agent

arXiv:2509.09151v2 Announce Type: replace-cross Abstract: Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While exi

Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

SafetyDGX agent

arXiv:2606.08828v1 Announce Type: new Abstract: Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains cha

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Model ReleasesDGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

ViMax: Agentic Video Generation

AgentsDGX agent

arXiv:2606.07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing meth

Virtual-point-based Solutions to Handle Generalized Absolute Pose Problem

AgentsDGX agent

arXiv:2606.09294v1 Announce Type: new Abstract: Multi-camera systems are increasingly adopted in robotics and autonomous navigation for their wide field of view, flexibility, and fault tolerance. Neve

Vision-Based Early Fault Diagnosis and Self-Recovery for Strawberry Harvesting Robots

ResearchDGX agent

arXiv:2601.02085v3 Announce Type: replace-cross Abstract: Strawberry-harvesting robots faced challenges such as poor visual perception, gripper misalignment, empty grasp/misgrasp, and slippage, which

Vision-Guided Dual-Arm Humanoid Robotic Disassembly of End-of-Life 18650 Lithium-ion Battery Packs

TutorialsDGX agent

arXiv:2606.08152v1 Announce Type: new Abstract: The growing volume of retired lithium-ion battery packs from electric vehicles and portable electronics calls for automated disassembly that is safe, fl

Vision-Language Asymmetry in Bistable Image Captioning

SafetyDGX agent

arXiv:2606.08031v1 Announce Type: new Abstract: Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment t

Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating

TutorialsDGX agent

arXiv:2606.09167v1 Announce Type: new Abstract: Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for

Vision Language Model Helps Private Information De-Identification in Vision Data

TutorialsDGX agent

arXiv:2606.09132v1 Announce Type: new Abstract: Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

SafetyDGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

SafetyDGX agent

arXiv:2606.09290v1 Announce Type: new Abstract: Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perce

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

ApplicationsDGX agent

arXiv:2606.09670v1 Announce Type: cross Abstract: Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these

Visual Template Inference for Data Extraction from Documents

Model ReleasesDGX agent

arXiv:2501.06659v2 Announce Type: replace-cross Abstract: Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, t

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

Model ReleasesDGX agent

arXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking extern

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Model ReleasesDGX agent

arXiv:2606.07723v1 Announce Type: new Abstract: Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning

Voting Protocols as Coordination Mechanisms for Role-Constrained Multi-Agent Tutoring Systems

AgentsDGX agent

arXiv:2606.08030v1 Announce Type: cross Abstract: Agentic tutoring systems introduce a coordination challenge: multiple agents may propose different but reasonable interventions, yet only one response

WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

SafetyDGX agent

arXiv:2606.08670v1 Announce Type: new Abstract: Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full-resolution 3D brain MRI synthesis can support data

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

TutorialsDGX agent

arXiv:2602.08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow hi

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Model ReleasesDGX agent

arXiv:2606.09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and ext

Web Agents Should Use Typed Actions Instead of Click-Based Browsing

SafetyDGX agent

arXiv:2602.17245v2 Announce Type: replace Abstract: This position paper argues that building a reliable agentic Web requires shifting from low-level interaction primitives to typed actions supported b

Wedge Sampling: Efficient Tensor Completion with Nearly-Linear Sample Complexity

ResearchDGX agent

arXiv:2602.05869v2 Announce Type: replace-cross Abstract: We introduce Wedge Sampling, a new non-adaptive sampling scheme for low-rank tensor completion. We study recovery of an order-k low-rank tenso

Weighted universal approximation of differentiable maps on infinite-dimensional manifolds

ResearchDGX agent

arXiv:2606.09820v1 Announce Type: cross Abstract: We generalize the universal approximation theorem for functional input neural networks (FNN) to differentiable maps by including the approximation of

What Makes a Desired Graph for Relational Deep Learning?

SafetyDGX agent

arXiv:2606.08491v1 Announce Type: new Abstract: Relational deep learning (RDL) converts relational databases (RDBs) into heterogeneous graphs, but graphs derived directly from database schemas are oft

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

ResearchDGX agent

arXiv:2606.07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-

What neurosurgeons need to see: synthetic intra-operative MRI from ultrasound for brain-shift compensation in brain tumour surgery

ResearchDGX agent

arXiv:2606.07658v1 Announce Type: new Abstract: Maximal safe resection is the primary objective in glioma surgery. Neuronavigation guidance is progressively degraded by brain shift after dural opening

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

ResearchDGX agent

arXiv:2606.09700v1 Announce Type: cross Abstract: Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content. However, these systems pr

What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing

ResearchDGX agent

arXiv:2606.08056v1 Announce Type: cross Abstract: Sign language models are predominantly trained with gloss-sequence or text supervision, thereby under-modeling non-lexical and productive construction

When Are Neural Interaction Discoveries Real? Identifiability, Recoverability, and a Pre-Fit Diagnostic

ResearchDGX agent

arXiv:2606.08390v1 Announce Type: new Abstract: When a neural time-series model reports that one variable modulates another's effect on a target, is the discovered interaction a property of the data o

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Model ReleasesDGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

When Do Local Score Models Extrapolate Across Size? A Diagnostic Theory and Benchmark

Model ReleasesDGX agent

arXiv:2606.09705v1 Announce Type: new Abstract: Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones. While translation-invar

When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference

ResearchDGX agent

arXiv:2606.08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference. We show that piping the signals every sampl

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding

ResearchDGX agent

arXiv:2606.08239v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made substantial advancements in video understanding, yet the reliability of their responses remains under

When Should an AI Scientist Stop? Verifiable Experiment Steering and Refusal for Autonomous Discovery

AgentsDGX agent

arXiv:2606.07576v1 Announce Type: new Abstract: We present CARTOGRAPH, a verification layer for AI scientists that couples unresolved-subspace experiment steering (select), explicit ambiguity closure

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA

AgentsDGX agent

arXiv:2606.08542v1 Announce Type: cross Abstract: Exploratory manipulation often turns an apparent failed attempt into the key evidence for what to do next. For example, a robot pulls a locked cabinet

When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

Local AiDGX agent

arXiv:2606.08918v1 Announce Type: new Abstract: Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existing methods often mislocalize images by matc

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.09644v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models

Model ReleasesDGX agent

arXiv:2606.07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the mod

Where the Score Lives: A Wavelet View of Diffusion

ResearchDGX agent

arXiv:2606.08309v1 Announce Type: cross Abstract: Score-based generative models have had remarkable success over the last decade in generating a diverse set of visually plausible images. A variety of

WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing

SafetyDGX agent

arXiv:2606.07710v1 Announce Type: cross Abstract: The autoregressive nature of large language models (LLMs) remains a significant bottleneck for inference, particularly in complex agentic workloads. W

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

Model ReleasesDGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning

ResearchDGX agent

arXiv:2606.07720v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks. The CoCoNuT (Chain of Contin

Wispy to Voluminous: Prior-free Multi-view Capture of Strand-level Facial Hair

ApplicationsDGX agent

arXiv:2606.08041v1 Announce Type: cross Abstract: Facial hair is a defining trait of personal identity, yet remains a critical bottleneck for digital avatars. Recent volumetric methods achieve photore

X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting

SafetyDGX agent

arXiv:2606.07934v1 Announce Type: new Abstract: Whole-body teleoperation is essential for scalable robot data collection in loco-manipulation tasks, yet existing approaches relying on exoskeleton suit

X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication

ApplicationsDGX agent

arXiv:2606.08437v1 Announce Type: cross Abstract: Palmprint modality offers a privacy-preserving biometric solution, yet its deployment is hindered by the domain gap between controlled enrollment and

← Previous
1…457458459460461…1049
Next →