AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
Human
86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
61,498 results
19 May 2026

Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research

ResearchDGX agent

arXiv:2605.18490v1 Announce Type: new Abstract: We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: a single-round Vector RAG system and an LLM-comp

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

ResearchDGX agent

arXiv:2602.22941v2 Announce Type: replace Abstract: Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard

Venom: A PyTorch Generative Modeling Toolkit

TutorialsDGX agent

arXiv:2605.17605v1 Announce Type: new Abstract: Modern generative modeling has grown into a broad collection of related but often separately implemented paradigms, including denoising diffusion models

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

HardwareDGX agent

arXiv:2605.17613v1 Announce Type: cross Abstract: The large size of the KV cache has become a major bottleneck for serving LLMs with increasing context lengths. In response, many KV cache compression

Verifier-Guided Code Translation via Meta-Step Decoding

Model ReleasesDGX agent

arXiv:2605.17626v1 Announce Type: new Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

Model ReleasesDGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

VeriHGN: Heterogeneous Graph-Based Congestion Prediction for Chip Layout Verification

ResearchDGX agent

arXiv:2603.11075v2 Announce Type: replace-cross Abstract: As Very Large Scale Integration (VLSI) designs continue to scale in size and complexity, layout verification has become a central challenge in

VGGT-CD: Training-Free Robust Registration for 3D Change Detection

Model ReleasesDGX agent

arXiv:2605.16859v1 Announce Type: cross Abstract: 3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods p

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

Model ReleasesDGX agent

arXiv:2605.16911v1 Announce Type: new Abstract: 3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subseq

Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance

SafetyDGX agent

arXiv:2605.16420v1 Announce Type: new Abstract: This paper addresses the problem of reconstructing missing or dropped frames in top-down drone video of autonomous surface vehicles performing structure

VideoNeuMat: Neural Material Extraction from Generative Video Models

ResearchDGX agent

arXiv:2602.07272v2 Announce Type: replace Abstract: Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currentl

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

ResearchDGX agent

arXiv:2605.16338v1 Announce Type: cross Abstract: The large-scale digitization of historical archives has created a paradox: 'dark data'-digital objects lacking metadata for retrieval. Manual archival

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

SafetyDGX agent

arXiv:2605.18192v1 Announce Type: new Abstract: Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existi

Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities

ResearchDGX agent

arXiv:2605.16880v1 Announce Type: new Abstract: Multimodal magnetic resonance imaging (MRI) is crucial for brain tumor segmentation, with many methods leveraging its four key modalities to capture com

Virtues of Ordered Chaos: Planning with Topple Actions in Tabletop Stack Rearrangement

ResearchDGX agent

arXiv:2605.17815v1 Announce Type: cross Abstract: Efficient object manipulation strategies have significant impact in automation applications. In this work, the stack rearrangement in tabletop setting

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

ApplicationsDGX agent

arXiv:2605.18547v1 Announce Type: new Abstract: Emotion Recognition in Conversation (ERC) is essential for effective human-machine interaction, aiming to identify speakers' emotional states in multi-t

Vision Foundation Models as Generalist Tokenizers for Image Generation

ResearchDGX agent

arXiv:2605.18390v1 Announce Type: new Abstract: In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Local AiDGX agent

arXiv:2605.18740v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evide

Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation

SafetyDGX agent

arXiv:2605.16393v1 Announce Type: cross Abstract: Semantic segmentation is essential for analysing anatomical features in biomedical research, yet a performance gap remains for Vision Transformers (Vi

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Model ReleasesDGX agent

arXiv:2602.04802v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing bench

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

ResearchDGX agent

arXiv:2605.17312v1 Announce Type: new Abstract: Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advan

VISTA: Variance-Gated Inter-Sequence Test-Time Adaptation for Multi-Sequence MRI Segmentation

Local AiDGX agent

arXiv:2605.17433v1 Announce Type: new Abstract: Deploying multi-sequence magnetic resonance imaging (MRI) segmentation models to new clinical environments is challenging due to variations in scanners

Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval

Model ReleasesDGX agent

arXiv:2605.16481v1 Announce Type: cross Abstract: Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps

Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting

SafetyDGX agent

arXiv:2605.17556v1 Announce Type: cross Abstract: Clay sculpting is a nuanced, artistic task involving dexterous manipulation with long-horizon planning to achieve high-level goals. As a robotics prob

Visual Search Patterns in 3D Pancreatic Imaging: An Eye Tracking Study

ResearchDGX agent

arXiv:2605.16408v1 Announce Type: new Abstract: Eye tracking has emerged as a powerful tool for examining visual perception and search strategies in various domains, including medicine. While it is re

Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC

ResearchDGX agent

arXiv:2605.17095v1 Announce Type: cross Abstract: Law enforcement agencies are accumulating vast amounts of body-worn camera (BWC) footage. However, this remains operationally opaque. That is, analyst

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

Model ReleasesDGX agent

arXiv:2605.18172v1 Announce Type: new Abstract: Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

SafetyDGX agent

arXiv:2603.18178v2 Announce Type: replace-cross Abstract: The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-co

Voice ''Cloning'' is Style Transfer

ResearchDGX agent

arXiv:2605.16578v1 Announce Type: cross Abstract: Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation

Voices in the Loop: Mapping Participatory AI

SafetyDGX agent

arXiv:2605.16827v1 Announce Type: new Abstract: Participatory approaches to artificial intelligence are increasingly documented across public, civic, and humanitarian settings, but evidence about how

VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment

SafetyDGX agent

arXiv:2605.16775v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has advanced medical image analysis be enabling learning form large unlabelled data. However, in brain magnetic resonan

VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

ResearchDGX agent

arXiv:2605.17102v1 Announce Type: cross Abstract: We present VoxScene, a novel anchor-conditioned voxel diffusion framework tailored for 3D scene synthesis. Current data-driven layout generation techn

VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption

ResearchDGX agent

arXiv:2605.17345v1 Announce Type: new Abstract: The release of public 3D medical image segmentation (MIS) datasets accelerates clinical research but simultaneously heightens risks of unauthorized AI m

VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

ApplicationsDGX agent

arXiv:2605.17584v1 Announce Type: new Abstract: Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often c

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

ResearchDGX agent

arXiv:2605.16364v1 Announce Type: cross Abstract: Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition erro

Wasserstein bounds for denoising diffusion probabilistic models via the Follmer process

ResearchDGX agent

arXiv:2605.18069v1 Announce Type: cross Abstract: This paper studies sampling error bounds for denoising diffusion probabilistic models (DDPMs) in the 2-Wasserstein distance. Our contributions are thr

Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.18313v1 Announce Type: cross Abstract: Small vision-language models (2-8B) are well-suited for clin- ical deployment due to privacy constraints, limited connectivity, and low-latency requir

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy

ResearchDGX agent

arXiv:2605.16796v1 Announce Type: cross Abstract: Watermarking combines an imperceptible change to an input image that will trigger a detector, to assert provenance and protect intellectual property.

Wavelet Flow Matching for Multi-Scale Physics Emulation

ResearchDGX agent

arXiv:2605.16573v1 Announce Type: cross Abstract: Accurate emulation of multi-scale physical systems governed by PDEs demands models that remain stable over long autoregressive rollouts while preservi

WavFlow: Audio Generation in Waveform Space

Model ReleasesDGX agent

arXiv:2605.18749v1 Announce Type: cross Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this wo

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

Model ReleasesDGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

SafetyDGX agent

arXiv:2605.17314v1 Announce Type: cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g.

Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation

ApplicationsDGX agent

arXiv:2605.18507v1 Announce Type: new Abstract: Due to the difficulty of obtaining ground-truth data for 4D radar scene flow estimation, previous methods typically rely on either self-supervised losse

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

Model ReleasesDGX agent

arXiv:2605.17637v1 Announce Type: new Abstract: Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate tr

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

Model ReleasesDGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

Weighted Flow Matching and Physics-Informed Nonlinear Filtering for Parameter Estimation in Digital Twins

Model ReleasesDGX agent

arXiv:2605.17146v1 Announce Type: cross Abstract: Digital twins (DTs) rely on continuous synchronization between physical systems and their virtual counterparts through online parameter estimation und

Weighted Reverse Convolution for Feature Upsampling

ResearchDGX agent

arXiv:2605.17472v1 Announce Type: new Abstract: Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting thei

Weisfeiler and Leman Follow the Arrow of Time: Expressive Power of Message Passing in Temporal Event Graphs

ResearchDGX agent

arXiv:2505.24438v3 Announce Type: replace Abstract: An important characteristic of temporal graphs is how the directed arrow of time influences their causal topology, i.e., which nodes can possibly in

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

Model ReleasesDGX agent

arXiv:2510.15221v2 Announce Type: replace Abstract: Affective computing has matured rapidly in laboratory settings, yet no prior dataset combines (i) months-to-years of duration, (ii) a naturalistic w

What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models

Model ReleasesDGX agent

arXiv:2605.18738v1 Announce Type: new Abstract: Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

ApplicationsDGX agent

arXiv:2512.24497v3 Announce Type: replace Abstract: A long-standing challenge in AI is to develop agents capable of solving a wide range of physical tasks and generalizing to new, unseen tasks and env

What is Holding Back Latent Visual Reasoning?

ResearchDGX agent

arXiv:2605.18445v1 Announce Type: cross Abstract: Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired b

What is the long-run distribution of stochastic gradient descent? A large deviations analysis

ResearchDGX agent

arXiv:2406.09241v3 Announce Type: replace-cross Abstract: In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we see

What Matters for Grocery Product Retrieval with Open Source Vision Language Models

ResearchDGX agent

arXiv:2605.18029v1 Announce Type: new Abstract: Multimodal product retrieval (MPR) underpins checkout-free retail and automated inventory systems, yet it demands fine-grained SKU discrimination that s

When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering

ResearchDGX agent

arXiv:2605.17658v1 Announce Type: new Abstract: Different age-related regulations have been proposed to protect minors from harmful content and interactions online. Automated age estimation is central

When Accuracy Is Not Enough: Uncertainty Collapse between Noisy Label Learning and Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2605.17795v1 Announce Type: cross Abstract: Learning with noisy labels (LNL) is typically benchmarked by closed-set classification accuracy, yet deployment often requires classifiers to reject o

When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

AgentsDGX agent

arXiv:2605.16312v1 Announce Type: cross Abstract: We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unl

← Previous
1…669670671672673…1025
Next →