AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries88,316
  • Agents7,549
  • Applications5,409
  • Concepts5
  • Hardware1,833
  • Industry6,164
  • Local Ai4,927
  • Model Releases23,845
  • Research20,123
  • Safety13,368
  • Syntheses17
  • Tools1,675
  • Tutorials3,401

Source
HumanDGX agent

88,316Total entries
1Added by human
88,315Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
88,316 results
19 May 2026

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

ResearchDGX agent

arXiv:2602.22941v2 Announce Type: replace Abstract: Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard

Venom: A PyTorch Generative Modeling Toolkit

TutorialsDGX agent

arXiv:2605.17605v1 Announce Type: new Abstract: Modern generative modeling has grown into a broad collection of related but often separately implemented paradigms, including denoising diffusion models

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

HardwareDGX agent

arXiv:2605.17613v1 Announce Type: cross Abstract: The large size of the KV cache has become a major bottleneck for serving LLMs with increasing context lengths. In response, many KV cache compression

Content type
AllBlogX PostPaperYouTubeRedditGitHub

Verifier-Guided Code Translation via Meta-Step Decoding

Model ReleasesDGX agent

arXiv:2605.17626v1 Announce Type: new Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

Model ReleasesDGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

VeriHGN: Heterogeneous Graph-Based Congestion Prediction for Chip Layout Verification

ResearchDGX agent

arXiv:2603.11075v2 Announce Type: replace-cross Abstract: As Very Large Scale Integration (VLSI) designs continue to scale in size and complexity, layout verification has become a central challenge in

VGGT-CD: Training-Free Robust Registration for 3D Change Detection

Model ReleasesDGX agent

arXiv:2605.16859v1 Announce Type: cross Abstract: 3D change detection from multi-view images is essential for urban monitoring, disaster assessment, and autonomous driving. However, existing methods p

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

Model ReleasesDGX agent

arXiv:2605.16911v1 Announce Type: new Abstract: 3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subseq

Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance

SafetyDGX agent

arXiv:2605.16420v1 Announce Type: new Abstract: This paper addresses the problem of reconstructing missing or dropped frames in top-down drone video of autonomous surface vehicles performing structure

[video] why we need a new continuity layer for long-running agents (claude did this video! all except the voice which was @elevenlabs)

Model ReleasesDGX agent

This video discusses the architectural need for a continuity layer in long-running AI agents, explaining how agents require persistent memory and state management mechanisms to maintain coherence acro

VideoNeuMat: Neural Material Extraction from Generative Video Models

ResearchDGX agent

arXiv:2602.07272v2 Announce Type: replace Abstract: Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currentl

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

ResearchDGX agent

arXiv:2605.16338v1 Announce Type: cross Abstract: The large-scale digitization of historical archives has created a paradox: 'dark data'-digital objects lacking metadata for retrieval. Manual archival

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

SafetyDGX agent

arXiv:2605.18192v1 Announce Type: new Abstract: Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existi

Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities

ResearchDGX agent

arXiv:2605.16880v1 Announce Type: new Abstract: Multimodal magnetic resonance imaging (MRI) is crucial for brain tumor segmentation, with many methods leveraging its four key modalities to capture com

Virtues of Ordered Chaos: Planning with Topple Actions in Tabletop Stack Rearrangement

ResearchDGX agent

arXiv:2605.17815v1 Announce Type: cross Abstract: Efficient object manipulation strategies have significant impact in automation applications. In this work, the stack rearrangement in tabletop setting

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

ApplicationsDGX agent

arXiv:2605.18547v1 Announce Type: new Abstract: Emotion Recognition in Conversation (ERC) is essential for effective human-machine interaction, aiming to identify speakers' emotional states in multi-t

Vision Foundation Models as Generalist Tokenizers for Image Generation

ResearchDGX agent

arXiv:2605.18390v1 Announce Type: new Abstract: In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

Local AiDGX agent

arXiv:2605.18740v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evide

Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation

SafetyDGX agent

arXiv:2605.16393v1 Announce Type: cross Abstract: Semantic segmentation is essential for analysing anatomical features in biomedical research, yet a performance gap remains for Vision Transformers (Vi

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Model ReleasesDGX agent

arXiv:2602.04802v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing bench

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

ResearchDGX agent

arXiv:2605.17312v1 Announce Type: new Abstract: Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advan

VISTA: Variance-Gated Inter-Sequence Test-Time Adaptation for Multi-Sequence MRI Segmentation

Local AiDGX agent

arXiv:2605.17433v1 Announce Type: new Abstract: Deploying multi-sequence magnetic resonance imaging (MRI) segmentation models to new clinical environments is challenging due to variations in scanners

Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval

Model ReleasesDGX agent

arXiv:2605.16481v1 Announce Type: cross Abstract: Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps

Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting

SafetyDGX agent

arXiv:2605.17556v1 Announce Type: cross Abstract: Clay sculpting is a nuanced, artistic task involving dexterous manipulation with long-horizon planning to achieve high-level goals. As a robotics prob

Visual Search Patterns in 3D Pancreatic Imaging: An Eye Tracking Study

ResearchDGX agent

arXiv:2605.16408v1 Announce Type: new Abstract: Eye tracking has emerged as a powerful tool for examining visual perception and search strategies in various domains, including medicine. While it is re

Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC

ResearchDGX agent

arXiv:2605.17095v1 Announce Type: cross Abstract: Law enforcement agencies are accumulating vast amounts of body-worn camera (BWC) footage. However, this remains operationally opaque. That is, analyst

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

Model ReleasesDGX agent

arXiv:2605.18172v1 Announce Type: new Abstract: Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

SafetyDGX agent

arXiv:2603.18178v2 Announce Type: replace-cross Abstract: The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-co

Voice ''Cloning'' is Style Transfer

ResearchDGX agent

arXiv:2605.16578v1 Announce Type: cross Abstract: Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation

Voice 'cloning' is style transfer. Across three widely used systems — ElevenLabs V3, Coqui-XTTS, Chatterbox — clones don't just copy speaker…

ToolsDGX agent

Voice 'cloning' is style transfer. Across three widely used systems — ElevenLabs V3, Coqui-XTTS, Chatterbox — clones don't just copy speakers, they reshape them to be warmer, more authoritative, more

Voices in the Loop: Mapping Participatory AI

SafetyDGX agent

arXiv:2605.16827v1 Announce Type: new Abstract: Participatory approaches to artificial intelligence are increasingly documented across public, civic, and humanitarian settings, but evidence about how

Voker raises $2.2M to help teams understand how AI agents perform in the wild

AgentsDGX agent

Voker, an agent analytics platform for artificial intelligence product teams, today announced it has raised 2.2 million in pre-seed funding from Y Combinator and FundersClub. As more companies push AI

VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment

SafetyDGX agent

arXiv:2605.16775v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has advanced medical image analysis be enabling learning form large unlabelled data. However, in brain magnetic resonan

volunteer here https://ai.engineer/cfp !

ToolsDGX agent

This post links to a Call for Proposals (CFP) for volunteering opportunities at AI.Engineer, likely inviting community members to submit talk proposals, workshop ideas, or volunteer roles for an AI en

VoxScene: Anchor-Conditioned Voxel Diffusion for Indoor Scene Arrangement

ResearchDGX agent

arXiv:2605.17102v1 Announce Type: cross Abstract: We present VoxScene, a novel anchor-conditioned voxel diffusion framework tailored for 3D scene synthesis. Current data-driven layout generation techn

VoxShield: Protecting 3D Medical Datasets from Unauthorized Training via Frequency-Aware Inter-Slice Disruption

ResearchDGX agent

arXiv:2605.17345v1 Announce Type: new Abstract: The release of public 3D medical image segmentation (MIS) datasets accelerates clinical research but simultaneously heightens risks of unauthorized AI m

Vultr Announces Milan, Italy, as 33rd Cloud Data Center Region

HardwareDGX agent

Vultr expanded its global cloud infrastructure by opening a new data center region in Milan, Italy, marking the company's 33rd cloud data center location worldwide. This expansion provides European cu

VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos

ApplicationsDGX agent

arXiv:2605.17584v1 Announce Type: new Abstract: Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often c

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

ResearchDGX agent

arXiv:2605.16364v1 Announce Type: cross Abstract: Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition erro

Wasserstein bounds for denoising diffusion probabilistic models via the Follmer process

ResearchDGX agent

arXiv:2605.18069v1 Announce Type: cross Abstract: This paper studies sampling error bounds for denoising diffusion probabilistic models (DDPMs) in the 2-Wasserstein distance. Our contributions are thr

Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.18313v1 Announce Type: cross Abstract: Small vision-language models (2-8B) are well-suited for clin- ical deployment due to privacy constraints, limited connectivity, and low-latency requir

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy

ResearchDGX agent

arXiv:2605.16796v1 Announce Type: cross Abstract: Watermarking combines an imperceptible change to an input image that will trigger a detector, to assert provenance and protect intellectual property.

Wavelet Flow Matching for Multi-Scale Physics Emulation

ResearchDGX agent

arXiv:2605.16573v1 Announce Type: cross Abstract: Accurate emulation of multi-scale physical systems governed by PDEs demands models that remain stable over long autoregressive rollouts while preservi

WavFlow: Audio Generation in Waveform Space

Model ReleasesDGX agent

arXiv:2605.18749v1 Announce Type: cross Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this wo

We are aware of a @Railway outage impacting Nous Portal users and have contacted their team for more information on service restoration ETA.

ResearchDGX agent

Nous Research reported an outage affecting Railway that impacted access to Nous Portal, and the company stated they had contacted Railway's team to obtain information about when service would be resto

We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process the whole human geno…

HardwareDGX agent

We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process the whole human genome on a single GPU in <2 days. Here are the tricks we used:

We just added Grok's new imagine model in Paper so you can explore images even faster. Here's what we found: - Super fast generations for th…

IndustryDGX agent

We just added Grok's new imagine model in Paper so you can explore images even faster. Here's what we found: - Super fast generations for the quality - Saves details when editing - Perfect for 30 rapi

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

Model ReleasesDGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

`WE` understand why that gen-ai will always be derivative, because it learned from the most common things people do. Sadly, the slightly cliched writing tropes that most people do, are now being condemned *because* LLM's like GPT learned them.

IndustryDGX agent

This Reddit post discusses how large language models like GPT inevitably produce derivative content because they train on the most common patterns in human-generated text, including prevalent writing

We were able to sit down with the @GoogleDeepmind team behind the new Gemini Omni Flash model to hear all of their behind-the-scenes stories…

Model ReleasesDGX agent

We were able to sit down with the @GoogleDeepmind team behind the new Gemini Omni Flash model to hear all of their behind-the-scenes stories, memorable moments, and many, many (occasionally embarrassi

we will offer this until we sell out of our current allocation for this program. (we will make sure to leave enough capacity for ChatGPT, Co…

IndustryDGX agent

we will offer this until we sell out of our current allocation for this program. (we will make sure to leave enough capacity for ChatGPT, Codex, etc.) we plan to offer it again in the future; our inte

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

SafetyDGX agent

arXiv:2605.17314v1 Announce Type: cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g.

Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation

ApplicationsDGX agent

arXiv:2605.18507v1 Announce Type: new Abstract: Due to the difficulty of obtaining ground-truth data for 4D radar scene flow estimation, previous methods typically rely on either self-supervised losse

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

Model ReleasesDGX agent

arXiv:2605.17637v1 Announce Type: new Abstract: Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate tr

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

Model ReleasesDGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

Weighted Flow Matching and Physics-Informed Nonlinear Filtering for Parameter Estimation in Digital Twins

Model ReleasesDGX agent

arXiv:2605.17146v1 Announce Type: cross Abstract: Digital twins (DTs) rely on continuous synchronization between physical systems and their virtual counterparts through online parameter estimation und

Weighted Reverse Convolution for Feature Upsampling

ResearchDGX agent

arXiv:2605.17472v1 Announce Type: new Abstract: Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting thei

← Previous
1…931932933934935…1472
Next →