AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
19 May 2026

Today, we launched a brand-new intelligent Search box. Here's what that means: An upgrade to the Search experience with our most advanced Ge…

Model ReleasesDGX agent

Today, we launched a brand-new intelligent Search box. Here's what that means: An upgrade to the Search experience with our most advanced Gemini 3.5 models, bringing with them our latest agentic capab

Use your LM Studio models to code locally in @zeddotdev 🚀

Local AiDGX agent

Use your LM Studio models to code locally in @zeddotdev 🚀 Local model usage grew 3x in Zed's agent in the last 10 weeks. Cameron Mcloughlin on why he prefers local: 'I worry about over-reliance on pro

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

ApplicationsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.24497v3 Announce Type: replace Abstract: A long-standing challenge in AI is to develop agents capable of solving a wide range of physical tasks and generalizing to new, unseen tasks and env

When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State

Model ReleasesDGX agent

arXiv:2605.18580v1 Announce Type: new Abstract: Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral discipline. In hot

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Model ReleasesDGX agent

arXiv:2605.16402v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interf

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Model ReleasesDGX agent

arXiv:2605.17912v1 Announce Type: cross Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about envir

18 May 2026

Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning

Model ReleasesDGX agent

arXiv:2605.15967v1 Announce Type: new Abstract: We study event-graph substrates: a class of world models that represent agent state as an append-only log of typed RDF triples and answer counterfactual

DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding

ResearchDGX agent

arXiv:2605.15542v1 Announce Type: new Abstract: GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions.

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling

Model ReleasesDGX agent

arXiv:2605.15486v1 Announce Type: cross Abstract: This study introduces intelligent frameworks that use Large Language Models (LLMs) to improve task scheduling for construction robots. The LLM is fed

Improved Bounds for Reward-Agnostic and Reward-Free Exploration

Model ReleasesDGX agent

arXiv:2602.16363v2 Announce Type: replace Abstract: We study reward-free and reward-agnostic exploration in episodic finite-horizon Markov decision processes (MDPs), where an agent explores an unknown

Large Language Models as Optimization Controllers: Adaptive Continuation for SIMP Topology Optimization

Model ReleasesDGX agent

arXiv:2603.25099v2 Announce Type: replace-cross Abstract: We present a framework in which a large language model (LLM) acts as an online adaptive controller for SIMP topology optimization, replacing c

Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning

SafetyDGX agent

arXiv:2605.15975v1 Announce Type: new Abstract: We tackle the challenge of building embodied AI agents that can reliably solve long-horizon planning problems. Imitation learning from demonstrations ha

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.15342v1 Announce Type: new Abstract: Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation

ScreenSearch: Uncertainty-Aware OS Exploration

SafetyDGX agent

arXiv:2605.16024v1 Announce Type: new Abstract: Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plau

17 May 2026

And, yes, our experiments used a mix of GPT-4 & GPT-4o (publishing takes awhile). I think we would see much larger results with more recent …

Model ReleasesDGX agent

And, yes, our experiments used a mix of GPT-4 & GPT-4o (publishing takes awhile). I think we would see much larger results with more recent models, let alone recent agentic tools. 'The Cybernetic Team

15 May 2026

AI radio hosts demonstrate why AI can’t be trusted alone

Model ReleasesDGX agent

Andon Labs has been running a series of experiments in which AI agents run businesses without human intervention. Its latest is a quartet of radio stations run by some of the most popular AI models ou

Anthropic just went after the 44% of U.S. GDP that enterprise AI has mostly ignored. Claude for Small Business launched this week with 15 pr…

Model ReleasesDGX agent

Anthropic just went after the 44% of U.S. GDP that enterprise AI has mostly ignored. Claude for Small Business launched this week with 15 prebuilt agentic workflows and 15 skills connected directly in

BOOKMARKS: Efficient Active Storyline Memory for Role-playing

ResearchDGX agent

arXiv:2605.14169v1 Announce Type: new Abstract: Memory systems are critical for role-playing agents (RPAs) to maintain long-horizon consistency. However, existing RPA memory methods (e.g., profiling)

Dataiku launches governed AI workflow builder inside Snowflake

Model ReleasesDGX agent

Dataiku Inc. today announced a deeper integration with Snowflake Inc. aimed at simplifying the creation of enterprise AI agents while preserving governance and operational controls, a growing concern

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

ResearchDGX agent

arXiv:2605.14742v1 Announce Type: new Abstract: Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing m

Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI

HardwareDGX agent

arXiv:2605.14665v1 Announce Type: new Abstract: Legal reasoning is not semantic similarity search. A court judgment encodes constrained symbolic reasoning: precedent propagation, procedural state tran

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

Local AiDGX agent

arXiv:2605.14037v1 Announce Type: cross Abstract: Under modern test-time compute and agentic paradigms, language models process ever-longer sequences. Efficient text generation with transformer archit

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

SafetyDGX agent

arXiv:2605.15044v1 Announce Type: cross Abstract: As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-L

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

Model ReleasesDGX agent

arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work

VerbalValue: A Socially Intelligent Virtual Host for Sales-Driven Live Commerce

Model ReleasesDGX agent

arXiv:2605.14542v1 Announce Type: new Abstract: A skilled live-commerce host is not merely a narrator, but a sales agent who converts viewer curiosity into purchase intent through expert product knowl

14 May 2026

big launches today!

Model ReleasesDGX agent

big launches today! 🚀Launching: LangSmith Engine LangSmith Engine is an agent that sits on top of your traces It runs in the background and automatically identifies issues It then proactively suggests

Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling

SafetyDGX agent

arXiv:2605.13646v1 Announce Type: cross Abstract: End-to-end autonomous driving, which bypasses traditional modular pipelines by directly predicting future trajectories from sensor inputs, has recentl

Differentiable Evolutionary Reinforcement Learning

SafetyDGX agent

arXiv:2512.13399v2 Announce Type: replace Abstract: Crafting effective reward signals remains a central challenge in Reinforcement Learning (RL), especially for complex reasoning tasks. Existing autom

EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing

ApplicationsDGX agent

arXiv:2605.13041v1 Announce Type: new Abstract: With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applica

Embodied Neurocomputation: A Framework for Interfacing Biological Neural Cultures with Scaled Task-Driven Validation

Model ReleasesDGX agent

arXiv:2605.13315v1 Announce Type: cross Abstract: Biological neural networks (BNNs) have been established as a powerful and adaptive substrate that offer the potential for incredibly energy and data e

Grok Build

IndustryDGX agent

Grok Build Grok Build is a fully interactive CLI, which means you can actually use your mouse to click. No flickers. Especially useful as I find myself running 5+ agents at a time and jumping between

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

Model ReleasesDGX agent

arXiv:2603.03295v2 Announce Type: replace-cross Abstract: Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

SafetyDGX agent

arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo

Mitchell's post here reminded me of a similar conversation I had recently about how cheap it can be to port native mobile apps to React Nati…

ToolsDGX agent

Mitchell's post here reminded me of a similar conversation I had recently about how cheap it can be to port native mobile apps to React Native using coding agents... and then port them back again late

Probabilistic Prediction Markets with Intermittent Contributions

TutorialsDGX agent

arXiv:2510.13385v3 Announce Type: replace Abstract: Although both data availability and the demand for accurate forecasts are increasing, collaboration between stakeholders is often constrained by dat

the langsmith team cooked on this one @jakebroekhuizen this + claude code + langsmith cli has been 🔥

Model ReleasesDGX agent

the langsmith team cooked on this one @jakebroekhuizen this + claude code + langsmith cli has been 🔥 🚀Launching: LangSmith Engine LangSmith Engine is an agent that sits on top of your traces It runs i

ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues

Model ReleasesDGX agent

arXiv:2605.12521v1 Announce Type: cross Abstract: Multi-turn tool calling is essential for LLMs to function as autonomous agents, yet synthesizing the training data required for these capabilities rem

13 May 2026

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

Model ReleasesDGX agent

arXiv:2605.11367v1 Announce Type: new Abstract: Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame wo

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Model ReleasesDGX agent

arXiv:2605.12501v1 Announce Type: new Abstract: Computer-use agents (CUAs) automate on-screen work, as illustrated by GPT-5.4 and Claude. Yet their reliability on complex, low-frequency interactions i

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

Model ReleasesDGX agent

arXiv:2512.24985v4 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabili

Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback

Model ReleasesDGX agent

arXiv:2506.13163v3 Announce Type: replace Abstract: We study the Logistic Contextual Slate Bandit problem, where, at each round, an agent selects a slate of N items from an exponentially large set (of

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

Model ReleasesDGX agent

arXiv:2603.28561v2 Announce Type: replace Abstract: The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction

Nautilus: From One Prompt to Plug-and-Play Robot Learning

Model ReleasesDGX agent

arXiv:2605.11665v1 Announce Type: new Abstract: Robot learning research is fragmented across policy families, benchmark suites, and real robots; each implementation is entangled with the others in a c

NVIDIA, Ineffable Intelligence Team Up to Build the Future of Reinforcement Learning Infrastructure

HardwareDGX agent

Reinforcement-learning agents — AI systems that learn by trial and error — can convert computation into new knowledge. That’s the focus of a new engineering-level collaboration between NVIDIA and Inef

p2p ai inference mesh

Local AiDGX agent

A peer-to-peer AI inference mesh is a decentralized network architecture where agents connect directly to discover peers and communicate through distributed protocols to share inference workloads . Su

PriorZero: Bridging Language Priors and World Models for Decision Making

SafetyDGX agent

arXiv:2605.12289v1 Announce Type: new Abstract: Leveraging the rich world knowledge of Large Language Models (LLMs) to enhance Reinforcement Learning (RL) agents offers a promising path toward general

Shapley Value Approximation Based on k-Additive Games

ResearchDGX agent

arXiv:2502.04763v2 Announce Type: replace-cross Abstract: The Shapley value is the prevalent solution for fair division problems in which a payout is to be divided among multiple agents. By adopting a

The new era of SaMD: Why cloud infrastructure is the foundation for digital health in 2026

SafetyDGX agent

In the healthcare and life sciences industries, speed saves lives, but meeting regulatory requirements and other administrative burdens often pumps the brakes for manufacturers of software as a medica

Transferable Delay-Aware Reinforcement Learning via Implicit Causal Graph Modeling

SafetyDGX agent

arXiv:2605.12312v1 Announce Type: new Abstract: Random delays weaken the temporal correspondence between actions and subsequent state feedback, making it difficult for agents to identify the true prop

TriBand-BEV: Real-Time LiDAR-Only 3D Pedestrian Detection via Height-Aware BEV and High-Resolution Feature Fusion

HardwareDGX agent

arXiv:2605.12220v1 Announce Type: new Abstract: Safe autonomous agents and mobile robots need fast real time 3D perception, especially for vulnerable road users (VRUs) such as pedestrians. We introduc

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

Model ReleasesDGX agent

arXiv:2605.12237v1 Announce Type: new Abstract: Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scal

12 May 2026

Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10034v1 Announce Type: new Abstract: Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driv

Beyond Thinking: Imagining in 360^irc for Humanoid Visual Search

ResearchDGX agent

arXiv:2605.09146v1 Announce Type: new Abstract: Humanoid Visual Search (HVS) requires agents to actively explore immersive 360^irc environments. While prior methods treat this as a monolithic task rel

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

Model ReleasesDGX agent

arXiv:2605.08462v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI sys

Efficient Estimation of Kernel Surrogate Models for Task Attribution

TutorialsDGX agent

arXiv:2602.03783v2 Announce Type: replace-cross Abstract: Modern AI agents such as large language models are trained on diverse tasks -- translation, code generation, mathematical reasoning, and text

Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis

SafetyDGX agent

arXiv:2605.10910v1 Announce Type: cross Abstract: We consider the problem of synthesizing Clifford quantum circuits for devices with all-to-all qubit connectivity. We approach this task as a reinforce

Execution Envelopes: A Shared Admission Contract for Backend AI Execution Requests

SafetyDGX agent

arXiv:2605.08267v1 Announce Type: cross Abstract: Enterprise AI backends increasingly admit heterogeneous execution requests across model deployment, inference, evaluation, data movement, and agentic

Governed Metaprogramming for Intelligent Systems: Reclassifying Eval as a Governed Effect

SafetyDGX agent

arXiv:2605.05248v2 Announce Type: replace-cross Abstract: AI systems increasingly synthesize executable structure at runtime: LLMs generate programs, agents construct workflows,self-improving systems

Higher Resolution, Better Generalization: Unlocking Visual Scaling in Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.10546v1 Announce Type: new Abstract: Pixel-based deep reinforcement learning agents are typically trained on heavily downsampled visual observations, a convention inherited from early bench

← Previous
1…233234235236237…297
Next →