AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
6 Jun 2026

When Should Memory Stay Silent: Measuring Memory-Use Boundaries in Memory-Augmented Conversational Agents

Model ReleasesDGX agent

arXiv:2606.06055v1 Announce Type: new Abstract: Long-term memory enables language model agents to support personalized interactions, but it remains unclear when available memories warrant integration

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

Model ReleasesDGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

5 Jun 2026

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

Model ReleasesDGX agent

arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

SafetyDGX agent

arXiv:2606.05894v1 Announce Type: new Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses ans

Storyboarding in Grok @imagine was pretty fun. I really like how the agents help visualize some of the important historical events. I was tr…

IndustryDGX agent

Storyboarding in Grok @imagine was pretty fun. I really like how the agents help visualize some of the important historical events. I was trying to iterate through the clothing and voices from the anc

Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning

SafetyDGX agent

arXiv:2601.21700v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly support culturally sensitive decision making, yet often exhibit misalignment due to skewed pretraining dat

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories

SafetyDGX agent

arXiv:2606.05414v1 Announce Type: new Abstract: Early failure alerting requires deciding, while a dialog or agent trajectory is still unfolding, whether to flag it as likely to fail. This is challengi

Working with agents should feel like working with a colleague. You should be able “speak to” them not just with text chats, but by gesturing…

IndustryDGX agent

Working with agents should feel like working with a colleague. You should be able “speak to” them not just with text chats, but by gesturing at a screen together, talking live, etc. With Design Mode,

4 Jun 2026

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filte…

Model ReleasesDGX agent

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% - surp

Adaptive Minds: Empowering Agents with LoRA-as-Tools

Model ReleasesDGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

An interview with Naomi Gleit, Meta's head of product who joined the company 20 years ago, on Zuckerberg's 'unfair' reputation, AI agents' capabilities, more (Zoe Kleinman/BBC)

IndustryDGX agent

Zoe Kleinman / BBC: An interview with Naomi Gleit, Meta's head of product who joined the company 20 years ago, on Zuckerberg's “unfair” reputation, AI agents' capabilities, more — When Naomi Gleit joi

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

Model ReleasesDGX agent

arXiv:2606.04141v1 Announce Type: cross Abstract: LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt inj

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100m…

Model ReleasesDGX agent

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward

OpenAI banned my account the day after I paid. 3 years of work, 30-40 Codex agents, all my client income — locked. No reason given. I'm the sole provider for my family. Has this happened to anyone?

IndustryDGX agent

This post describes a user's account suspension by OpenAI shortly after making a payment, resulting in loss of access to multiple Codex agents and client-dependent income, with no explanation provided

Shared my first trace from @NanoClaw_AI to @huggingface yesterday. Very cool! By default, all agents should store their traces on HF (in pri…

IndustryDGX agent

Shared my first trace from @NanoClaw_AI to @huggingface yesterday. Very cool! By default, all agents should store their traces on HF (in private) so that you can keep a history of them, analyze them,.

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and…

HardwareDGX agent

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and TensorRT engines plus event-driven streaming I/O for low-la

3 Jun 2026

ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents

Local AiDGX agent

arXiv:2606.03239v1 Announce Type: new Abstract: LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on o

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

Model ReleasesDGX agent

arXiv:2507.21638v2 Announce Type: replace Abstract: The development of reinforcement learning (RL) algorithms has been largely driven by ambitious challenge tasks and benchmarks. Games have dominated

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reache…

Model ReleasesDGX agent

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reached 18/100 all-pass versus 14/100 for Opus alone, at 39% of th

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:1…

Model ReleasesDGX agent

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:12b-mlx Claude Code: ollama launch claude --model gemma4:12b-

Margin Play: A Multi-Agent System For Public Policy Analysis In The Brazilian Equatorial Margin

SafetyDGX agent

arXiv:2606.02614v1 Announce Type: cross Abstract: The Brazilian Equatorial Margin (BEM) is Brazil's next offshore oil frontier, with operations expected to begin in 2026 in the Foz do Amazonas basin.

MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency

HardwareDGX agent

arXiv:2606.03014v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs. Efficiently exec

NVIDIA Enables the Next Era Of Physical AI Research With Agent Skills For Autonomous Vehicles, Robotics And Vision AI

HardwareDGX agent

At CVPR, NVIDIA is unveiling new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems. The core challenge in physica

OpenAI ran a hiring challenge, but the top candidate was one they couldn’t hire: our autonomous research agent, Aiden. In Parameter Golf, Ai…

Model ReleasesDGX agent

OpenAI ran a hiring challenge, but the top candidate was one they couldn’t hire: our autonomous research agent, Aiden. In Parameter Golf, Aiden ran for 22 days, and out-outperformed all 1,016 other re

Pinecone Nexus Now Integrates with Microsoft OneLake, Bringing AI Agents Directly to Enterprise Data

ApplicationsDGX agent

Pinecone Nexus has integrated with Microsoft OneLake, enabling AI agents to access and work directly with enterprise data stored in OneLake without requiring data movement or copying. This integration

RGMem: Renormalization Group-inspired Memory Evolution for Language Agents

ResearchDGX agent

arXiv:2510.16392v3 Announce Type: replace Abstract: Personalized and continuous interactions are critical for LLM-based conversational agents, yet finite context windows and static parametric memory h

the best agents aren't just built with the best models: they're built with harnesses purpose-built for the task at hand here's a guide on ho…

TutorialsDGX agent

the best agents aren't just built with the best models: they're built with harnesses purpose-built for the task at hand here's a guide on how to build a harness that's really good at feeding the model

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

Model ReleasesDGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at th…

ToolsDGX agent

Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at the value Uber thinks these tools are providing https://simonw

2 Jun 2026

Build on-device personal AI agents on Windows PCs with new tools from NVIDIA and Microsoft, including secure sandboxing, faster local infere…

Local AiDGX agent

Build on-device personal AI agents on Windows PCs with new tools from NVIDIA and Microsoft, including secure sandboxing, faster local inference, multi-GPU support, and RTX acceleration for Windows AI

Cisco bets on AI agents to redefine workplace collaboration

IndustryDGX agent

Cisco Systems Inc. today unveiled a set of collaboration and customer experience products that transform its Webex videoconferencing platform into what executives describe as an intelligent operating

Conductor's parallel coding agents were local-only, but now they can run remotely on Vercel. 'Our customers can't tell the difference becaus…

ToolsDGX agent

Conductor's parallel coding agents were local-only, but now they can run remotely on Vercel. 'Our customers can't tell the difference because Vercel's Sandboxes are so fast.' https://vercel.com/blog/h

Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based Agents

Model ReleasesDGX agent

arXiv:2606.02204v1 Announce Type: new Abstract: Large language model agents achieve strong performance on text-based benchmarks but incur prohibitive inference costs, motivating the use of compact neu

Deploy Agentic-Ready AI at the Edge with Memory Efficiency in NVIDIA JetPack 7.2

HardwareDGX agent

NVIDIA JetPack 7.2 enables deployment of AI agents to edge devices with optimized memory and performance for real-world applications. The release directly supports one-command deployment of NVIDIA Nem

Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents

ResearchDGX agent

arXiv:2606.00476v1 Announce Type: new Abstract: Do LLM agents act on the reasoning they state? This question of process fidelity is central to using LLMs in social simulation, yet it is hard to measur

Holo3.1: Fast & Local Computer Use Agents

ToolsDGX agent

Holo3.1 is a computer use agent developed by Hugging Face that enables fast, local execution of tasks on computing systems without requiring cloud infrastructure. The model is designed to interpret an

Leyline: KV Cache Directives for Agentic Inference

SafetyDGX agent

arXiv:2606.01065v1 Announce Type: cross Abstract: Modern KV cache management assumes the chatbot workload: prompts arrive once and the cache grows append-only, so prefix caching and forward-only evict

LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]

ResearchDGX agent

Research demonstrates that LLM-based agents can generate functionally correct patches that pass all tests while still containing security vulnerabilities, challenging the assumption that test-passing

LLM Consortium for Software Design Refinement: A Controlled Experiment on Multi-Agent Collaboration Topologies

Model ReleasesDGX agent

arXiv:2606.01490v1 Announce Type: cross Abstract: We present a controlled experiment evaluating 12 multi-agent LLM collaboration topologies for software architecture design. Using a 2imes2imes2 factor

MobEvolve: An Agentic Self-Evolving Heuristic System for Interpretable Human Mobility Generation

SafetyDGX agent

arXiv:2606.01640v1 Announce Type: new Abstract: Human mobility generation aims to synthesize realistic trip chains for target populations based on individual features. Existing paradigms, including de

NVIDIA Partners With Microsoft on Unified Stack for Agentic AI Deployment, From Windows Devices to Cloud to Local

Local AiDGX agent

The agentic AI moment has arrived, but delivering on its promise requires more than good models. It also takes fast hardware, secure runtimes, a responsive data layer and models tuned for long-running

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving

HardwareDGX agent

arXiv:2606.01839v1 Announce Type: cross Abstract: LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when t

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

SafetyDGX agent

arXiv:2606.00135v1 Announce Type: cross Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This pa

RescueBench: Can Embodied Agents Save Lives in the Wild ?

Model ReleasesDGX agent

arXiv:2606.01848v1 Announce Type: new Abstract: Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and

Today we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model runnin…

Local AiDGX agent

Today we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This kee

Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool 'to responsibly encourage agentic AI adoption' (Natalie Lung/Bloomberg)

Model ReleasesDGX agent

Natalie Lung / Bloomberg: Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool “to responsibly encourage agentic AI adoption” — Uber Technologies Inc. has set

Using Parallel Agents to Move Faster in Replit https://x.com/i/broadcasts/1NxarrEMVOnKj

ToolsDGX agent

This broadcast discusses techniques for improving performance and speed in Replit by utilizing parallel agents or concurrent processing methods. The content likely covers how developers can leverage p

1 Jun 2026

A new beginning of PC starts with @NVIDIARTXSpark, supercharging what's possible in Hermes Agent.

HardwareDGX agent

A new beginning of PC starts with @NVIDIARTXSpark, supercharging what's possible in Hermes Agent. This is the NVIDIA RTX Spark Superchip. A new beginning for personal computers. Designed for creators,

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

SafetyDGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

Model ReleasesDGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

SafetyDGX agent

arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

Model ReleasesDGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

ResearchDGX agent

arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as 'agents' for decision-making (DM) in interactive and dynamic environments. Yet, since they

Secure AI agents with Policy and Lambda interceptors in Amazon Bedrock AgentCore gateway

SafetyDGX agent

In this post, we use a lakehouse data agent to demonstrate how you can use Policy for deterministic access control and Lambda interceptors for dynamic validation. We then show how to combine Lambda in

Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

Model ReleasesDGX agent

arXiv:2605.31408v1 Announce Type: cross Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granulari

The gains are most obvious in coding and operational areas. Even conservative organizations are adopting coding agents fast. That doesn't me…

ApplicationsDGX agent

The gains are most obvious in coding and operational areas. Even conservative organizations are adopting coding agents fast. That doesn't mean AI adoption doesn't come with huge challenges and cost co

We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with t…

HardwareDGX agent

We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with the new OpenShell runtime, which connects Hermes to @Microsof

← Previous
1…130131132133134…300
Next →