AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
7 Jul 2026

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2602.21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interacti

LLMoxie: Exploring Agentic AI for Scientific Software Development

AgentsDGX agent

arXiv:2607.02703v1 Announce Type: cross Abstract: In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a Lite

MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control

SafetyDGX agent

arXiv:2509.23960v2 Announce Type: replace-cross Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi-ag

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

my hot take is that evals are the only part of agent engineering that require real thinking

AgentsDGX agent

Harrison Chase argues that evaluations (evals) are the most intellectually demanding aspect of agent engineering, distinguishing them from other engineering tasks that may be more routine or formulaic

OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement

AgentsDGX agent

arXiv:2607.05346v1 Announce Type: new Abstract: We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-r

Radware adds Claude Code protection and compliance reporting to agent security

Model ReleasesDGX agent

Application and network security company Radware Ltd. today expanded its Agentic AI Protection product with compliance reporting, deeper visibility into artificial intelligence agent activity, and new

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies

SafetyDGX agent

arXiv:2607.03423v1 Announce Type: cross Abstract: Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool guardrails

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Model ReleasesDGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

AgentsDGX agent

arXiv:2607.04235v1 Announce Type: new Abstract: Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address th

Storage gets promoted in the agentic AI era

AgentsDGX agent

The year 2026 could be remembered as the moment when storage technology received a massive promotion. The reason is that the current transition from simple chatbots to agentic AI systems has raised th

Strategic Buying Agents

SafetyDGX agent

arXiv:2607.04708v1 Announce Type: cross Abstract: Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy

TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews

AgentsDGX agent

arXiv:2503.20666v2 Announce Type: replace-cross Abstract: Thematic analysis (TA) is a widely used qualitative approach for uncovering latent meanings in unstructured text data. TA provides valuable in

The 'I Don't Know' Filter: Enhancing Agentic Reliability in Function Calling

AgentsDGX agent

arXiv:2607.04034v1 Announce Type: cross Abstract: The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the traini

3 Jul 2026

Agent Runs now available in the Vercel MCP and CLI

AgentsDGX agent

Vercel has announced the availability of Agent Runs in both the Vercel MCP (Model Context Protocol) and CLI (Command Line Interface), expanding developer capabilities for AI-powered automation and dep

AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

Model ReleasesDGX agent

arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop rea

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

AgentsDGX agent

arXiv:2607.01846v1 Announce Type: new Abstract: Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This paper presents

ContextNest: Verifiable Context Governance for Autonomous AI Agent

AgentsDGX agent

arXiv:2607.02116v1 Announce Type: new Abstract: Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of prov

Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

AgentsDGX agent

arXiv:2501.07892v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simplicit

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Model ReleasesDGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

2 Jul 2026

A Task-State Representation for Long-Horizon Mobile GUI Agents

AgentsDGX agent

arXiv:2607.00502v1 Announce Type: new Abstract: While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient

Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework

AgentsDGX agent

arXiv:2607.01034v1 Announce Type: cross Abstract: Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change. Their cap

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

AgentsDGX agent

arXiv:2607.01087v1 Announce Type: cross Abstract: Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low

From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents

Model ReleasesDGX agent

arXiv:2607.00233v1 Announce Type: new Abstract: How do two agents invent a shared language from scratch? In a Lewis signaling game, a sender and receiver must coordinate on a code using only their int

SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents

AgentsDGX agent

arXiv:2607.00011v1 Announce Type: cross Abstract: Reusable skill libraries are becoming important infrastructure for large language model (LLM) agents, yet existing selection methods often treat skill

1 Jul 2026

AI-Assisted Discovery of Convex Relaxations via Dual Agents

AgentsDGX agent

arXiv:2606.31182v1 Announce Type: new Abstract: Recent work shows that LLM agents can improve sharp-constant inequalities by searching for extremal constructions, which yield upper bounds. We address

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation

SafetyDGX agent

arXiv:2603.26266v3 Announce Type: replace Abstract: Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction. However, due to i

if i were building an agent from scratch with memory i would use a wiki structure, it's simple and extensible

AgentsDGX agent

Harrison Chase suggests using a wiki structure as the foundation for building AI agents with memory systems, citing its simplicity and extensibility as key advantages. This approach would organize age

if you want agents to do work at scale (like security triage, trace analysis, document parsing), they need structured workflows enforced w/ …

AgentsDGX agent

if you want agents to do work at scale (like security triage, trace analysis, document parsing), they need structured workflows enforced w/ code map reduce is a great example -- exactly the kind of pa

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

Model ReleasesDGX agent

arXiv:2606.31504v1 Announce Type: new Abstract: We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own s

We gave a 2 hr deepdive on how to build inference engines that handle trillion token agentic workloads at @aiDotEngineer. Will drop slides a…

AgentsDGX agent

Together AI presented a 2-hour technical deep dive on building inference engines capable of handling trillion-token agentic workloads, covering architecture and optimization strategies for large-scale

30 Jun 2026

Couchbase’s AI Data Plane aims to turn fragmented data into real enterprise agent memory

AgentsDGX agent

Couchbase Inc. is trying to solve one of the hardest problems in enterprise artificial intelligence today: turning brittle, chat-style pilots into production-grade agents capable of remembering, reaso

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

AgentsDGX agent

arXiv:2606.30296v1 Announce Type: new Abstract: Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode:

Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

SafetyDGX agent

arXiv:2603.05786v2 Announce Type: replace-cross Abstract: As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which int

Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution

AgentsDGX agent

arXiv:2606.28971v1 Announce Type: new Abstract: Real-world image restoration (IR) remains challenging due to complex and coupled degradations. While recent agentic IR frameworks leverage Large Languag

29 Jun 2026

Semantic search alone doesn't cut it. Neither does brute-force grep. Agents need both. Today we're shipping the Retrieval Harness in LlamaPa…

AgentsDGX agent

Semantic search alone doesn't cut it. Neither does brute-force grep. Agents need both. Today we're shipping the Retrieval Harness in LlamaParse Index: semantic search, server-side grep, and file-level

26 Jun 2026

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Model ReleasesDGX agent

arXiv:2606.14397v2 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabi

When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework

SafetyDGX agent

arXiv:2606.26400v1 Announce Type: new Abstract: Agentic systems are changing how complex operational tasks are coordinated, introducing a new paradigm for connecting heterogeneous data sources and aut

25 Jun 2026

deployment cookbook for langchain agents!

ApplicationsDGX agent

deployment cookbook for langchain agents! Agents are easy to demo locally. The hard part is shipping them inside a real app. We published a deployment cookbook for @LangChain agents: full-stack exampl

Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents

AgentsDGX agent

arXiv:2601.03785v3 Announce Type: replace Abstract: Long-term human-agent dialogues are organized by topic continuity: adjacent turns often develop the same goal, plan, problem, or event, while relate

Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound

Model ReleasesDGX agent

arXiv:2606.25978v1 Announce Type: cross Abstract: Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis sp

Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows

AgentsDGX agent

arXiv:2604.25345v2 Announce Type: replace Abstract: Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently

24 Jun 2026

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

SafetyDGX agent

arXiv:2606.24428v1 Announce Type: new Abstract: Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experien

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

AgentsDGX agent

arXiv:2606.24839v1 Announce Type: new Abstract: Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evalu

GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.24551v1 Announce Type: new Abstract: Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Model ReleasesDGX agent

arXiv:2606.24595v1 Announce Type: new Abstract: Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interactio

'Most agents don't learn, they just leave traces.' In 12 minutes, @jakebroekhuizen breaks down how to actually close the loop. Surface issue…

AgentsDGX agent

'Most agents don't learn, they just leave traces.' In 12 minutes, @jakebroekhuizen breaks down how to actually close the loop. Surface issues with LangSmith Engine Write memory updates back to Context

Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games

SafetyDGX agent

arXiv:2509.18371v2 Announce Type: replace-cross Abstract: Multi-agent games in dynamic nonlinear settings are challenging due to the time-varying interactions among the agents and the non-stationarity

pretty sick that i get to work with Jake every day on making continual learning + memory accessible at scale for every single agent one comm…

AgentsDGX agent

pretty sick that i get to work with Jake every day on making continual learning + memory accessible at scale for every single agent one common thread here is...the Trace a large part of Continual Lear

SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation

Model ReleasesDGX agent

arXiv:2606.24626v1 Announce Type: new Abstract: As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even

Skills for the future software profession: beyond agentic AI!

AgentsDGX agent

arXiv:2606.21894v2 Announce Type: replace-cross Abstract: As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engineer

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

AgentsDGX agent

arXiv:2606.23743v1 Announce Type: cross Abstract: Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration me

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

AgentsDGX agent

arXiv:2511.07397v2 Announce Type: replace Abstract: Voice agents face a fundamental tension: the reasoning, retrieval, and tool use that make foundation models capable are iterative and slow, while co

23 Jun 2026

A trillion tokens a day and 200k GitHub stars Very proud of the Hermes Agent team and what we've built at @NousResearch 'Better today than y…

AgentsDGX agent

Nous Research announced that their Hermes Agent has achieved a trillion tokens processed daily and reached 200,000 GitHub stars, reflecting significant adoption and performance milestones for their AI

GRADE: Graph Representation of LLM Agent Dependency and Execution

AgentsDGX agent

arXiv:2606.22741v1 Announce Type: new Abstract: Can one graph represent every kind of LLM agent's run? A trace records what each step did, never what it relied on, the state it read, and the results i

Hermes Agent can now /learn from anything: feed it directories of any source material (code, API docs, manuals, PDFs, configs) and it distil…

AgentsDGX agent

Hermes Agent has been updated to accept diverse source materials including code repositories, API documentation, manuals, PDFs, and configuration files as input, with the ability to distill and learn

PolicyGuard: Towards Test-time and Step-level Adversary (Backdoor) Defense for Reinforcement Learning Agent

AgentsDGX agent

arXiv:2606.12896v2 Announce Type: replace Abstract: While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention an

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

AgentsDGX agent

arXiv:2606.20474v2 Announce Type: replace Abstract: Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrency determi

22 Jun 2026

Hermes Agent now supports computer use via @trycua on Windows and Linux in addition to existing macOS support

AgentsDGX agent

Hermes Agent, developed by Nous Research, has expanded its computer use capabilities to support Windows and Linux operating systems in addition to its existing macOS functionality. This enhancement, i

11 Jun 2026

APPO: Agentic Procedural Policy Optimization

SafetyDGX agent

arXiv:2606.12384v1 Announce Type: cross Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents

ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories

AgentsDGX agent

arXiv:2606.11520v1 Announce Type: cross Abstract: Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution

← Previous
1…5657585960…297
Next →