AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
7 Aug 2026

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Local AiDGX agent

arXiv:2608.05784v1 Announce Type: new Abstract: Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the

Agentic Software Issue Resolution with Large Language Models: A Survey

AgentsDGX agent

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users,

ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

AgentsDGX agent

arXiv:2608.05201v1 Announce Type: cross Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lac

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics

AgentsDGX agent

arXiv:2601.21800v4 Announce Type: replace Abstract: We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks.

Learning Globally Reusable Skills for Coding Agents

AgentsDGX agent

arXiv:2608.06153v1 Announce Type: cross Abstract: Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches

6 Aug 2026

Agent Skills for Automated Reasoning policies in Amazon Bedrock

SafetyDGX agent

Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, and validates a custo

AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection

AgentsDGX agent

arXiv:2608.04053v1 Announce Type: cross Abstract: Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous e

Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

AgentsDGX agent

arXiv:2608.04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the oth

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

AgentsDGX agent

arXiv:2604.07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair,

Securing AI agents with temporal policies in Amazon Bedrock AgentCore

AgentsDGX agent

Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that evaluate authorization based on an agent's session history. Learn how to enforce workflow sequencing, prevent data fabr

5 Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

AgentsDGX agent

arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is los

Cloudflare launches Cloudflare OS: an open-source AI agentic workspace for the enterprise

Model ReleasesDGX agent

Cloudflare Inc. today announced the launch of Cloudflare OS, an open-source artificial intelligence agentic workspace available through the browser, filled with custom shared micro-applications for en

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release — Hark, a startup

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

AgentsDGX agent

arXiv:2608.03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. Howe

S^3: Improving Agent Safety through Multi-Stage Defense

Model ReleasesDGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach

AgentsDGX agent

arXiv:2608.02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrast

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

AgentsDGX agent

arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

Model ReleasesDGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

4 Aug 2026

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch

AgentsDGX agent

arXiv:2608.00316v1 Announce Type: new Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.01918v1 Announce Type: cross Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable

How to debug production AI agents with Signal in Arize AX

TutorialsDGX agent

Learn how Arize Signal turns production traces into ranked issues, proposed fixes, regression datasets, and reviewable pull requests for AI agents. The post How to debug production AI agents with Sign

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

AgentsDGX agent

arXiv:2512.10739v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

Model ReleasesDGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

3 Aug 2026

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Model ReleasesDGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Model ReleasesDGX agent

arXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect an

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

SafetyDGX agent

arXiv:2607.29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, ex

2 Aug 2026

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published wher…

AgentsDGX agent

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, 'Active Graph Agent Runtime (BabyAGI 4)', on

1 Aug 2026

datasette-apps 0.2a0

AgentsDGX agent

Release: datasette-apps 0.2a0 Changes that improve Datasette Apps when created and edited using Datasette Agent: New app_debug() tool allowing agent to open an app (invisibly) and test it using JavaSc

31 Jul 2026

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

HardwareDGX agent

arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large langu

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Model ReleasesDGX agent

arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outc

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Model ReleasesDGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

SafetyDGX agent

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents

LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

SafetyDGX agent

arXiv:2607.27690v1 Announce Type: new Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolv

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

AgentsDGX agent

arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist

30 Jul 2026

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

Local AiDGX agent

arXiv:2607.26300v1 Announce Type: new Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematic

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

SafetyDGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

Echoverse: Deep, evolving environments for computer-use agents

ResearchDGX agent

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping t

29 Jul 2026

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

AgentsDGX agent

arXiv:2607.25560v1 Announce Type: new Abstract: Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and priv

AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading

SafetyDGX agent

arXiv:2605.05580v2 Announce Type: replace Abstract: Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio execution. Howev

Distributing Security Controls Through Harness Engineering

AgentsDGX agent

arXiv:2607.25890v1 Announce Type: new Abstract: AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizatio

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

AgentsDGX agent

arXiv:2607.25877v1 Announce Type: new Abstract: This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focu

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

AgentsDGX agent

ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughp

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

AgentsDGX agent

arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

AgentsDGX agent

arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing t

28 Jul 2026

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Local AiDGX agent

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research traj

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

Model ReleasesDGX agent

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

SafetyDGX agent

arXiv:2607.24097v1 Announce Type: new Abstract: Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evi

Perplexity brings its Personal Computer AI agent to Windows

Model ReleasesDGX agent

Perplexity AI Inc. today released a Windows version of Personal Computer, expanding its agentic automation software beyond the original Macintosh platform and making it available to more than 1 billio

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

Model ReleasesDGX agent

arXiv:2607.22571v1 Announce Type: new Abstract: Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing ag

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Model ReleasesDGX agent

arXiv:2607.24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However,

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

AgentsDGX agent

arXiv:2607.18080v2 Announce Type: replace-cross Abstract: Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associ

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

SafetyDGX agent

arXiv:2607.22758v1 Announce Type: cross Abstract: The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication t

27 Jul 2026

as the progenitor of the agent lab thesis which got the evals/routing/interactivity/ROI focus right i gotta say the biggest argument against…

Model ReleasesDGX agent

as the progenitor of the agent lab thesis which got the evals/routing/interactivity/ROI focus right i gotta say the biggest argument against myself is that Claude Code got accidentally 'open sourced'

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

Model ReleasesDGX agent

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation a

love this frame. calls to mind the role of the hippocampus in human navigation (via place cells and grid cells), and how navigation is, in a…

Model ReleasesDGX agent

love this frame. calls to mind the role of the hippocampus in human navigation (via place cells and grid cells), and how navigation is, in a sense, what makes agents *agents* vs plain old LLM calls in

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

AgentsDGX agent

arXiv:2607.21887v1 Announce Type: cross Abstract: Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferat

25 Jul 2026

“one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s inte…

AgentsDGX agent

“one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s internal constraints, per sources” I don’t think this kind of pr

24 Jul 2026

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

Model ReleasesDGX agent

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

No sé nada de Ollama, ni programación ni idea, pero estoy creando un agente evolutivo

Local AiDGX agent

Con ayuda de ChatGPT y con el modelo de Ollama, Qwen3:14b estoy creando un agente que corre local y tiene la iniciativa para pensar, investigar, aprender, generar propuestas y esperar mi autorización

Nutanix and AMD build enterprise AI stack to take agents from pilot to production

AgentsDGX agent

Enterprise AI ambition is far outpacing the readiness of the enterprise AI stack, and the growing rift between proof-of-concept deployments and production-scale agent operations is where most organiza

← Previous
1…4647484950…297
Next →