AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
Model Releases

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

DGX agent

Ivan Mehta / TechCrunch: Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release — Hark, a startup

model-releasestechmeme
5 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

DGX agent

arXiv:2608.03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. Howe

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

S^3: Improving Agent Safety through Multi-Stage Defense

DGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach

DGX agent

arXiv:2608.02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrast

agentsarxiv-cs-ai
5 Aug 2026
Agents

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

DGX agent

arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central

agentsarxiv-cs-ai
5 Aug 2026
Model Releases

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

DGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch

DGX agent

arXiv:2608.00316v1 Announce Type: new Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by

agentsarxiv-cs-lg
4 Aug 2026
Model Releases

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

DGX agent

arXiv:2608.01918v1 Announce Type: cross Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable

model-releasesarxiv-cs-cl
4 Aug 2026
Tutorials

How to debug production AI agents with Signal in Arize AX

DGX agent

Learn how Arize Signal turns production traces into ranked issues, proposed fixes, regression datasets, and reviewable pull requests for AI agents. The post How to debug production AI agents with Sign

tutorialsarize-ai
4 Aug 2026
Agents

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

DGX agent

arXiv:2512.10739v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

DGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

DGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

DGX agent

arXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect an

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

DGX agent

arXiv:2607.29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, ex

safetyarxiv-cs-ai
3 Aug 2026
Agents

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published wher…

DGX agent

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, 'Active Graph Agent Runtime (BabyAGI 4)', on

agentsyohei-nakajima--x
2 Aug 2026
Agents

datasette-apps 0.2a0

DGX agent

Release: datasette-apps 0.2a0 Changes that improve Datasette Apps when created and edited using Datasette Agent: New app_debug() tool allowing agent to open an app (invisibly) and test it using JavaSc

agentssimon-willison
1 Aug 2026
Hardware

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

DGX agent

arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large langu

hardwarearxiv-cs-ai
31 Jul 2026
Model Releases

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

DGX agent

arXiv:2607.28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outc

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

DGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

model-releasesarxiv-cs-ai
31 Jul 2026
Safety

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

DGX agent

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents

safetyarxiv-cs-cv
31 Jul 2026
Safety

LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

DGX agent

arXiv:2607.27690v1 Announce Type: new Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolv

safetyarxiv-cs-ro
31 Jul 2026
Agents

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

DGX agent

arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist

agentsarxiv-cs-cl
31 Jul 2026
Local Ai

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

DGX agent

arXiv:2607.26300v1 Announce Type: new Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematic

local-aiarxiv-cs-cl
30 Jul 2026
Safety

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

DGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

safetyarxiv-cs-cl
30 Jul 2026
Research

Echoverse: Deep, evolving environments for computer-use agents

DGX agent

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping t

researchmicrosoft-research
30 Jul 2026
Agents

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

DGX agent

arXiv:2607.25560v1 Announce Type: new Abstract: Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and priv

agentsarxiv-cs-ai
29 Jul 2026
Safety

AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading

DGX agent

arXiv:2605.05580v2 Announce Type: replace Abstract: Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio execution. Howev

safetyarxiv-cs-ai
29 Jul 2026
Agents

Distributing Security Controls Through Harness Engineering

DGX agent

arXiv:2607.25890v1 Announce Type: new Abstract: AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizatio

agentsarxiv-cs-ai
29 Jul 2026
Agents

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

DGX agent

arXiv:2607.25877v1 Announce Type: new Abstract: This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focu

agentsarxiv-cs-ai
29 Jul 2026
Agents

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

DGX agent

ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughp

agentstogether-ai-blog
29 Jul 2026
Agents

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

DGX agent

arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas

agentsarxiv-cs-ai
29 Jul 2026
Agents

VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents

DGX agent

arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing t

agentsarxiv-cs-ai
29 Jul 2026
Local Ai

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

DGX agent

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research traj

local-aiarxiv-cs-ai
28 Jul 2026
Model Releases

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

DGX agent

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden informat

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

DGX agent

arXiv:2607.24097v1 Announce Type: new Abstract: Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evi

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

Perplexity brings its Personal Computer AI agent to Windows

DGX agent

Perplexity AI Inc. today released a Windows version of Personal Computer, expanding its agentic automation software beyond the original Macintosh platform and making it available to more than 1 billio

model-releasessiliconangle
28 Jul 2026
Model Releases

SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs

DGX agent

arXiv:2607.22571v1 Announce Type: new Abstract: Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing ag

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

DGX agent

arXiv:2607.24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However,

model-releasesarxiv-cs-ai
28 Jul 2026
Agents

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

DGX agent

arXiv:2607.18080v2 Announce Type: replace-cross Abstract: Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associ

agentsarxiv-cs-ai
28 Jul 2026
Safety

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

DGX agent

arXiv:2607.22758v1 Announce Type: cross Abstract: The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication t

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

as the progenitor of the agent lab thesis which got the evals/routing/interactivity/ROI focus right i gotta say the biggest argument against…

DGX agent

as the progenitor of the agent lab thesis which got the evals/routing/interactivity/ROI focus right i gotta say the biggest argument against myself is that Claude Code got accidentally 'open sourced'

model-releasesswyx--x
27 Jul 2026
Model Releases

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

DGX agent

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation a

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

love this frame. calls to mind the role of the hippocampus in human navigation (via place cells and grid cells), and how navigation is, in a…

DGX agent

love this frame. calls to mind the role of the hippocampus in human navigation (via place cells and grid cells), and how navigation is, in a sense, what makes agents *agents* vs plain old LLM calls in

model-releasesyohei-nakajima--x
27 Jul 2026
Agents

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

DGX agent

arXiv:2607.21887v1 Announce Type: cross Abstract: Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferat

agentsarxiv-cs-cl
27 Jul 2026
Agents

“one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s inte…

DGX agent

“one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s internal constraints, per sources” I don’t think this kind of pr

agentsgary-marcus--x
25 Jul 2026
Model Releases

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

DGX agent

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

model-releasesarxiv-cs-ai
24 Jul 2026
Local Ai

No sé nada de Ollama, ni programación ni idea, pero estoy creando un agente evolutivo

DGX agent

Con ayuda de ChatGPT y con el modelo de Ollama, Qwen3:14b estoy creando un agente que corre local y tiene la iniciativa para pensar, investigar, aprender, generar propuestas y esperar mi autorización

local-air-ollama
24 Jul 2026
Agents

Nutanix and AMD build enterprise AI stack to take agents from pilot to production

DGX agent

Enterprise AI ambition is far outpacing the readiness of the enterprise AI stack, and the growing rift between proof-of-concept deployments and production-scale agent operations is where most organiza

agentssiliconangle
24 Jul 2026
← Previous
1…5859606162…371
Next →