AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
Agents

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

DGX agent

arXiv:2608.06171v1 Announce Type: new Abstract: Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across

agentsarxiv-cs-cl
7 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

DGX agent

arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks.

agentsarxiv-cs-cl
7 Aug 2026
Agents

Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning

DGX agent

arXiv:2508.13661v4 Announce Type: replace Abstract: Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to a

agentsarxiv-cs-lg
6 Aug 2026
Safety

Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

DGX agent

Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give y

safetyaws-ml-blog
6 Aug 2026
Agents

Fully onboard with productizing hillclimbing as an automated service for any agentic task. I've had the fortune of knowing @silennai since t…

DGX agent

Fully onboard with productizing hillclimbing as an automated service for any agentic task. I've had the fortune of knowing @silennai since the AutoGPT days, and I know that him and Kion are going to d

agentsjerry-liu--x
6 Aug 2026
Agents

State2State: Environment-Derived Mid-Training for LLM Agents

DGX agent

arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with

agentsarxiv-cs-cl
6 Aug 2026
Agents

stratum: A System Infrastructure for Massive Agent-Centric ML Workloads

DGX agent

arXiv:2603.03589v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated. LLMs enable a new t

agentsarxiv-cs-lg
6 Aug 2026
Agents

What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills

DGX agent

arXiv:2608.04562v1 Announce Type: new Abstract: Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study

agentsarxiv-cs-ai
6 Aug 2026
Agents

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

DGX agent

arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations b

agentsarxiv-cs-ai
5 Aug 2026
Tools

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's a…

DGX agent

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's also how a billion users will soon use ChatGPT by default. I

toolsswyx--x
5 Aug 2026
Model Releases

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

DGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

model-releasesarxiv-cs-cl
5 Aug 2026
Agents

Traceable Multi-Agent System for Knowledge-Based Forecasting

DGX agent

arXiv:2608.03339v1 Announce Type: new Abstract: Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this

agentsarxiv-cs-ai
5 Aug 2026
Agents

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

DGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

agentsarxiv-cs-ai
5 Aug 2026
Agents

ArmorCode targets runaway AI costs with four new remediation agents

DGX agent

Exposure management startup ArmorCode Inc. today used Black Hat USA 2026 in Las Vegas to detail an expansion of its agentic artificial intelligence platform, adding four planned agents and three new s

agentssiliconangle
4 Aug 2026
Model Releases

Control Under Compression: Reliability Frontiers for Tool-Using Agents

DGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Cross-Domain Hybrid OPD for Generalizable Search Agents

DGX agent

arXiv:2608.02101v1 Announce Type: new Abstract: Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

DGX agent

arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

DGX agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

DGX agent

arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Su

safetyarxiv-cs-cl
4 Aug 2026
Agents

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

DGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

agentsarxiv-cs-cl
4 Aug 2026
Agents

Unpacking ChatGPT Work: the Agent for a Billion Users

DGX agent

ChatGPT Work was launched by OpenAI on July 9, 2026 as an agent‑oriented knowledge‑work platform that combines chat, Codex tools and cloud agents across fourteen model configurations. Within three wee

agentslatent-space
4 Aug 2026
Agents

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

DGX agent

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an ext

agentsarxiv-cs-cl
3 Aug 2026
Safety

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

DGX agent

arXiv:2607.29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gra

safetyarxiv-cs-ai
3 Aug 2026
Agents

Announcing the Agentic Catalog Experience in Amazon Quick

DGX agent

Amazon Quick introduces the Agentic Catalog Experience, an AI-powered workflow for data curators to discover upstream catalog assets in natural language and auto-create Datasets and Topics with inheri

agentsaws-ml-blog
31 Jul 2026
Model Releases

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

DGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

model-releasesarxiv-cs-ai
31 Jul 2026
Model Releases

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

DGX agent

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisti

model-releasesarxiv-cs-ai
31 Jul 2026
Agents

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

DGX agent

arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity with

agentsarxiv-cs-cl
30 Jul 2026
Model Releases

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

DGX agent

arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

DGX agent

arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental too

agentsarxiv-cs-lg
30 Jul 2026
Agents

From Signal to PR: What if your agents got better every time they failed?

DGX agent

Signal, a managed agent built into Arize AX, continuously reviews production traces, surfaces ranked issues with evidence and proposed fixes, and — with Managed Agents — can carry investigations into

agentsarize-ai
29 Jul 2026
Model Releases

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

DGX agent

arXiv:2607.25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (

model-releasesarxiv-cs-ai
29 Jul 2026
Agents

Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation

DGX agent

arXiv:2607.24006v1 Announce Type: cross Abstract: Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through legitimate

agentsarxiv-cs-ai
28 Jul 2026
Model Releases

Diagrid Catalyst 2.0 adds durable execution to more than 10 agent frameworks

DGX agent

Agent infrastructure startup Diagrid Inc. today released Catalyst 2.0, an update to its managed workflow engine that adds automatic failure recovery and cryptographic verification to artificial intell

model-releasessiliconangle
28 Jul 2026
Agents

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

DGX agent

arXiv:2607.24419v1 Announce Type: new Abstract: Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and i

agentsarxiv-cs-ai
28 Jul 2026
Local Ai

Falsifiable Commitment Planning for Self-Correcting Web Agents

DGX agent

arXiv:2607.24167v1 Announce Type: new Abstract: Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, o

local-aiarxiv-cs-ai
28 Jul 2026
Agents

Since launching Qwen3.8-Max-Preview, we've received valuable feedback from developers! To help users better explore Qwen3.8's agentic capabi…

DGX agent

Since launching Qwen3.8-Max-Preview, we've received valuable feedback from developers! To help users better explore Qwen3.8's agentic capabilities, we're officially launching the #QwenGrowthPlan today

agentsqwen--x
28 Jul 2026
Model Releases

Agentic Evaluation of Copyright Law Compliance

DGX agent

arXiv:2607.21799v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate,

model-releasesarxiv-cs-cl
27 Jul 2026
Agents

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

DGX agent

arXiv:2508.12393v3 Announce Type: replace Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet curr

agentsarxiv-cs-cl
27 Jul 2026
Hardware

Six Agent Harness Capabilities for Higher Model Performance

DGX agent

The performance of AI agents depends not only on the underlying models but also heavily on their “harness”—the surrounding architecture that supplies context, state management, action execution, and t

hardwarenvidia-developer
27 Jul 2026
Agents

Way Security, which uses AI-driven automation and agentic workflows to help deploy IAM systems, raised a $20M seed from Insight Partners and Glilot Capital (Chris Metinko/Axios)

DGX agent

Chris Metinko / Axios: Way Security, which uses AI-driven automation and agentic workflows to help deploy IAM systems, raised a 20M seed from Insight Partners and Glilot Capital — Way Security raised

agentstechmeme
27 Jul 2026
Agents

We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its …

DGX agent

We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast

agentskimi-moonshot--x
27 Jul 2026
Model Releases

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

DGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

model-releasesr-localllama
25 Jul 2026
Model Releases

Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants

DGX agent

arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

DGX agent

arXiv:2607.21522v1 Announce Type: cross Abstract: Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graph

agentsarxiv-cs-ai
24 Jul 2026
Model Releases

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

DGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

DGX agent

arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel thr

model-releasesarxiv-cs-ai
23 Jul 2026
Agents

Environment-free Synthetic Data Generation for API-Calling Agents

DGX agent

arXiv:2607.16900v2 Announce Type: replace Abstract: Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale

agentsarxiv-cs-ai
23 Jul 2026
Agents

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems avai…

DGX agent

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems available to everyone are getting extremely powerful (even as th

agentsethan-mollick--x
23 Jul 2026
← Previous
1…4950515253…371
Next →