AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
6 Aug 2026

stratum: A System Infrastructure for Massive Agent-Centric ML Workloads

AgentsDGX agent

arXiv:2603.03589v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated. LLMs enable a new t

What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills

AgentsDGX agent

arXiv:2608.04562v1 Announce Type: new Abstract: Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study

5 Aug 2026

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

Agents
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations b

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's a…

ToolsDGX agent

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's also how a billion users will soon use ChatGPT by default. I

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Model ReleasesDGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

Traceable Multi-Agent System for Knowledge-Based Forecasting

AgentsDGX agent

arXiv:2608.03339v1 Announce Type: new Abstract: Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

AgentsDGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

4 Aug 2026

ArmorCode targets runaway AI costs with four new remediation agents

AgentsDGX agent

Exposure management startup ArmorCode Inc. today used Black Hat USA 2026 in Las Vegas to detail an expansion of its agentic artificial intelligence platform, adding four planned agents and three new s

Control Under Compression: Reliability Frontiers for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2608.01056v1 Announce Type: cross Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments,

Cross-Domain Hybrid OPD for Generalizable Search Agents

SafetyDGX agent

arXiv:2608.02101v1 Announce Type: new Abstract: Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Model ReleasesDGX agent

arXiv:2608.00355v1 Announce Type: new Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

Model ReleasesDGX agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

SafetyDGX agent

arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Su

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

AgentsDGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

Unpacking ChatGPT Work: the Agent for a Billion Users

AgentsDGX agent

ChatGPT Work was launched by OpenAI on July 9, 2026 as an agent‑oriented knowledge‑work platform that combines chat, Codex tools and cloud agents across fourteen model configurations. Within three wee

3 Aug 2026

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

AgentsDGX agent

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an ext

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

SafetyDGX agent

arXiv:2607.29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gra

31 Jul 2026

Announcing the Agentic Catalog Experience in Amazon Quick

AgentsDGX agent

Amazon Quick introduces the Agentic Catalog Experience, an AI-powered workflow for data curators to discover upstream catalog assets in natural language and auto-create Datasets and Topics with inheri

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Model ReleasesDGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Model ReleasesDGX agent

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisti

30 Jul 2026

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

AgentsDGX agent

arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity with

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Model ReleasesDGX agent

arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

AgentsDGX agent

arXiv:2607.27083v1 Announce Type: new Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental too

29 Jul 2026

From Signal to PR: What if your agents got better every time they failed?

AgentsDGX agent

Signal, a managed agent built into Arize AX, continuously reviews production traces, surfaces ranked issues with evidence and proposed fixes, and — with Managed Agents — can carry investigations into

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

Model ReleasesDGX agent

arXiv:2607.25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (

28 Jul 2026

Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation

AgentsDGX agent

arXiv:2607.24006v1 Announce Type: cross Abstract: Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through legitimate

Diagrid Catalyst 2.0 adds durable execution to more than 10 agent frameworks

Model ReleasesDGX agent

Agent infrastructure startup Diagrid Inc. today released Catalyst 2.0, an update to its managed workflow engine that adds automatic failure recovery and cryptographic verification to artificial intell

Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers

AgentsDGX agent

arXiv:2607.24419v1 Announce Type: new Abstract: Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and i

Falsifiable Commitment Planning for Self-Correcting Web Agents

Local AiDGX agent

arXiv:2607.24167v1 Announce Type: new Abstract: Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, o

Since launching Qwen3.8-Max-Preview, we've received valuable feedback from developers! To help users better explore Qwen3.8's agentic capabi…

AgentsDGX agent

Since launching Qwen3.8-Max-Preview, we've received valuable feedback from developers! To help users better explore Qwen3.8's agentic capabilities, we're officially launching the #QwenGrowthPlan today

27 Jul 2026

Agentic Evaluation of Copyright Law Compliance

Model ReleasesDGX agent

arXiv:2607.21799v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate,

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

AgentsDGX agent

arXiv:2508.12393v3 Announce Type: replace Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet curr

Six Agent Harness Capabilities for Higher Model Performance

HardwareDGX agent

The performance of AI agents depends not only on the underlying models but also heavily on their “harness”—the surrounding architecture that supplies context, state management, action execution, and t

Way Security, which uses AI-driven automation and agentic workflows to help deploy IAM systems, raised a $20M seed from Insight Partners and Glilot Capital (Chris Metinko/Axios)

AgentsDGX agent

Chris Metinko / Axios: Way Security, which uses AI-driven automation and agentic workflows to help deploy IAM systems, raised a 20M seed from Insight Partners and Glilot Capital — Way Security raised

We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its …

AgentsDGX agent

We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast

25 Jul 2026

CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows

Model ReleasesDGX agent

If you run local agentic coding harnesses (Aider, Claude Code, etc.), prompt evaluation usually eats up most of your execution time. Every turn re-evaluates thousands of identical prefix tokens_system

24 Jul 2026

Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants

Model ReleasesDGX agent

arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

AgentsDGX agent

arXiv:2607.21522v1 Announce Type: cross Abstract: Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graph

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

Model ReleasesDGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

23 Jul 2026

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel thr

Environment-free Synthetic Data Generation for API-Calling Agents

AgentsDGX agent

arXiv:2607.16900v2 Announce Type: replace Abstract: Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems avai…

AgentsDGX agent

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems available to everyone are getting extremely powerful (even as th

21 Jul 2026

We're hosting the August LA Agentic AI Meetup on Tuesday, August 11, 5 to 7pm at Gulp in Playa Vista. May drew 70 people. July brought 80. L…

AgentsDGX agent

We're hosting the August LA Agentic AI Meetup on Tuesday, August 11, 5 to 7pm at Gulp in Playa Vista. May drew 70 people. July brought 80. Let's break 100 in August. Come hang out with builders, engin

16 Jul 2026

Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes

Model ReleasesDGX agent

arXiv:2607.13071v1 Announce Type: cross Abstract: Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper docume

Set-shifting Behavioral Test for Harnessed Agents

Model ReleasesDGX agent

arXiv:2607.13396v1 Announce Type: new Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psyc

15 Jul 2026

Agents-A1-4B (Qwen3.7-4B ???) : Scaling the Horizon, Not the Parameters

Model ReleasesDGX agent

MODEL + GGUF : https://huggingface.co/InternScience/models?search=a1-4b Technical Report Benchmark Qwen3.5-4B Agents-A1-4B Qwen3.5 Qwen3.6 Nex-N2-mini Agents-A1 🧠 Dense Models (~4B) 🔀 MoE Models (35B-

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

Model ReleasesDGX agent

arXiv:2607.12058v1 Announce Type: cross Abstract: Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe oper

EFLUX: Elastic Multi-Robot Formation Navigation and Adaptation with Agentic LLMs

AgentsDGX agent

arXiv:2607.12050v1 Announce Type: new Abstract: Multi-robot teams operating in confined or cluttered environments must adapt both their formation geometry and group topology to navigate through comple

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

AgentsDGX agent

arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions

SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning

AgentsDGX agent

arXiv:2607.12042v1 Announce Type: new Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing m

14 Jul 2026

How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo

HardwareDGX agent

Autonomous coding agents such as Codex (GPT‑5.5) can fully automate reinforcement‑learning research workflows by provisioning GPU‑hosted environments, orchestrating experiments, and iteratively optimi

13 Jul 2026

Building an agentic AI solution at Bluesight with Amazon Bedrock

AgentsDGX agent

In this post, we describe how Bluesight used two AWS engagements and Amazon Bedrock AgentCore to evolve from a single-product AI prototype to Prism, a unified agentic AI solution spanning six healthca

10 Jul 2026

Calibrated Stackelberg Games: Learning Optimal Commitments Against Calibrated Agents

AgentsDGX agent

arXiv:2306.02704v2 Announce Type: replace-cross Abstract: We introduce Calibrated Stackelberg Games (CSGs), a generalization of the standard Stackelberg Games (SGs) framework. In CSGs, a principal rep

The agentic marketing stack starts with the data layer

AgentsDGX agent

Agentic marketing leverages AI agents to automate marketing tasks, and this approach requires a robust data foundation as its core component. Databricks discusses how organizations need unified data p

Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

AgentsDGX agent

arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference

9 Jul 2026

Holy…? First public demo of an iPhone agent ordering me a donut by @agi_inc & @divgarg. Just late night roommate sessions at @AGIHouseSF

AgentsDGX agent

A demonstration of an AI agent with iPhone integration that autonomously ordered a donut, showcasing early progress in agentic AI capabilities. The demo was presented publicly by AGI Inc and Div Garg,

TIL about soak testing my agents

AgentsDGX agent

Soak testing for AI agents involves running agents continuously over extended periods to identify performance degradation, resource leaks, stability issues, and edge cases that may not appear during s

Towards Agentic AI Governance: A Preliminary Assessment

AgentsDGX agent

arXiv:2607.07612v1 Announce Type: cross Abstract: Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely charact

8 Jul 2026

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

Model ReleasesDGX agent

arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and ope

7 Jul 2026

CONTRA: Red-Teaming Configurations of Personalizable Agents

SafetyDGX agent

arXiv:2607.03220v1 Announce Type: cross Abstract: Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These sy

← Previous
1…3940414243…296
Next →