AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
11 Aug 2026

Software Engineering for and with GUI Agent

SafetyDGX agent

arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity

SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users (Zac Hall/9to5Mac)

AgentsDGX agent

Zac Hall / 9to5Mac: SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users — SpaceXAI and Cursor

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework

AgentsDGX agent

arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbook

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
10 Aug 2026

A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers

Model ReleasesDGX agent

arXiv:2608.06694v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is la

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

Model ReleasesDGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery

Model ReleasesDGX agent

arXiv:2608.07126v1 Announce Type: cross Abstract: Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178 missions f

Plan-and-Avoid: Real-Time Aircraft Trajectory Coordination in a Multi-Agent Environment

AgentsDGX agent

arXiv:2608.06648v1 Announce Type: new Abstract: This paper presents a real-time Plan-and-Avoid (PAA framework for coordinating cooperative multi-agent airspace operations around a declared priority tr

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

Model ReleasesDGX agent

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multim…

AgentsDGX agent

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multimodal workloads. Start building: https://www.together.ai/mode

7 Aug 2026

Comparative Approaches to Agent Retrieval over Large Skill Libraries

AgentsDGX agent

arXiv:2608.06196v1 Announce Type: new Abstract: Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and prov

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

AgentsDGX agent

arXiv:2608.05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Spli

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

Model ReleasesDGX agent

arXiv:2608.05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal r

This might be a good time to mention my recent keynotes titled: 'Agentic AI is Neurosymbolic AI'

AgentsDGX agent

This might be a good time to mention my recent keynotes titled: 'Agentic AI is Neurosymbolic AI' I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known a

WorldClaw: Agentic 3D Open-World Generation at Scale

AgentsDGX agent

arXiv:2608.05248v1 Announce Type: new Abstract: Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coher

6 Aug 2026

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

ResearchDGX agent

arXiv:2608.05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. Ho

Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills

AgentsDGX agent

arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while

Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking

AgentsDGX agent

arXiv:2608.04768v1 Announce Type: new Abstract: Automated cooking robots have traditionally relied on predefined procedures and rule-based control, ensuring stable execution but offering limited perso

InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

Model ReleasesDGX agent

arXiv:2608.04761v1 Announce Type: cross Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience

OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee (Zac Hall/9to5Mac)

Model ReleasesDGX agent

Zac Hall / 9to5Mac: OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee — OpenAI's GPT-5 tur

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Model ReleasesDGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

SafetyDGX agent

arXiv:2608.04436v1 Announce Type: new Abstract: Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understandi

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Model ReleasesDGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

SafetyDGX agent

arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens

5 Aug 2026

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

AgentsDGX agent

arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these syst

ETA: A New Agentic Paradigm for Embodied Tasks

AgentsDGX agent

arXiv:2608.03924v1 Announce Type: new Abstract: When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environm

Field Aware Agent Skill Retrieval

AgentsDGX agent

arXiv:2608.02880v1 Announce Type: cross Abstract: As lifelong learning agents accumulate lifelong growing skill banks, retrieving the correct skill becomes an increasingly important bottleneck. Most c

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

SafetyDGX agent

arXiv:2608.02670v1 Announce Type: cross Abstract: Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root ex

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Local AiDGX agent

arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often

4 Aug 2026

AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents

Model ReleasesDGX agent

arXiv:2608.00832v1 Announce Type: new Abstract: Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a cand

Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale

Model ReleasesDGX agent

arXiv:2608.00101v1 Announce Type: cross Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different fr

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

AgentsDGX agent

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las

BiCAA: Bidirectional Credit Assignment for Search-Augmented Agent

SafetyDGX agent

arXiv:2608.01321v1 Announce Type: new Abstract: Multi-step search is a fundamental capability for search agents, enabling them to iteratively acquire, refine, and integrate external evidence for compl

Global Optimization and Inference-Time Region Grafting for Agentic Workflows

Model ReleasesDGX agent

arXiv:2608.02353v1 Announce Type: new Abstract: Recent advances in agentic workflow optimization automate workflow design through task-specific workflow search or input-conditioned architecture select

HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning

Model ReleasesDGX agent

arXiv:2608.01358v1 Announce Type: new Abstract: Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question

Picking the right agent harness is now a crucial skill for any AI engineer. Imagine using the same model, same task, and same prompt. Now mo…

Model ReleasesDGX agent

Picking the right agent harness is now a crucial skill for any AI engineer. Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success c

3 Aug 2026

Finally a good paper testing whether agent memory needs an LLM at all. Production memory stacks spend extra model calls on summarizing inter…

AgentsDGX agent

Finally a good paper testing whether agent memory needs an LLM at all. Production memory stacks spend extra model calls on summarizing interactions, writing records, and reranking retrievals. Every on

GOP AGs warn OpenAI's Altman to preserve records in AI agent hacking probe https://www.foxbusiness.com/technology/gop-ags-warn-openai-altman…

AgentsDGX agent

GOP AGs warn OpenAI's Altman to preserve records in AI agent hacking probe https://www.foxbusiness.com/technology/gop-ags-warn-openai-altman-preserve-records-ai-agent-hacking-probe?%3Fintcmp=tw_fbn&ta

Self-Supervised Skill Optimization

AgentsDGX agent

arXiv:2607.28777v1 Announce Type: new Abstract: Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized

Transcript-Managed Transformers: Monotone Multi-Agent Collapse and Universality with Two Pop-Enabled Transcripts

AgentsDGX agent

arXiv:2607.29496v1 Announce Type: new Abstract: We study transcript management for fixed, finite-precision causal Transformers. A transcript is partitioned into channels of bounded blocks. Each transi

Try Qwen3.8-Max on Hermes Agent and you will have to doubt on how much these open frontier models have caught up with frontier closed models…

AgentsDGX agent

Try Qwen3.8-Max on Hermes Agent and you will have to doubt on how much these open frontier models have caught up with frontier closed models. These new open models are insanely good. Meet Qwen3.8-Max:

When an agent starts with shared context, more people can get reliable answers. When truth is shared, AI becomes infrastucture. Read about h…

AgentsDGX agent

When an agent starts with shared context, more people can get reliable answers. When truth is shared, AI becomes infrastucture. Read about how our internal truth layer drives our teams at Replit: http

2 Aug 2026

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skil…

AgentsDGX agent

I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skills from OpenAI, Anthropic, HuggingFace & more. Thank me late

31 Jul 2026

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

AgentsDGX agent

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume t

Procedural Fairness in Multi-Agent Bandits

SafetyDGX agent

arXiv:2601.10600v2 Announce Type: replace-cross Abstract: In the context of multi-agent multi-armed bandits (MA-MAB), fairness is often reduced to outcomes: maximizing welfare, reducing inequality, or

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

Model ReleasesDGX agent

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

30 Jul 2026

Just for reference, from what I observed in my Using Local Coding Agents blog article last month: https://x.com/rasbt/status/207051816739969…

Model ReleasesDGX agent

Just for reference, from what I observed in my Using Local Coding Agents blog article last month: https://x.com/rasbt/status/2070518167399698490?s=20 'I tried to analyze why Claude Code uses more toke

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

SafetyDGX agent

arXiv:2607.26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learnin

Software Engineers: Do you honestly get anything useful out of LLMs?

AgentsDGX agent

For 6 months now I've been trying to make agentic coding work for me, using Pi and a handful 30-120B models (Qwens, Nemotrons, Leguna...etc). I'm not greedy either, I stick to decent quants, never qua

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

SafetyDGX agent

arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained in

29 Jul 2026

After a few more hours, I think I've figured out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models a…

Model ReleasesDGX agent

After a few more hours, I think I've figured out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~1…

AgentsDGX agent

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~10x lower P50 latency at high concurrency. ThunderAgent was a

Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

AgentsDGX agent

arXiv:2607.25340v1 Announce Type: new Abstract: The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: i

ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

SafetyDGX agent

arXiv:2607.25283v1 Announce Type: new Abstract: This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes th

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

Model ReleasesDGX agent

arXiv:2607.25400v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify no

From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance

AgentsDGX agent

arXiv:2607.24791v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive impl

How agentic AI can help telecom finance teams protect the margin when every moment matters

AgentsDGX agent

Agentic AI empowers telecom finance teams to preserve margins by rapidly detecting and preventing revenue leakage across billing, provisioning, and cost‑allocation processes. The technology automates

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Model ReleasesDGX agent

arXiv:2607.25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluat

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.25891v1 Announce Type: new Abstract: Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on nar

← Previous
1…7677787980…299
Next →