AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
28 Jul 2026

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

AgentsDGX agent

arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. Howeve

Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure

Model ReleasesDGX agent

arXiv:2607.22611v1 Announce Type: new Abstract: The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access contro

False Prophets: On the Security of World Models in Agentic Systems

Model ReleasesDGX agent

arXiv:2607.23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

How Databricks manages its own coding agent spend with Unity AI Gateway Budgets

AgentsDGX agent

Databricks employs Unity AI Gateway budgets to oversee spending on its coding agents, enabling tighter cost controls and optimized use of AI resources. The approach helps the organization balance budg

Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2607.23605v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across

I really hope we get details from @OpenAI on the task that as specifies to their rogue agent I'm guessing it was given the full ExploitGym s…

AgentsDGX agent

I really hope we get details from @OpenAI on the task that as specifies to their rogue agent I'm guessing it was given the full ExploitGym suite and told to solve it, with an option to run 5.6-Sol sub

PD^3: A Project Duplication Detection Framework via Adapted Multi-Agent Debate

AgentsDGX agent

arXiv:2505.17492v2 Announce Type: replace Abstract: Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing methods u

Scaling GUI Agents with Visual State Transitions

AgentsDGX agent

arXiv:2607.24112v1 Announce Type: new Abstract: We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

SafetyDGX agent

arXiv:2607.24300v1 Announce Type: new Abstract: Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-au

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

AgentsDGX agent

arXiv:2607.23438v1 Announce Type: new Abstract: As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they

Sources: the OpenAI agent that breached Hugging Face also compromised a customer at AI infrastructure company Modal Labs (Reuters)

AgentsDGX agent

Reuters: Sources: the OpenAI agent that breached Hugging Face also compromised a customer at AI infrastructure company Modal Labs — The rogue agent that escaped from OpenAI and went on a days-long hac

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

AgentsDGX agent

arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes i

27 Jul 2026

Microsoft’s introduces its first agent-powered cybersecurity model

AgentsDGX agent

Microsoft Corp. today introduced its first in-house cybersecurity model, MAI-Cyber-1-Flash, and a companion agentic system called Project Perception that fields teams of artificial intelligence agents

26 Jul 2026

again, the log is the agent

AgentsDGX agent

The thread argues that applied AI has entered a distributed‑systems phase, with event‑driven architectures treating logs as “agents” that consume data rather than serve as endpoints. It notes that alm

24 Jul 2026

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

Model ReleasesDGX agent

arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta

I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]

Model ReleasesDGX agent

Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

Model ReleasesDGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

23 Jul 2026

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

SafetyDGX agent

arXiv:2607.19190v2 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a str

16 Jul 2026

CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems

Model ReleasesDGX agent

arXiv:2607.13716v1 Announce Type: new Abstract: Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateway

Self-Improvements in Modern Agentic Systems: A Survey

AgentsDGX agent

arXiv:2607.13104v1 Announce Type: new Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, fro

Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

AgentsDGX agent

arXiv:2607.13091v1 Announce Type: cross Abstract: LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feed

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

Model ReleasesDGX agent

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training

15 Jul 2026

Speculate with Memory: Lossless Acceleration for LLM Agents

AgentsDGX agent

arXiv:2607.12236v1 Announce Type: cross Abstract: Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle.

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

Model ReleasesDGX agent

arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace

14 Jul 2026

Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants

AgentsDGX agent

Proactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Exi

13 Jul 2026

ICYMI - you can now build recursive language models (RLMs) with deepagents: the main agent can write custom code to recursively call subagen…

AgentsDGX agent

ICYMI - you can now build recursive language models (RLMs) with deepagents: the main agent can write custom code to recursively call subagents! this is super flexible -- the harness can take any shape

9 Jul 2026

Effective Strategies for Asynchronous Software Engineering Agents

AgentsDGX agent

arXiv:2603.21489v2 Announce Type: replace-cross Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon t

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

AgentsDGX agent

arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks

7 Jul 2026

AppAgent: Multimodal Agents as Smartphone Users

AgentsDGX agent

arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i

Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems

SafetyDGX agent

arXiv:2607.04433v1 Announce Type: cross Abstract: The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward aut

Give your eve agent GitHub tools

AgentsDGX agent

This Vercel changelog entry announces GitHub tools integration for Eve agents, enabling developers to connect their AI agents with GitHub repositories and workflows. The tools likely allow Eve agents

Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild

AgentsDGX agent

arXiv:2607.03316v1 Announce Type: cross Abstract: Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development workflows, yet

Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents

AgentsDGX agent

arXiv:2508.01858v3 Announce Type: replace-cross Abstract: Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital envi

2 Jul 2026

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

SafetyDGX agent

arXiv:2607.00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous

30 Jun 2026

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

AgentsDGX agent

arXiv:2606.29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill librari

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

SafetyDGX agent

arXiv:2606.28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defen

29 Jun 2026

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

SafetyDGX agent

arXiv:2606.28270v1 Announce Type: new Abstract: The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has funda

26 Jun 2026

A Deterministic Control Plane for LLM Coding Agents

AgentsDGX agent

arXiv:2606.26924v1 Announce Type: cross Abstract: LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-spe

25 Jun 2026

Retrofit, don’t rebuild: Agentic overlays for transforming legacy enterprise services

AgentsDGX agent

In this technical collaboration between AWS and the authors, we present a pragmatic solution: agentic overlays. Agentic overlays are thin wrapper layers that transform traditional REST-based services

The first WF2026 keynotes are premiering now! Congratulations to: - @techgirl1908, Agentic AI Foundation: Build Systems, Not Code - Nishant …

AgentsDGX agent

The first WF2026 keynotes are premiering now! Congratulations to: - @techgirl1908, Agentic AI Foundation: Build Systems, Not Code - Nishant Gupta, Meta Superintelligence Labs: Production Evals For Age

24 Jun 2026

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

AgentsDGX agent

arXiv:2606.24649v1 Announce Type: new Abstract: Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tel…

AgentsDGX agent

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tell the agent so many rich details via audio. The longer and t

23 Jun 2026

One thing we’ve seen pretty consistently: shipping the first version is only a small part of the work A key part of building reliable agents…

AgentsDGX agent

One thing we’ve seen pretty consistently: shipping the first version is only a small part of the work A key part of building reliable agents is having a repeatable lifecycle for improving them over ti

11 Jun 2026

Hermes Agent sets you free

AgentsDGX agent

Nous Research announced Hermes Agent, a system designed to provide greater autonomy and flexibility for AI agents in task execution and decision-making. The post likely highlights how this agent frame

9 Jun 2026

Gemini for Government: Your blueprint for mission impact

Model ReleasesDGX agent

The public sector has reached a critical inflection point. For years, organizations have explored what’s possible through isolated AI pilots and experimentation. Today, the question has shifted to “wh

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

TutorialsDGX agent

arXiv:2602.08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow hi

8 Jun 2026

From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability

AgentsDGX agent

arXiv:2606.07150v1 Announce Type: cross Abstract: Agent-interoperability protocols such as A2A and MCP standardize what agents say to one another, but assume address-based transport over HTTP(S). Such

Snowflake and 1Password tackle the growing challenge of securing AI agents at scale

AgentsDGX agent

As AI agents gain access to sensitive enterprise data, AI agent security is becoming a top priority for organizations. The challenge is no longer just protecting systems, but ensuring autonomous agent

7 Jun 2026

Great paper on self-improving agents:

AgentsDGX agent

Great paper on self-improving agents: This was one of the standout AI papers of the week. (bookmark it) It tackles a question most self-improving AI agents ignore: is the agent actually discovering an

6 Jun 2026

ADK Arena: Evaluating Agent Development Kits via LLM-as-a-Developer

Model ReleasesDGX agent

arXiv:2606.05548v1 Announce Type: cross Abstract: The rapid proliferation of Agent Development Kits (ADKs), SDK-level frameworks for building LLM-powered autonomous agents, has outpaced any empirical

5 Jun 2026

Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents

AgentsDGX agent

arXiv:2603.26233v2 Announce Type: replace Abstract: As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspeci

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

SafetyDGX agent

arXiv:2606.05985v1 Announce Type: new Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural back

Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

AgentsDGX agent

arXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of age

4 Jun 2026

Building the AI factory for self-improving agents: What’s new in Arize AX

AgentsDGX agent

Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. The post Building the

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

AgentsDGX agent

arXiv:2603.02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works

Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents

Local AiDGX agent

arXiv:2606.04555v1 Announce Type: cross Abstract: Long-horizon conversational agents need to interact with users through evolving events, tasks, and goals. Such histories are naturally temporal, yet m

3 Jun 2026

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

AgentsDGX agent

arXiv:2606.02859v1 Announce Type: cross Abstract: How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedric

Overlaying Governance: A Compositional Authorization Framework for Delegation and Scope in Agentic AI

AgentsDGX agent

arXiv:2606.03518v1 Announce Type: new Abstract: As AI systems evolve from passive models into autonomous active agents capable of initiating actions, collaborating, and delegating tasks, the tradition

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

← Previous
1…2627282930…296
Next →