AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
Model Releases

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

DGX agent

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes i

model-releasesarxiv-cs-ai
28 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Microsoft’s introduces its first agent-powered cybersecurity model

DGX agent

Microsoft Corp. today introduced its first in-house cybersecurity model, MAI-Cyber-1-Flash, and a companion agentic system called Project Perception that fields teams of artificial intelligence agents

agentssiliconangle
27 Jul 2026
Agents

again, the log is the agent

DGX agent

The thread argues that applied AI has entered a distributed‑systems phase, with event‑driven architectures treating logs as “agents” that consume data rather than serve as endpoints. It notes that alm

agentsyohei-nakajima--x
26 Jul 2026
Model Releases

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

DGX agent

arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]

DGX agent

Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t

model-releasesr-machinelearning
24 Jul 2026
Model Releases

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

DGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

DGX agent

arXiv:2607.19190v2 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a str

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems

DGX agent

arXiv:2607.13716v1 Announce Type: new Abstract: Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateway

model-releasesarxiv-cs-ai
16 Jul 2026
Agents

Self-Improvements in Modern Agentic Systems: A Survey

DGX agent

arXiv:2607.13104v1 Announce Type: new Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, fro

agentsarxiv-cs-ai
16 Jul 2026
Agents

Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

DGX agent

arXiv:2607.13091v1 Announce Type: cross Abstract: LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feed

agentsarxiv-cs-ai
16 Jul 2026
Model Releases

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

DGX agent

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training

model-releasesarxiv-cs-ai
16 Jul 2026
Agents

Speculate with Memory: Lossless Acceleration for LLM Agents

DGX agent

arXiv:2607.12236v1 Announce Type: cross Abstract: Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle.

agentsarxiv-cs-cl
15 Jul 2026
Model Releases

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

DGX agent

arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace

model-releasesarxiv-cs-ai
15 Jul 2026
Agents

Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants

DGX agent

Proactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Exi

agentsapple-ml-research
14 Jul 2026
Agents

ICYMI - you can now build recursive language models (RLMs) with deepagents: the main agent can write custom code to recursively call subagen…

DGX agent

ICYMI - you can now build recursive language models (RLMs) with deepagents: the main agent can write custom code to recursively call subagents! this is super flexible -- the harness can take any shape

agentsharrison-chase--x
13 Jul 2026
Agents

Effective Strategies for Asynchronous Software Engineering Agents

DGX agent

arXiv:2603.21489v2 Announce Type: replace-cross Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon t

agentsarxiv-cs-ai
9 Jul 2026
Agents

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

DGX agent

arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks

agentsarxiv-cs-ai
9 Jul 2026
Agents

AppAgent: Multimodal Agents as Smartphone Users

DGX agent

arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i

agentsarxiv-cs-cv
7 Jul 2026
Safety

Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems

DGX agent

arXiv:2607.04433v1 Announce Type: cross Abstract: The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward aut

safetyarxiv-cs-cl
7 Jul 2026
Agents

Give your eve agent GitHub tools

DGX agent

This Vercel changelog entry announces GitHub tools integration for Eve agents, enabling developers to connect their AI agents with GitHub repositories and workflows. The tools likely allow Eve agents

agentsvercel-blog
7 Jul 2026
Agents

Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild

DGX agent

arXiv:2607.03316v1 Announce Type: cross Abstract: Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development workflows, yet

agentsarxiv-cs-ai
7 Jul 2026
Agents

Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents

DGX agent

arXiv:2508.01858v3 Announce Type: replace-cross Abstract: Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital envi

agentsarxiv-cs-ai
7 Jul 2026
Safety

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

DGX agent

arXiv:2607.00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous

safetyarxiv-cs-ai
2 Jul 2026
Agents

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

DGX agent

arXiv:2606.29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill librari

agentsarxiv-cs-ai
30 Jun 2026
Safety

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

DGX agent

arXiv:2606.28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defen

safetyarxiv-cs-ai
30 Jun 2026
Safety

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

DGX agent

arXiv:2606.28270v1 Announce Type: new Abstract: The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has funda

safetyarxiv-cs-ai
29 Jun 2026
Agents

A Deterministic Control Plane for LLM Coding Agents

DGX agent

arXiv:2606.26924v1 Announce Type: cross Abstract: LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions, IDE-spe

agentsarxiv-cs-ai
26 Jun 2026
Agents

Retrofit, don’t rebuild: Agentic overlays for transforming legacy enterprise services

DGX agent

In this technical collaboration between AWS and the authors, we present a pragmatic solution: agentic overlays. Agentic overlays are thin wrapper layers that transform traditional REST-based services

agentsaws-ml-blog
25 Jun 2026
Agents

The first WF2026 keynotes are premiering now! Congratulations to: - @techgirl1908, Agentic AI Foundation: Build Systems, Not Code - Nishant …

DGX agent

The first WF2026 keynotes are premiering now! Congratulations to: - @techgirl1908, Agentic AI Foundation: Build Systems, Not Code - Nishant Gupta, Meta Superintelligence Labs: Production Evals For Age

agentsswyx--x
25 Jun 2026
Agents

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

DGX agent

arXiv:2606.24649v1 Announce Type: new Abstract: Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language

agentsarxiv-cs-cv
24 Jun 2026
Agents

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tel…

DGX agent

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tell the agent so many rich details via audio. The longer and t

agentsdair-ai--x
24 Jun 2026
Agents

One thing we’ve seen pretty consistently: shipping the first version is only a small part of the work A key part of building reliable agents…

DGX agent

One thing we’ve seen pretty consistently: shipping the first version is only a small part of the work A key part of building reliable agents is having a repeatable lifecycle for improving them over ti

agentsharrison-chase--x
23 Jun 2026
Agents

Hermes Agent sets you free

DGX agent

Nous Research announced Hermes Agent, a system designed to provide greater autonomy and flexibility for AI agents in task execution and decision-making. The post likely highlights how this agent frame

agentsnous-research--x
11 Jun 2026
Model Releases

Gemini for Government: Your blueprint for mission impact

DGX agent

The public sector has reached a critical inflection point. For years, organizations have explored what’s possible through isolated AI pilots and experimentation. Today, the question has shifted to “wh

model-releasesgoogle-cloud-ai
9 Jun 2026
Tutorials

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

DGX agent

arXiv:2602.08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow hi

tutorialsarxiv-cs-ai
9 Jun 2026
Agents

From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability

DGX agent

arXiv:2606.07150v1 Announce Type: cross Abstract: Agent-interoperability protocols such as A2A and MCP standardize what agents say to one another, but assume address-based transport over HTTP(S). Such

agentsarxiv-cs-ai
8 Jun 2026
Agents

Snowflake and 1Password tackle the growing challenge of securing AI agents at scale

DGX agent

As AI agents gain access to sensitive enterprise data, AI agent security is becoming a top priority for organizations. The challenge is no longer just protecting systems, but ensuring autonomous agent

agentssiliconangle
8 Jun 2026
Agents

Great paper on self-improving agents:

DGX agent

Great paper on self-improving agents: This was one of the standout AI papers of the week. (bookmark it) It tackles a question most self-improving AI agents ignore: is the agent actually discovering an

agentsdair-ai--x
7 Jun 2026
Model Releases

ADK Arena: Evaluating Agent Development Kits via LLM-as-a-Developer

DGX agent

arXiv:2606.05548v1 Announce Type: cross Abstract: The rapid proliferation of Agent Development Kits (ADKs), SDK-level frameworks for building LLM-powered autonomous agents, has outpaced any empirical

model-releasesarxiv-cs-ai
6 Jun 2026
Agents

Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents

DGX agent

arXiv:2603.26233v2 Announce Type: replace Abstract: As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspeci

agentsarxiv-cs-cl
5 Jun 2026
Safety

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

DGX agent

arXiv:2606.05985v1 Announce Type: new Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural back

safetyarxiv-cs-cl
5 Jun 2026
Agents

Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

DGX agent

arXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of age

agentsarxiv-cs-cv
5 Jun 2026
Agents

Building the AI factory for self-improving agents: What’s new in Arize AX

DGX agent

Arize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. The post Building the

agentsarize-ai
4 Jun 2026
Agents

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

DGX agent

arXiv:2603.02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works

agentsarxiv-cs-ai
4 Jun 2026
Local Ai

Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents

DGX agent

arXiv:2606.04555v1 Announce Type: cross Abstract: Long-horizon conversational agents need to interact with users through evolving events, tasks, and goals. Such histories are naturally temporal, yet m

local-aiarxiv-cs-ai
4 Jun 2026
Agents

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

DGX agent

arXiv:2606.02859v1 Announce Type: cross Abstract: How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedric

agentsarxiv-cs-ai
3 Jun 2026
Agents

Overlaying Governance: A Compositional Authorization Framework for Delegation and Scope in Agentic AI

DGX agent

arXiv:2606.03518v1 Announce Type: new Abstract: As AI systems evolve from passive models into autonomous active agents capable of initiating actions, collaborating, and delegating tasks, the tradition

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

DGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…3334353637…370
Next →