AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,911 results
8 Jul 2026

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce A

Delay-Aware Active Triangulation with Uncertainty-Driven Multi-Agent Reinforcement Learning for Counter-UAS

AgentsDGX agent

arXiv:2607.05957v1 Announce Type: new Abstract: Multi-agent active visual triangulation enables precise 3D localization of aerial targets by coordinating mobile observers with controllable cameras. Ho

Love partnering with baseten to make sure everyone can use open weight models in deep agents

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

This post discusses Baseten's partnership efforts to democratize access to open-weight models for use in AI agents, making advanced model capabilities available to a broader audience. The initiative a

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2607.06008v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external env

Solidigm targets the intelligence layer as agentic inference pushes storage to center stage

AgentsDGX agent

The shift from model training to agentic inference is forcing a fundamental rethink of how artificial intelligence infrastructure is built and which components carry the most strategic weight. What wa

The agent is the user now: lessons from the founder of WorkOS

AgentsDGX agent

WorkOS founder Michael Grinich explains why the next era of AI engineering depends on the systems around agents: identity, permissions, evals, memory, and feedback loops that keep autonomous software

7 Jul 2026

Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence

AgentsDGX agent

arXiv:2602.24115v2 Announce Type: replace Abstract: Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder t

An Exploration of Agentic Information Fusion for Test Maintenance Prediction

AgentsDGX agent

arXiv:2607.04786v1 Announce Type: cross Abstract: Test maintenance is a critical, yet costly, activity - particularly as codebases rapidly evolve. To assist, we present MAST, a multi-agent framework t

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games

Model ReleasesDGX agent

arXiv:2607.04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally reli

Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality

Model ReleasesDGX agent

arXiv:2607.03691v1 Announce Type: cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middlewa

Expanding Managed Agents in Gemini API: background tasks, remote MCP and more

Model ReleasesDGX agent

Google expands Managed Agents in Gemini API with support for background execution , and adds the ability to register remote Model Context Protocol (MCP) servers to extend agent capabilities . Develope

FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

AgentsDGX agent

arXiv:2607.04718v1 Announce Type: new Abstract: Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This work

kAgent: An execution-guided crash resolution agent for the Linux kernel

AgentsDGX agent

arXiv:2504.20412v3 Announce Type: replace-cross Abstract: Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive. However

NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads

HardwareDGX agent

NVIDIA Vera is a CPU designed to help AI factories scale agentic AI and reinforcement learning by shortening CPU execution time, increasing task throughput, and enabling smarter, longer-thinking agent

OpenTinker: Separating Concerns in Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2601.07376v2 Announce Type: replace Abstract: We introduce extsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared exe

ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration

Model ReleasesDGX agent

arXiv:2607.03730v1 Announce Type: new Abstract: Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentally passive and reactive: they respond to explici

SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery

Local AiDGX agent

arXiv:2607.02807v1 Announce Type: new Abstract: Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems. However, they tend to converge onto a s

Toward Efficient Agents: Memory, Tool learning, and Planning

Model ReleasesDGX agent

arXiv:2601.14192v2 Announce Type: replace Abstract: Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has conti

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

SafetyDGX agent

arXiv:2607.04425v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform int

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

SafetyDGX agent

arXiv:2607.05132v1 Announce Type: cross Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents tha

6 Jul 2026

AI agent exploits Langflow in first fully autonomous ransomware attack

AgentsDGX agent

Cloud security company Sysdig Inc. has documented what it says is the first ransomware operation carried out from start to finish by an autonomous artificial intelligence agent, a campaign it calls Ja

Own the loop: A field guide to agent harnesses

AgentsDGX agent

As models become cheaper and more interchangeable, the durable advantage shifts to the agent harness: the loop, tools, memory, permissions, and workflow you can own and refine. The post Own the loop:

Soon most agents for work will run in the cloud and communication with them will happen almost exclusively in your existing work channels (S…

AgentsDGX agent

Soon most agents for work will run in the cloud and communication with them will happen almost exclusively in your existing work channels (Slack, Teams, ...) This hasn't become the norm yet because lo

5 Jul 2026

ActiveGraph makes the agent trace the runtime: Yohei Nakajima's open-source Python runtime treats an append-only event log as the source of …

AgentsDGX agent

ActiveGraph makes the agent trace the runtime: Yohei Nakajima's open-source Python runtime treats an append-only event log as the source of truth, enabling replay, forking, and lineage for long-runnin

3 Jul 2026

A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development

AgentsDGX agent

arXiv:2603.04390v2 Announce Type: replace Abstract: WebGIS development requires consistency, yet agentic AI often fails due to LLM context constraints, forgetting, stochasticity, instruction failure,

BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_

Coding-agents can replicate scientific machine learning papers

AgentsDGX agent

arXiv:2607.02134v1 Announce Type: new Abstract: Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predic

In this interview at @aiDotEngineer World's Fair, @vercel chief of software @andrewqu explains why agents represent a new form of software, …

AgentsDGX agent

In this interview at @aiDotEngineer World's Fair, @vercel chief of software @andrewqu explains why agents represent a new form of software, what Vercel learned from building its own, and why Vercel it

Mark Zuckerberg says Meta’s agentic AI efforts aren’t progressing as fast as he had hoped

AgentsDGX agent

Meta Platforms Inc. Chief Executive Mark Zuckerberg told employees at an internal town hall meeting that the company’s work on artificial intelligence agents hasn’t progressed as quickly as he had hop

MMAO-Cls: Metabolic Multi-Agent Optimization for Joint Feature Selection and Classifier Tuning

AgentsDGX agent

arXiv:2607.01539v1 Announce Type: cross Abstract: This paper studies whether the Metabolic Multi-Agent Optimizer (MMAO) can act as a credible outer-loop optimizer for classification model selection. W

Simulation Based Reward Function Validation for Multi-Agent On Orbit Inspection

AgentsDGX agent

arXiv:2607.01367v1 Announce Type: cross Abstract: A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MARL). While MARL has already been employe

Steerability via constraints: a substrate for scalable oversight of coding agents

Model ReleasesDGX agent

arXiv:2607.02389v1 Announce Type: new Abstract: Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human

The Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2607.01415v1 Announce Type: new Abstract: Coding-agent reinforcement learning treats execution infrastructure as a background implementation detail, despite relying on large numbers of interacti

2 Jul 2026

At a town hall, Mark Zuckerberg said Meta's AI agent development has not accelerated as expected and its reorganization was not as 'clean' as it could have been (Katie Paul/Reuters)

AgentsDGX agent

Katie Paul / Reuters: At a town hall, Mark Zuckerberg said Meta's AI agent development has not accelerated as expected and its reorganization was not as “clean” as it could have been — Meta (META.O) C

BaRA: BFS-and-Reflection Web Data Collection Agent

AgentsDGX agent

arXiv:2607.00007v1 Announce Type: cross Abstract: Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, ret

Fugu is now available on OpenCode! ✨ When our team was developing Fugu’s multi-agent orchestration, OpenCode was our tool of choice to verif…

AgentsDGX agent

Fugu is now available on OpenCode! ✨ When our team was developing Fugu’s multi-agent orchestration, OpenCode was our tool of choice to verify our models. We share a core philosophy with the OpenCode t

GameDevBench: Evaluating Agentic Capabilities Through Game Development

Model ReleasesDGX agent

arXiv:2602.11103v2 Announce Type: replace Abstract: Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

AgentsDGX agent

arXiv:2601.04424v2 Announce Type: replace Abstract: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclea

I really like this 'understand to participate' framing of the cognitive debt problem when working with coding agents

AgentsDGX agent

I really like this 'understand to participate' framing of the cognitive debt problem when working with coding agents That's where another answer comes in: we can understand to participate. You can lea

I’ll be in room 2005 (graph track) at 11:10am to talk about ActiveGraph: event-sourced graph runtime for building auditable agents! Been hav…

AgentsDGX agent

Yohei Nakajima will present ActiveGraph, an event-sourced graph runtime designed for building auditable agents, at 11:10am in room 2005 of a conference's graph track. The presentation focuses on lever

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection

AgentsDGX agent

arXiv:2607.00035v1 Announce Type: new Abstract: LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, bro

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

AgentsDGX agent

arXiv:2607.00597v1 Announce Type: new Abstract: Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, an

NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents

AgentsDGX agent

arXiv:2601.14660v2 Announce Type: replace-cross Abstract: Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enabling t

On the last day of @aiDotEngineer Worlds Fair SF, I am SO excited to be releasing the latest episode of the Agentic Review podcast, featurin…

AgentsDGX agent

On the last day of @aiDotEngineer Worlds Fair SF, I am SO excited to be releasing the latest episode of the Agentic Review podcast, featuring @PaulDuvall, author of Continuous Integration, and AI-nati

Pinecone releases Nexus into public preview to bring business knowledge to AI agents

AgentsDGX agent

Pinecone Systems Inc., an artificial intelligence infrastructure company providing fully managed vector databases, Wednesday launched the public preview of Pinecone Nexus, which curates and distribute

Self-GC: Self-Governing Context for Long-Horizon LLM Agents

AgentsDGX agent

arXiv:2607.00692v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. C

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests

Model ReleasesDGX agent

arXiv:2607.00990v1 Announce Type: cross Abstract: Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue re

WorkBench Revisited: Workplace Agents Two Years On

Model ReleasesDGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

1 Jul 2026

Autoresearch: The feedback loop behind self-improving agents

AgentsDGX agent

Autoresearch describes a feedback mechanism where AI agents can evaluate their own outputs and use that introspection to iteratively improve their reasoning and decision-making capabilities. This conc

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

AgentsDGX agent

arXiv:2606.19501v2 Announce Type: replace Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

AgentsDGX agent

arXiv:2606.31980v1 Announce Type: new Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a mul

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents

AgentsDGX agent

arXiv:2606.31046v1 Announce Type: new Abstract: Artificial life has explored life-like behavior on many computational substrates, but mostly in researcher-designed closed worlds. We argue that large l

The best players want to be coached. You build trust, then you push hard. Now here's the thing: your agent has no skin in the game. no ego t…

AgentsDGX agent

The best players want to be coached. You build trust, then you push hard. Now here's the thing: your agent has no skin in the game. no ego to protect, no trust to earn first. So you can and should ski

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

Model ReleasesDGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

AgentsDGX agent

arXiv:2606.30801v1 Announce Type: new Abstract: Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors

We’re also publishing extensive documentation and technical materials about Agentic MapReduce, including a deep-dive on our evals. Read our …

AgentsDGX agent

We’re also publishing extensive documentation and technical materials about Agentic MapReduce, including a deep-dive on our evals. Read our announcement: https://cognition.com/blog/introducing-devin-s

30 Jun 2026

A living map of everything your Hermes agent has learned; every memory & skill Press play to watch it unfold, @NousResearch Copy yours into …

AgentsDGX agent

This post from Nous Research demonstrates a visualization tool or system that maps the accumulated knowledge, memories, and skills learned by a Hermes AI agent during its training or operation. The in

An expanded Vercel Agent: chat, investigations, and approved actions, now in public beta

AgentsDGX agent

Vercel has released an expanded version of its Vercel Agent in public beta, introducing new capabilities including chat functionality, investigation tools, and an approved actions feature. The update

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Model ReleasesDGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

← Previous
1…6667686970…299
Next →