AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,919 results
Agents

Improving agents The old way: Manually reading traces, looking for patterns, writing evals, and creating fixes. The better way: Letting Lang…

DGX agent

This post from LangChain's Harrison Chase contrasts traditional manual methods of improving AI agents (tracing execution, identifying patterns, writing evaluations, and implementing fixes) with a more

agentsharrison-chase--x
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

DGX agent

arXiv:2605.29625v1 Announce Type: new Abstract: The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, m

agentsarxiv-cs-ai
29 May 2026
Safety

Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents

DGX agent

arXiv:2605.30190v1 Announce Type: new Abstract: Diffusion-based planning has achieved strong results in single-agent offline reinforcement learning, yet scaling to many-agent systems remains intractab

safetyarxiv-cs-lg
29 May 2026
Model Releases

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

DGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents

DGX agent

arXiv:2605.29534v1 Announce Type: new Abstract: Recent advances in mobile GUI agents have shown strong potential for automating mobile tasks, but most effective systems still depend on large vision-la

local-aiarxiv-cs-ai
29 May 2026
Agents

Asana acquires StackAI to run AI agent workflows across enterprise systems

DGX agent

Work management software company Asana Inc. today said it has completed the acquisition of StackAI Inc., a no-code platform for building artificial intelligence agents, in a deal that adds the ability

agentssiliconangle
28 May 2026
Agents

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

DGX agent

arXiv:2605.27864v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on p

agentsarxiv-cs-ai
28 May 2026
Model Releases

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

DGX agent

arXiv:2605.27922v1 Announce Type: new Abstract: LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, perform

model-releasesarxiv-cs-ai
28 May 2026
Agents

If you're building your own cloud agent like Devin or Ramp Inspect, there's lots of great details here on setting up VMs, computer use, memo…

DGX agent

If you're building your own cloud agent like Devin or Ramp Inspect, there's lots of great details here on setting up VMs, computer use, memory, and more. Fun deep dive with the creator of OpenInspect

agentsswyx--x
28 May 2026
Model Releases

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

DGX agent

arXiv:2605.28721v1 Announce Type: new Abstract: Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diag

model-releasesarxiv-cs-ai
28 May 2026
Agents

Long-Term Mapping of the Douro River Plume with Multi-Agent Reinforcement Learning

DGX agent

arXiv:2510.03534v5 Announce Type: replace-cross Abstract: We study the problem of long-term (multiple days) mapping of a river plume using multiple autonomous underwater vehicles (AUVs), focusing on t

agentsarxiv-cs-lg
28 May 2026
Agents

Modiqo raises $3M in pre-seed funding to help AI agents learn by Rote

DGX agent

Agentic artificial intelligence infrastructure startup Modiqo Inc. says it can assist companies in moving from experimental workflows to repeatable and production-ready systems after raising 3 million

agentssiliconangle
28 May 2026
Model Releases

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

DGX agent

arXiv:2605.27853v1 Announce Type: new Abstract: We present MolLingo, a multi-agent system that emulates the reasoning process of a chemist to automate molecular design. Existing LLM-based approaches e

model-releasesarxiv-cs-ai
28 May 2026
Agents

Plan Before Search: Search Agents Need Plan

DGX agent

arXiv:2605.28354v1 Announce Type: new Abstract: Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a

agentsarxiv-cs-ai
28 May 2026
Model Releases

SkillGrad: Optimizing Agent Skills Like Gradient Descent

DGX agent

arXiv:2605.27760v1 Announce Type: new Abstract: Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However,

model-releasesarxiv-cs-ai
28 May 2026
Agents

SynthTools: A Framework for Scaling Synthetic Tools for Agent Development

DGX agent

arXiv:2511.09572v2 Announce Type: replace Abstract: For agentic systems to use external tools to solve complex, long-horizon tasks, we need a large set of diverse and controllable tool-use environment

agentsarxiv-cs-ai
28 May 2026
Agents

Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem

DGX agent

arXiv:2605.28588v1 Announce Type: cross Abstract: We analyzed 3,984 AI agent skills from major marketplaces and found 76 confirmed malicious payloads, including credential theft, backdoor installation

agentsarxiv-cs-ai
28 May 2026
Agents

The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

DGX agent

This episode of the Latent Space podcast features Walden Yan from Cognition and Cole Murray from OpenInspect discussing the emerging paradigm of asynchronous AI agents—autonomous systems that operate

agentslatent-space
28 May 2026
Agents

UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind

DGX agent

arXiv:2605.27721v1 Announce Type: cross Abstract: Understanding what a user believes and intends is central to building effective agent assistants. This ability is often evaluated through Theory-of-Mi

agentsarxiv-cs-ai
28 May 2026
Safety

Voluntary Collusion with Secret Tools in Competing LLM Agents

DGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

safetyarxiv-cs-ai
28 May 2026
Agents

When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?

DGX agent

arXiv:2605.28224v1 Announce Type: new Abstract: Multi-trajectory inference for tool-use LLM agents - generating multiple reasoning attempts and selecting among them - benefits from transferring knowle

agentsarxiv-cs-ai
28 May 2026
Agents

384 people in the Hermes Agent Jam Session in our discord, happening now! Come check out the latest projects people are working on and showi…

DGX agent

Nous Research announced a live Discord event called the Hermes Agent Jam Session with 384 participants, where developers showcased ongoing projects and innovations. The event provided a venue for the

agentsnous-research--x
27 May 2026
Agents

ATOM: Instantiating Budget-Controllable Multi-Agent Collaboration via Nucleus-Electron Hierarchy

DGX agent

arXiv:2605.26178v1 Announce Type: cross Abstract: Large Language Model (LLM)-based multi-agent systems rely on optimized collaboration topologies to balance performance and communication costs. Howeve

agentsarxiv-cs-lg
27 May 2026
Agents

Excited to dive into this - an open source agent designed with memory/continual learning in mind

DGX agent

This post likely announces or discusses an open-source AI agent framework that incorporates memory and continual learning capabilities, enabling the system to retain information across interactions an

agentsharrison-chase--x
27 May 2026
Agents

📹 How to manage context the right way with LangSmith Context Hub Context Hub gives your team and your agent a shared place to store, edit, …

DGX agent

📹 How to manage context the right way with LangSmith Context Hub Context Hub gives your team and your agent a shared place to store, edit, version, and retrieve context from It can be used for skills,

agentsharrison-chase--x
27 May 2026
Agents

Improving your agent has been a manual process of: ✅ Reading traces ✅ Looking for patterns ✅ Writing evals ✅ Creating fixes Now, LangSmith E…

DGX agent

LangSmith has introduced automated tools to streamline agent improvement, eliminating the manual workflow of reading execution traces, identifying patterns, writing evaluations, and implementing fixes

agentsharrison-chase--x
27 May 2026
Agents

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

DGX agent

arXiv:2603.01131v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in clinical diagnosis but remain limited by unreliable report generation, weak evidence ground

agentsarxiv-cs-ai
27 May 2026
Agents

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

DGX agent

arXiv:2605.26691v1 Announce Type: new Abstract: Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume t

agentsarxiv-cs-ai
27 May 2026
Local Ai

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

DGX agent

arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens wh

local-aiarxiv-cs-ai
27 May 2026
Agents

the future of Continual Learning will rely on building to systems to ingest, understand, & apply knowledge from Agent Traces at scale had a …

DGX agent

the future of Continual Learning will rely on building to systems to ingest, understand, & apply knowledge from Agent Traces at scale had a blast presenting LangSmith Engine with @bentannyhill to show

agentsharrison-chase--x
27 May 2026
Agents

this is a fantastic example of a robust agentic memory system built on langgraph! there's 4 key parts 1. retrieval (rag) 2. storage, fragmen…

DGX agent

this is a fantastic example of a robust agentic memory system built on langgraph! there's 4 key parts 1. retrieval (rag) 2. storage, fragmented by memory type 3. reasoning* 4. learning (often called d

agentsharrison-chase--x
27 May 2026
Agents

Today's Hermes Agent Jam starts in 2 hours, see you soon!

DGX agent

Nous Research announced an upcoming 'Hermes Agent Jam' event starting in 2 hours from the time of posting. The event likely involves developers and AI researchers collaborating on or demonstrating app

agentsnous-research--x
27 May 2026
Model Releases

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

DGX agent

arXiv:2605.26144v1 Announce Type: cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

DGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

model-releasesarxiv-cs-ai
27 May 2026
Agents

We launched Context Hub as a way to manage skills, AGENTS.md files, and other context files an agent might need You can easily use it as a v…

DGX agent

We launched Context Hub as a way to manage skills, AGENTS.md files, and other context files an agent might need You can easily use it as a virtual filesystem in deepagents See this video for more info

agentsharrison-chase--x
27 May 2026
Agents

XGrammar-2: Efficient Dynamic Structured Generation Engine for Agentic LLMs

DGX agent

arXiv:2601.04426v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional structured gen

agentsarxiv-cs-ai
27 May 2026
Model Releases

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

DGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

model-releasesarxiv-cs-ai
26 May 2026
Safety

Auditing medical multi-agent AI reveals risks of false consensus

DGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

safetyarxiv-cs-ai
26 May 2026
Agents

MedBeads: An Agent-Native, Immutable Data Substrate for Trustworthy Medical AI

DGX agent

arXiv:2602.01086v2 Announce Type: replace Abstract: Background: As of 2026, Large Language Models (LLMs) demonstrate expert-level medical knowledge. However, deploying them as autonomous 'Clinical Age

agentsarxiv-cs-ai
26 May 2026
Agents

MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing

DGX agent

arXiv:2605.23986v1 Announce Type: cross Abstract: Memory is a fundamental component for enabling long-context LLM agents, supporting persistent state across interactions through a continuous serve-and

agentsarxiv-cs-ai
26 May 2026
Agents

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation

DGX agent

arXiv:2605.24785v1 Announce Type: new Abstract: Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill di

agentsarxiv-cs-ai
26 May 2026
Agents

Practical Quantum CIM Empowerment via All-Domestic-Core Agentic Large Model

DGX agent

arXiv:2605.23934v1 Announce Type: new Abstract: Quantum computing devices are recognized as powerful tools for solving NP-complete problems. However, the intricacy of their modeling presents notable b

agentsarxiv-cs-ai
26 May 2026
Model Releases

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

DGX agent

arXiv:2605.24218v1 Announce Type: new Abstract: Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how human

model-releasesarxiv-cs-cl
26 May 2026
Agents

Scaling up Energy-Aware Multi-Agent Reinforcement Learning for Mission-Oriented Drone Networks with Individual Reward

DGX agent

arXiv:2605.24992v1 Announce Type: cross Abstract: Multi-agent reinforcement learning (MARL) has shown wide applicability in collaborative systems such as autonomous driving and smart cities for its ab

agentsarxiv-cs-ai
26 May 2026
Model Releases

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

DGX agent

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing b

model-releasesarxiv-cs-ai
26 May 2026
Safety

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

DGX agent

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks

safetyarxiv-cs-ai
26 May 2026
← Previous
1…8889909192…374
Next →