AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
11 Jun 2026

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

AgentsDGX agent

arXiv:2606.11830v1 Announce Type: new Abstract: Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical s

10 Jun 2026

Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents

AgentsDGX agent

arXiv:2606.10315v1 Announce Type: cross Abstract: LLM-as-judge is the default instrument for evaluating conversational agents, yet its reliability is almost always reported as agreement with human rat

Exclusive: Relai raises $6.9M to enable verifiable and continuous learning for AI agents

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

Artificial intelligence infrastructure startup Relai Inc. said today it has closed on 6.9 million in funding as it bids to ensure the reliability of autonomous AI agents for enterprises. The company a

From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents

AgentsDGX agent

arXiv:2606.09863v1 Announce Type: new Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise. We study this failure mode, false success, across

Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory

AgentsDGX agent

arXiv:2606.10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions. Existing memory systems often s

shoutouts: • why multi-agent LLM systems fail? (arXiv:2503.13657) — @mertcemri @melissapan + @istoica05 @matei_zaharia @profjoeyg @adityagp …

AgentsDGX agent

shoutouts: • why multi-agent LLM systems fail? (arXiv:2503.13657) — @mertcemri @melissapan + @istoica05 @matei_zaharia @profjoeyg @adityagp & team • DSPy (arXiv:2310.03714) — @lateinteraction + @hazyr

This is just awesomeness from @cohere, @nickfrosst, and team. I so badly want a coding agent that just runs on my local machine. We are not …

AgentsDGX agent

This is just awesomeness from @cohere, @nickfrosst, and team. I so badly want a coding agent that just runs on my local machine. We are not too far now! Excited to get this to work with my @dair_ai co

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Model ReleasesDGX agent

arXiv:2606.11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely

9 Jun 2026

Also, I found that Hermes Agent + Nemotron 3 Ultra is a mighty combo!

Model ReleasesDGX agent

Also, I found that Hermes Agent + Nemotron 3 Ultra is a mighty combo! Excited to launch a new way to upskill with AI agents. This is how we are making it possible for anyone to learn to build with cod

ConMem: Structured Memory-Guided Adaptation in Training-Free Multi-Agent Systems

AgentsDGX agent

arXiv:2606.08702v1 Announce Type: new Abstract: Recent advances have improved the adaptive capabilities of LLM-based multi-agent systems (MAS) through memory-, skill-, and learning-based approaches, y

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents

AgentsDGX agent

arXiv:2606.07904v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly rely on external APIs, but standard tool schemas describe how to call a tool, not when the tool

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

Model ReleasesDGX agent

arXiv:2606.08367v1 Announce Type: cross Abstract: Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this approach is mism

Goal-Oriented Reasoning for RAG-based Memory in Conversational Agentic LLM Systems

AgentsDGX agent

arXiv:2605.12213v2 Announce Type: replace Abstract: LLM-based conversational AI agents struggle to maintain coherent behavior over long horizons due to limited context. While RAG-based approaches are

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and re…

AgentsDGX agent

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and remain frozen or mostly unchanged. The harness, like the skill

SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation

AgentsDGX agent

arXiv:2606.09774v1 Announce Type: new Abstract: Advanced scientific simulators expose specialized input languages that turn simulation goals into executable configurations, but learning them can cost

Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing

AgentsDGX agent

arXiv:2606.08256v1 Announce Type: new Abstract: Verifiability, attribution, and reproducibility are foundational requirements of scientific knowledge, yet current publishing infrastructure does not en

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Model ReleasesDGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

8 Jun 2026

Autonomous computational catalysis through an agentic research system

AgentsDGX agent

arXiv:2601.13508v4 Announce Type: replace-cross Abstract: Autonomous agents are beginning to transform scientific research from tool-assisted workflows toward self-sustaining discovery processes. Comp

Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

AgentsDGX agent

arXiv:2606.06976v1 Announce Type: new Abstract: Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct response

How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope

AgentsDGX agent

arXiv:2606.07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute t

Off-Policy Evaluation with Strategic Agents via Local Disclosure

Local AiDGX agent

arXiv:2606.07308v1 Announce Type: new Abstract: We study off-policy evaluation (OPE) under strategic behavior where decision subjects (or agents) respond to a decision maker's policy by strategically

OpenSkill: Open-World Self-Evolution for LLM Agents

AgentsDGX agent

arXiv:2606.06741v1 Announce Type: new Abstract: Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful tra

Pega expands AI platform with agent orchestration, development tools and new pricing model

AgentsDGX agent

Workflow automation vendor Pegasystems Inc. today unveiled a broad set of artificial intelligence enhancements aimed at helping enterprises deploy AI agents in mission-critical business processes whil

The Agent Open: AI's Pickleball Tournament 🏓 Come put your code and backhand to the test and embrace the full Open experience. Custom built…

AgentsDGX agent

The Agent Open: AI's Pickleball Tournament 🏓 Come put your code and backhand to the test and embrace the full Open experience. Custom built out courts. Stadium seating. Exhibition matches by AI leader

The Open Source Community is backing OpenEnv for Agentic RL

AgentsDGX agent

OpenEnv is an open-source framework backed by the community for training and developing agentic reinforcement learning systems. The project represents collaborative efforts within the open-source ecos

7 Jun 2026

datasette-agent-edit 0.1a0

Model ReleasesDGX agent

Release: datasette-agent-edit 0.1a0 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, updating large SQL qu

6 Jun 2026

good dev tools are cached intelligence for agents!

IndustryDGX agent

good dev tools are cached intelligence for agents! Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write c

5 Jun 2026

Unsupervised Skill Discovery for Agentic Data Analysis

AgentsDGX agent

arXiv:2606.06416v1 Announce Type: cross Abstract: Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updati

4 Jun 2026

AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning

HardwareDGX agent

arXiv:2606.04484v1 Announce Type: new Abstract: We present AgentJet, a distributed swarm training framework for large language model (LLM) agent reinforcement learning. Unlike centralized frameworks t

From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents

AgentsDGX agent

arXiv:2606.04967v1 Announce Type: cross Abstract: AI tools for programming are no longer just autocomplete or chat assistants: they organize themselves as development frameworks, with process, roles,

Radiant Logic extends identity visibility platform to enterprise AI agents with real-time risk scoring

AgentsDGX agent

Radiant Logic Inc., a platform that provides identity visibility and intelligence, today announced it’s extending its services to agentic artificial intelligence to help companies control and govern t

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

SafetyDGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

Streaming Communication in Multi-Agent Reasoning

Model ReleasesDGX agent

arXiv:2606.05158v1 Announce Type: cross Abstract: Multi-agent reasoning systems adopt a 'generate-then-transfer' paradigm that forces end-to-end latency to scale linearly with pipeline depth. We intro

We have a Fleet agent called @docs_plz in our Slack that's made a very noticeable impact on our velocity of docs changes. In the chart below…

AgentsDGX agent

We have a Fleet agent called @docs_plz in our Slack that's made a very noticeable impact on our velocity of docs changes. In the chart below, you can clearly see that after it was added, the amount of

3 Jun 2026

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Model ReleasesDGX agent

arXiv:2602.12430v4 Announce Type: replace-cross Abstract: The transition from monolithic language models to modular, skill-equipped agents marks a defining shift in how large language models (LLMs) ar

Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization

AgentsDGX agent

arXiv:2604.17708v2 Announce Type: replace Abstract: Automating operations research (OR) with large language models (LLMs) remains limited by hand-crafted reasoning--execution workflows. Complex OR tas

FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration

AgentsDGX agent

arXiv:2512.11213v2 Announce Type: replace Abstract: Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training. However, e

Making agent memory more reliable, transparent, and production-ready

AgentsDGX agent

Memory has always mattered for personalization and continuity. But as customers move agents from demos into production, another requirement becomes just as important: reliability. Enterprise teams nee

Uncertainty-Aware Clarification in LLM Agents with Information Gain

AgentsDGX agent

arXiv:2606.03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous too

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

SafetyDGX agent

arXiv:2602.08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving com

2 Jun 2026

Acting with AI: An Interaction-Based Framework for Agentic Tort Liability

AgentsDGX agent

arXiv:2606.00518v1 Announce Type: new Abstract: Agentic AI systems can plan over multiple steps, use tools, and execute tasks over time. When such systems cause harm, tort law struggles to allocate re

Archestra raises $10M to broker AI agent access to corporate data

AgentsDGX agent

Archestra Inc., a U.K.-based startup whose open-source platform brokers access between artificial intelligence agents and sensitive enterprise data, today announced that it has raised 10 million in ne

Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows

AgentsDGX agent

arXiv:2602.14849v2 Announce Type: replace-cross Abstract: LLM agents execute multi-step workflows that mutate external state through tools. Common orchestrators treat tool return as the settlement tri

Beyond One-shot: AI Agents for Learning in Field Experiments

AgentsDGX agent

arXiv:2606.02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention d

Can LLM Agents Sustain Long-Horizon Organizational Dynamics?

AgentsDGX agent

arXiv:2606.01199v1 Announce Type: new Abstract: Large language agents are increasingly used for social simulation, yet it remains unclear whether they can sustain coherent behavior in structured organ

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

AgentsDGX agent

arXiv:2601.19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows tha

'Do Not Mention This to the User': Detecting and Understanding Malicious Agent Skills

AgentsDGX agent

arXiv:2602.06547v3 Announce Type: replace-cross Abstract: LLM-based coding agents increasingly rely on third-party extensions called skills, which bundle natural language instructions and helper scrip

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

AgentsDGX agent

arXiv:2606.01365v1 Announce Type: new Abstract: Tool-using multi-agent large language model (LLM) systems spend computation through model tokens, tool calls, retries, and code execution before produci

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Model ReleasesDGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

AgentsDGX agent

arXiv:2606.02483v1 Announce Type: cross Abstract: Tool-augmented language agents speculatively issue likely future tool calls to hide latency, but those calls leak inferred user intent to external ser

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems

SafetyDGX agent

arXiv:2606.01779v1 Announce Type: new Abstract: LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent

Iteris: Agentic Research Loops for Computational Mathematics

AgentsDGX agent

arXiv:2606.02484v1 Announce Type: new Abstract: Recent advances in large language models and agentic AI systems have enabled significant progress in mathematical discovery, from solving competition pr

MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems

SafetyDGX agent

arXiv:2601.14230v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) are emerging as promising socio-collaborative companions for emotional and cognitive support. However, existing syst

MemPro: Agentic Memory Systems as Evolvable Programs

AgentsDGX agent

arXiv:2606.00619v1 Announce Type: cross Abstract: Long-horizon autonomous agents require memory systems to retain historical information, track evolving states, and reuse relevant knowledge beyond fin

Microsoft unveils Project Solara, an Android-based platform for agent-first devices, with concept hardware and pilots planned at Best Buy, Target, and others (Todd Bishop/GeekWire)

AgentsDGX agent

Todd Bishop / GeekWire: Microsoft unveils Project Solara, an Android-based platform for agent-first devices, with concept hardware and pilots planned at Best Buy, Target, and others — [Editor's Note:

Monitoring Agentic Systems Before They're Reliable

AgentsDGX agent

arXiv:2606.02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the

Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design

Model ReleasesDGX agent

arXiv:2606.00555v1 Announce Type: new Abstract: Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two

Scaling Agentic Capabilities via Grounded Interaction Synthesis

AgentsDGX agent

arXiv:2606.02001v1 Announce Type: new Abstract: General agentic intelligence hinges on the ability to interact with diverse real-world tools to complete complex tasks, a capability fundamentally tied

Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration

AgentsDGX agent

arXiv:2410.02511v2 Announce Type: replace Abstract: With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing

1 Jun 2026

a parallel experiment building a coding agent on top of @activegraphai. you can see everything flattened down to a single event log trace

AgentsDGX agent

This post documents a parallel experiment where Yohei Nakajima built a coding agent using ActiveGraphAI, showcasing the system's architecture through a flattened event log trace that makes all operati

← Previous
1…5758596061…297
Next →