AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
3 Jun 2026

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

AgentsDGX agent

arXiv:2606.03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little

DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees

AgentsDGX agent

arXiv:2606.03083v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly rely on memory to learn from experiences over continual interactions. However, storing experiences

InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain

Local AiDGX agent

arXiv:2606.03329v1 Announce Type: new Abstract: Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

@jgreze will speak on this at https://ai.engineer/wf gathering all the top agent labs. lfg

AgentsDGX agent

Swyx announces that @jgreze will speak at the ai.engineer/wf gathering, which is positioned as an event bringing together leading agent research labs. The post expresses enthusiasm for the upcoming sp

Just pushed an update to help remote connecting with the Hermes Agent GUI over tailscale to function! Please update if you had any issues!

AgentsDGX agent

Nous Research released an update improving remote connectivity functionality for the Hermes Agent GUI when used over Tailscale, addressing previous compatibility issues. Users experiencing connection

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.03203v1 Announce Type: new Abstract: Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unv

PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search

AgentsDGX agent

arXiv:2606.03099v1 Announce Type: cross Abstract: Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-bas

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

Model ReleasesDGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

Model ReleasesDGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

SafetyDGX agent

arXiv:2606.03137v1 Announce Type: new Abstract: LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existi

Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

Local AiDGX agent

arXiv:2606.02862v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive

Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots…

AgentsDGX agent

Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots have now passed human traffic online for the first time in

What Makes Interaction Trajectories Effective for Training Terminal Agents?

Model ReleasesDGX agent

arXiv:2606.03461v1 Announce Type: new Abstract: Stronger code agents are commonly assumed to be superior teachers for post-training, yet this assumption remains poorly disentangled from task difficult

2 Jun 2026

A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems

AgentsDGX agent

arXiv:2606.00138v1 Announce Type: new Abstract: Finite element analysis (FEA) is the most important numerical approach for solid mechanics. Challenges of FEA include a steep learning curve for entry-l

ACON: Optimizing Context Compression for Long-horizon LLM Agents

Model ReleasesDGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

SafetyDGX agent

arXiv:2606.00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queu

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Model ReleasesDGX agent

arXiv:2606.02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future e

AMP: A Vendor-Neutral Wire Format for Agent Memory Operations

AgentsDGX agent

arXiv:2606.01138v1 Announce Type: cross Abstract: Agent-memory frameworks - mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor - each ship their own SDK, storage layout, and operational voc

ASE-26: a curriculum for agentic software engineering as a discipline

Model ReleasesDGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime. This extensibility also creates a supp

Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation

Model ReleasesDGX agent

arXiv:2606.01725v1 Announce Type: new Abstract: Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes. Despite its popularity, its system-level beha

🎉 Congratulations to @OdessiaTravel on their public launch! Odessia is an AI-powered travel agent that allows users to plan and book entire…

AgentsDGX agent

🎉 Congratulations to @OdessiaTravel on their public launch! Odessia is an AI-powered travel agent that allows users to plan and book entire trips in one conversation. The team used LangSmith and LangG

Coordinating Task Switching in a Robotics Multi-Agent System Using Behavior Trees

AgentsDGX agent

arXiv:2606.01170v1 Announce Type: cross Abstract: The application of multi-agent systems in robotics is a very challenging field. Several competitions involving such systems are proposed to foster res

CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation

Model ReleasesDGX agent

arXiv:2606.01815v1 Announce Type: new Abstract: Evaluating LLM agents in realistic service scenarios requires complex task dependencies, imperfect user behavior, and an evaluation that accommodates mu

datasette-agent-micropython 0.1a0

Model ReleasesDGX agent

Release: datasette-agent-micropython 0.1a0 I want Datasette Agent to be able to generate and execute Python code safely. This alpha is looking promising so far. GPT-5.5 has so far failed to break out

GitHub's plan for Agents — Kyle Daigle, GitHub

AgentsDGX agent

GitHub is developing AI agents to automate software development workflows and enhance developer productivity. The discussion likely covers GitHub's strategy for integrating autonomous AI capabilities

🚀 Go from prototype to production in one click. We added a new Deploy button to LangSmith Studio, so you can deploy your agent directly to …

AgentsDGX agent

LangSmith Studio added a new Deploy button feature that enables users to deploy agents directly to production with a single click, streamlining the workflow from prototype development to production de

Grok Build is genuinely amazing right now. It is not just another coding assistant. It is a full agentic system that can plan, write, refact…

AgentsDGX agent

Grok Build is genuinely amazing right now. It is not just another coding assistant. It is a full agentic system that can plan, write, refactor, debug, and build complete projects autonomously from a s

How Baz improved its AI Agent Code Review accuracy using Amazon Bedrock AgentCore

AgentsDGX agent

This post walks through how Baz built their Spec Review agent using Amazon Bedrock and Amazon Bedrock AgentCore. We'll cover the architecture decisions, implementation details, and the business outcom

'I Strongly Suspect This Website Is a Scam': Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

Model ReleasesDGX agent

arXiv:2606.00497v1 Announce Type: cross Abstract: Deceptive web content, widely instantiated across the internet and commonly known as extit{social-engineering attacks}, manipulates autonomous web age

II-Agent is now live on the App Store. Your sovereign AI workspace for building, researching, writing, designing, and automating from one in…

AgentsDGX agent

II-Agent is now live on the App Store. Your sovereign AI workspace for building, researching, writing, designing, and automating from one intelligent interface. Download it. Bring your own key. Build

Joint Agent Memory and Exploration Learning via Novelty Signals

SafetyDGX agent

arXiv:2606.01528v1 Announce Type: new Abstract: In open-ended environments, exploration is fundamental for autonomous agents, yet current language model agents struggle with this. Effective exploratio

MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2606.00610v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become an essential method for mitigating hallucinations in Large Language Models (LLMs) by leveraging extern

Multi-Agent Conformal Prediction with Personalized Statistical Validity

AgentsDGX agent

arXiv:2606.00717v1 Announce Type: cross Abstract: Uncertainty quantification is essential in high-stakes machine learning tasks. However, one of the principled solutions, conformal prediction, faces c

OctoT2I: A Self-Evolving Agentic Text-to-Image Router

AgentsDGX agent

arXiv:2606.01803v1 Announce Type: new Abstract: The explosive growth of Text-to-Image (T2I) models, from large-scale versions to lightweight, real-time ones, now faces diminishing marginal returns fro

Prompting agents to 'do better' is unreliable 🙅 Giving them a rubric, a grader, and a correction loop is much closer to how you get your ag…

Model ReleasesDGX agent

Prompting agents to 'do better' is unreliable 🙅 Giving them a rubric, a grader, and a correction loop is much closer to how you get your agent do what you want! Similar to /goal in Claude Code or othe

Read more about hybrid agentic inference in Perplexity Computer: https://www.perplexity.ai/hub/blog/the-data-center-moves-to-your-machine

AgentsDGX agent

Perplexity discusses hybrid agentic inference in Perplexity Computer, which likely describes a computational approach that combines processing between local machines and data centers. The concept sugg

RocketSmith: An Agentic System for High-Powered Rocket Design and Manufacturing

AgentsDGX agent

arXiv:2606.00097v1 Announce Type: new Abstract: This work presents RocketSmith, an agentic system capable of the design, manufacturing, and optimization processes in high powered rocket development. T

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

AgentsDGX agent

arXiv:2602.12984v2 Announce Type: replace Abstract: Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely ov

SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems

AgentsDGX agent

arXiv:2606.01314v1 Announce Type: new Abstract: Recent self-evolving agents have shown that skills can be discovered, refined, and accumulated through execution. However, existing skill-evolution fram

The next evolution of Hermes Agent is here! Introducing Hermes Desktop: everything you love about Hermes, now native on your machine. First …

AgentsDGX agent

The next evolution of Hermes Agent is here! Introducing Hermes Desktop: everything you love about Hermes, now native on your machine. First demoed in Jensen's GTC keynote, it's now in public preview.

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

1 Jun 2026

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

AgentsDGX agent

arXiv:2605.31468v1 Announce Type: new Abstract: Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Model ReleasesDGX agent

arXiv:2605.31170v1 Announce Type: cross Abstract: Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages

Enhancing Human-Likeness in Reinforcement Learning Agents via Hierarchical Macro Action Quantization

ResearchDGX agent

arXiv:2605.30928v1 Announce Type: new Abstract: Human-like agents are a long-standing goal of artificial intelligence. Despite strong performance, most reinforcement learning (RL) agents remain reward

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

Model ReleasesDGX agent

arXiv:2605.30947v1 Announce Type: new Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantit

MiniMax M3 imminent. Will be doing deep testing with it on my own coding agent and harness. Review coming soon.

AgentsDGX agent

MiniMax M3, an upcoming AI model, is expected to be released soon and will undergo comprehensive testing within a custom coding agent framework. A detailed technical review of the model's performance

.@MukilLoganathan’s Interrupt keynote on Sandboxes. https://youtu.be/IIchUA5T3gs In 20 minutes, you’ll learn how to run agent code safely. I…

AgentsDGX agent

.@MukilLoganathan’s Interrupt keynote on Sandboxes. https://youtu.be/IIchUA5T3gs In 20 minutes, you’ll learn how to run agent code safely. Isolated from your runtime, with network controls, persistent

NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents

AgentsDGX agent

arXiv:2601.21372v2 Announce Type: replace Abstract: We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimization implem

PithTrain: A Compact and Agent-Native MoE Training System

HardwareDGX agent

arXiv:2605.31463v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built opti

.@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. https://www.langchain.com/blog/ho…

AgentsDGX agent

.@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. https://www.langchain.com/blog/how-rippling-went-ai-native-across-every-product-in-6-months-w

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

SafetyDGX agent

arXiv:2605.30698v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spo

Sources: Tencent, which has fallen behind domestic rivals in AI models, plans to test an AI agent for WeChat with a small group of users before a phased rollout (Zijing Wu/Financial Times)

AgentsDGX agent

Zijing Wu / Financial Times: Sources: Tencent, which has fallen behind domestic rivals in AI models, plans to test an AI agent for WeChat with a small group of users before a phased rollout — Maker of

We're trending on @huggingface! 🥳 Tbh, we undersold this model. It's a lot more capable at agentic tasks than I expected. I keep discoverin…

AgentsDGX agent

We're trending on @huggingface! 🥳 Tbh, we undersold this model. It's a lot more capable at agentic tasks than I expected. I keep discovering new capabilities every day, it's crazy for 1B active parame

30 May 2026

Found a way to save everyone 14% on input tokens on average during read file operations in Hermes Agent! This is now on main. `hermes update…

AgentsDGX agent

Nous Research has optimized token efficiency in their Hermes Agent, achieving an average 14% reduction in input token usage during file read operations. This optimization has been merged to the main c

29 May 2026

Agent4Edu: Generating Learner Response Data by Generative Agents for Intelligent Education Systems

AgentsDGX agent

arXiv:2501.10332v2 Announce Type: replace-cross Abstract: Personalized learning represents a promising educational strategy within intelligent educational systems, aiming to enhance learners' practice

AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning

Model ReleasesDGX agent

arXiv:2602.23258v2 Announce Type: replace Abstract: While Multi-Agent Systems (MAS) excel in complex reasoning, they suffer from the cascading impact of erroneous information from individual agents. C

AnomalyAgent: Training-Free Agentic Models for Zero-/Few-Shot Anomaly Detection

AgentsDGX agent

arXiv:2605.30140v1 Announce Type: new Abstract: Benefiting from generalizability of vision-language models (VLMs) such as CLIP, many zero-/few-shot anomaly detection (AD) approaches have achieved impr

Cloud CISO Perspectives: How to build an AI-ready security program for the public sector

Model ReleasesDGX agent

Welcome to the second Cloud CISO Perspectives for May 2026. Today, Usman Chaudhary, Field CISO, Google Public Sector, offers a guide for CISOs protecting government agencies and critical infrastructur

← Previous
1…8283848586…300
Next →