AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,911 results
30 Jun 2026

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

Model ReleasesDGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

harbor is a great framework for running evals for long running, stateful agents its becoming industry standard, powering benchmarks like ter…

AgentsDGX agent

harbor is a great framework for running evals for long running, stateful agents its becoming industry standard, powering benchmarks like terminal bench 2 we've integrated deeply with harbor: across La

Hierarchical Experimentalist Agents

Model ReleasesDGX agent

arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametr

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning

AgentsDGX agent

arXiv:2606.28349v1 Announce Type: cross Abstract: Long-context reasoning requires models to access, retrieve, and integrate evidence scattered across documents, dialogues, and accumulated interaction

how do you run agent code without a full blown sandbox? we did a lot of work to harden the code interpreter runtime we use

AgentsDGX agent

Harrison Chase discusses techniques for executing agent code safely without requiring a complete sandbox environment, highlighting hardening measures implemented in their code interpreter runtime. The

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning

Local AiDGX agent

arXiv:2606.29425v1 Announce Type: new Abstract: Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordination patter

Monte Carlo Query Search: Active Capability Assessment of AI Agents

AgentsDGX agent

arXiv:2512.16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires metho

More on Hermes Agent web search & extraction: https://hermes-agent.nousresearch.com/docs/user-guide/features/web-search

AgentsDGX agent

This documentation page covers advanced features of Hermes Agent's web search and information extraction capabilities, explaining how users can leverage the tool to search the internet and extract rel

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…

Model ReleasesDGX agent

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI agents. LLMs suffer from all sorts of reward hacking issues. T

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Model ReleasesDGX agent

arXiv:2603.29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visuali

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

Model ReleasesDGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for t…

Model ReleasesDGX agent

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for the naturalness/latency Building voice agents can come with t

29 Jun 2026

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

AgentsDGX agent

arXiv:2606.13544v3 Announce Type: replace-cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor compe

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

AgentsDGX agent

arXiv:2606.27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by t

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

SafetyDGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

Exclusive: Agentic coding startup Baz brings code reviews to the planning stage as it extends seed funding to $17M

AgentsDGX agent

Agentic coding startup Baz Technologies Inc. said today it’s launching a new platform that sits between developers and the code bases they’re working on in order to catch software vulnerabilities befo

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

AgentsDGX agent

arXiv:2606.27974v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fi

Things are progressing... Purely made with Hermes Agent > MCP > Unreal Here is some quick walk around footage of a scene I created purely th…

AgentsDGX agent

Things are progressing... Purely made with Hermes Agent > MCP > Unreal Here is some quick walk around footage of a scene I created purely through guidance/prompting . Media Time to get Unreal with Her

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to c…

AgentsDGX agent

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to choose your winners for the Round of 32 by 9:00 am PT Tuesday

Training Observable Control Policies to Expose Agent State Through Actions

SafetyDGX agent

arXiv:2606.27609v1 Announce Type: new Abstract: Physical or operational constraints often impose communications limitations on autonomous agents. Such limitations complicate monitoring or multiagent c

27 Jun 2026

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, R…

AgentsDGX agent

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, Research Lead, RL at MiniMax, and @realDanFu, VP of Kernels a

26 Jun 2026

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

AgentsDGX agent

arXiv:2606.27251v1 Announce Type: cross Abstract: Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT)

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

SafetyDGX agent

arXiv:2606.26122v1 Announce Type: new Abstract: Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

Model ReleasesDGX agent

arXiv:2606.26356v1 Announce Type: new Abstract: Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despi

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

AgentsDGX agent

arXiv:2606.26793v1 Announce Type: cross Abstract: Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image inj

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

AgentsDGX agent

arXiv:2606.26669v1 Announce Type: new Abstract: Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explo

25 Jun 2026

AI SDK 7 is here. This release sets the foundation for agents and AI platforms in production: approvals, durability, telemetry, and more.

AgentsDGX agent

AI SDK 7 is here. This release sets the foundation for agents and AI platforms in production: approvals, durability, telemetry, and more. AI SDK 7 is now available. Introducing: reasoning control, age

Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks

AgentsDGX agent

Explore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maintaining flexibility to choose among more than 20 models. The p

Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory

Local AiDGX agent

arXiv:2606.25115v1 Announce Type: new Abstract: On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights. This memory is hard-bounded and

Neurosymbolic agents (here Codex) are absolutely crushing pure chatbots.

AgentsDGX agent

Neurosymbolic agents (here Codex) are absolutely crushing pure chatbots. This is a fascinating and important set of data which shows us where things are going, using OpenAI as a canary in the coal min

PhoneBuddy: Training Open Models for Agentic Phone Use

AgentsDGX agent

arXiv:2606.23049v2 Announce Type: replace Abstract: Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult bec

Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

SafetyDGX agent

arXiv:2606.25127v1 Announce Type: new Abstract: We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Per

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

SafetyDGX agent

arXiv:2606.20615v2 Announce Type: replace Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG

AgentsDGX agent

arXiv:2606.25191v1 Announce Type: cross Abstract: Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable mode

Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is mor…

AgentsDGX agent

Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is more complex, longer-running, and increasingly cross-functional

24 Jun 2026

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on worksh…

AgentsDGX agent

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on workshop on how inference engines work and what it takes to serve

AI agents are changing work — and Dell’s John Roese says it’s just beginning

AgentsDGX agent

To gain a better understanding of the longer-term impact that autonomous agents will have on the nature of work, Dell Technologies Inc. has been taking a closer look at how AI is already changing how

Are We Ready For An Agent-Native Memory System?

Model ReleasesDGX agent

arXiv:2606.24775v1 Announce Type: new Abstract: Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a data management system that supports

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

Model ReleasesDGX agent

arXiv:2606.24429v1 Announce Type: cross Abstract: Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly unde

Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification

AgentsDGX agent

arXiv:2606.23764v1 Announce Type: cross Abstract: Fei Xiaotong's Differential Order Pattern characterizes rural society as egocentric and relationally graded, with cooperation attenuating over social

Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures

AgentsDGX agent

A field guide to the new wave of long-horizon agent benchmarks: what each one actually measures, the realism-versus-verifiability bargain it strikes, and the seam where its score leaks. The post Long-

Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

AgentsDGX agent

arXiv:2606.24783v1 Announce Type: cross Abstract: Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We

Red-Teaming the Agentic Red-Team

SafetyDGX agent

arXiv:2606.24496v1 Announce Type: cross Abstract: The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, wh

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Model ReleasesDGX agent

arXiv:2606.24112v1 Announce Type: new Abstract: Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed proven

23 Jun 2026

AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents

AgentsDGX agent

arXiv:2606.20625v1 Announce Type: cross Abstract: LLM agents are promising for alpha mining via combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refine

Causal Discovery in the Era of Agents

AgentsDGX agent

arXiv:2606.23608v1 Announce Type: cross Abstract: Recent attempts to combine large language models (LLMs) with causal discovery ask models to infer pairwise directions, propose graph structures, or in

Darwin Mobile Agent: A Roadmap for Self-Evolution

SafetyDGX agent

arXiv:2606.20622v1 Announce Type: cross Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments. Guided by the 'Bitter Lesso

Distilling Collaborative Dynamics into Latent Space for Implicit Coordination in Decentralized Multi-Agent Manipulation

Model ReleasesDGX agent

arXiv:2606.22982v1 Announce Type: new Abstract: Multi-arm manipulation demands precise spatiotemporal coordination, yet many centralized approaches scale poorly as team size increases. To address this

Dynamic multi-agent deep reinforcement learning-based pricing and incentivization approach in multimodal transportation networks

AgentsDGX agent

arXiv:2606.23257v1 Announce Type: new Abstract: In multimodal transportation systems, shared mobility services (SMSs) are promoted for their potential to enhance flexibility and reduce congestion. How

Exabeam launches Praxen, an open-source tool to verify AI agent behavior

Model ReleasesDGX agent

Security intelligence and management solutions company Exabeam Inc. today introduced Agent Behavior Verification, a pre-deployment security discipline for artificial intelligence agents, and released

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.22995v1 Announce Type: new Abstract: Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy up

Highly-recommended read. It's exciting to see large-scale agentic RL becoming more accessible. Cool to see the infra layer for this is being…

AgentsDGX agent

Highly-recommended read. It's exciting to see large-scale agentic RL becoming more accessible. Cool to see the infra layer for this is being built and I think this plays an important role in self-impr

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

Model ReleasesDGX agent

arXiv:2602.09012v2 Announce Type: replace Abstract: The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a

Okta expands Cross App Access ecosystem to secure AI agent connections

AgentsDGX agent

Identity and access management company Okta Inc. today said more than 25 software makers have signed on to its Cross App Access framework, which routes the connections artificial intelligence agents m

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

Model ReleasesDGX agent

arXiv:2606.20512v2 Announce Type: replace-cross Abstract: LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test su

Skill Coverage: A Test Adequacy Metric for Agent Skills

Model ReleasesDGX agent

arXiv:2606.20659v1 Announce Type: cross Abstract: Agent skills encode reusable procedural knowledge that guides large language model agents across tasks and execution contexts. Existing evaluations pr

Towards Adaptive Categories: Dimensional Governance for Agentic AI

AgentsDGX agent

arXiv:2505.11579v3 Announce Type: replace-cross Abstract: As AI systems evolve from static tools to dynamic agents, traditional categorical governance frameworks -- based on fixed risk tiers, levels o

22 Jun 2026

Just a glimpse of what collective AI intelligence will bring. We haven’t truly cracked multi-agent orchestration but with every new frontier…

AgentsDGX agent

Just a glimpse of what collective AI intelligence will bring. We haven’t truly cracked multi-agent orchestration but with every new frontier model, intelligence should compound. Introducing Sakana Fug

The new /goal command in Grok Build is a huge update Until now, most coding agents have worked like enhanced chatbots You ask. It responds. …

AgentsDGX agent

The new /goal command in Grok Build is a huge update Until now, most coding agents have worked like enhanced chatbots You ask. It responds. You review. You guide it. Repeat /goal changes the entire pa

Use Case 2: Financial Time Series Prediction Can an AI agent navigate sequential, no-look-ahead market decisions? Just for fun, we tested Fu…

AgentsDGX agent

Use Case 2: Financial Time Series Prediction Can an AI agent navigate sequential, no-look-ahead market decisions? Just for fun, we tested Fugu Ultra on 50 weeks of historical data for an anonymized eq

← Previous
1…6768697071…299
Next →