AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,919 results
Model Releases

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

DGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

model-releasesarxiv-cs-ai
30 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

DGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

harbor is a great framework for running evals for long running, stateful agents its becoming industry standard, powering benchmarks like ter…

DGX agent

harbor is a great framework for running evals for long running, stateful agents its becoming industry standard, powering benchmarks like terminal bench 2 we've integrated deeply with harbor: across La

agentsharrison-chase--x
30 Jun 2026
Model Releases

Hierarchical Experimentalist Agents

DGX agent

arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametr

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning

DGX agent

arXiv:2606.28349v1 Announce Type: cross Abstract: Long-context reasoning requires models to access, retrieve, and integrate evidence scattered across documents, dialogues, and accumulated interaction

agentsarxiv-cs-ai
30 Jun 2026
Agents

how do you run agent code without a full blown sandbox? we did a lot of work to harden the code interpreter runtime we use

DGX agent

Harrison Chase discusses techniques for executing agent code safely without requiring a complete sandbox environment, highlighting hardening measures implemented in their code interpreter runtime. The

agentsharrison-chase--x
30 Jun 2026
Local Ai

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning

DGX agent

arXiv:2606.29425v1 Announce Type: new Abstract: Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordination patter

local-aiarxiv-cs-ai
30 Jun 2026
Agents

Monte Carlo Query Search: Active Capability Assessment of AI Agents

DGX agent

arXiv:2512.16733v3 Announce Type: replace Abstract: Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requires metho

agentsarxiv-cs-ai
30 Jun 2026
Agents

More on Hermes Agent web search & extraction: https://hermes-agent.nousresearch.com/docs/user-guide/features/web-search

DGX agent

This documentation page covers advanced features of Hermes Agent's web search and information extraction capabilities, explaining how users can leverage the tool to search the internet and extract rel

agentsnous-research--x
30 Jun 2026
Model Releases

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…

DGX agent

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI agents. LLMs suffer from all sorts of reward hacking issues. T

model-releasesdair-ai--x
30 Jun 2026
Model Releases

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

DGX agent

arXiv:2603.29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visuali

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

DGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for t…

DGX agent

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for the naturalness/latency Building voice agents can come with t

model-releasesharrison-chase--x
30 Jun 2026
Agents

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

DGX agent

arXiv:2606.13544v3 Announce Type: replace-cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor compe

agentsarxiv-cs-ai
29 Jun 2026
Agents

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

DGX agent

arXiv:2606.27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by t

agentsarxiv-cs-ai
29 Jun 2026
Safety

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

DGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

safetyarxiv-cs-cl
29 Jun 2026
Agents

Exclusive: Agentic coding startup Baz brings code reviews to the planning stage as it extends seed funding to $17M

DGX agent

Agentic coding startup Baz Technologies Inc. said today it’s launching a new platform that sits between developers and the code bases they’re working on in order to catch software vulnerabilities befo

agentssiliconangle
29 Jun 2026
Agents

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

DGX agent

arXiv:2606.27974v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fi

agentsarxiv-cs-ai
29 Jun 2026
Agents

Things are progressing... Purely made with Hermes Agent > MCP > Unreal Here is some quick walk around footage of a scene I created purely th…

DGX agent

Things are progressing... Purely made with Hermes Agent > MCP > Unreal Here is some quick walk around footage of a scene I created purely through guidance/prompting . Media Time to get Unreal with Her

agentsnous-research--x
29 Jun 2026
Agents

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to c…

DGX agent

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to choose your winners for the Round of 32 by 9:00 am PT Tuesday

agentstogether-ai--x
29 Jun 2026
Safety

Training Observable Control Policies to Expose Agent State Through Actions

DGX agent

arXiv:2606.27609v1 Announce Type: new Abstract: Physical or operational constraints often impose communications limitations on autonomous agents. Such limitations complicate monitoring or multiagent c

safetyarxiv-cs-lg
29 Jun 2026
Agents

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, R…

DGX agent

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, Research Lead, RL at MiniMax, and @realDanFu, VP of Kernels a

agentstogether-ai--x
27 Jun 2026
Agents

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

DGX agent

arXiv:2606.27251v1 Announce Type: cross Abstract: Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT)

agentsarxiv-cs-ai
26 Jun 2026
Safety

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

DGX agent

arXiv:2606.26122v1 Announce Type: new Abstract: Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples

safetyarxiv-cs-cv
26 Jun 2026
Model Releases

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

DGX agent

arXiv:2606.26356v1 Announce Type: new Abstract: Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despi

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG

DGX agent

arXiv:2606.26793v1 Announce Type: cross Abstract: Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image inj

agentsarxiv-cs-ai
26 Jun 2026
Agents

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

DGX agent

arXiv:2606.26669v1 Announce Type: new Abstract: Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explo

agentsarxiv-cs-ai
26 Jun 2026
Agents

AI SDK 7 is here. This release sets the foundation for agents and AI platforms in production: approvals, durability, telemetry, and more.

DGX agent

AI SDK 7 is here. This release sets the foundation for agents and AI platforms in production: approvals, durability, telemetry, and more. AI SDK 7 is now available. Introducing: reasoning control, age

agentsvercel--x
25 Jun 2026
Agents

Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks

DGX agent

Explore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maintaining flexibility to choose among more than 20 models. The p

agentsgithub-ai-blog
25 Jun 2026
Local Ai

Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory

DGX agent

arXiv:2606.25115v1 Announce Type: new Abstract: On-device language-model agents improve by accumulating experience in retrieved memory rather than by updating weights. This memory is hard-bounded and

local-aiarxiv-cs-lg
25 Jun 2026
Agents

Neurosymbolic agents (here Codex) are absolutely crushing pure chatbots.

DGX agent

Neurosymbolic agents (here Codex) are absolutely crushing pure chatbots. This is a fascinating and important set of data which shows us where things are going, using OpenAI as a canary in the coal min

agentsgary-marcus--x
25 Jun 2026
Agents

PhoneBuddy: Training Open Models for Agentic Phone Use

DGX agent

arXiv:2606.23049v2 Announce Type: replace Abstract: Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult bec

agentsarxiv-cs-cl
25 Jun 2026
Safety

Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

DGX agent

arXiv:2606.25127v1 Announce Type: new Abstract: We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Per

safetyarxiv-cs-lg
25 Jun 2026
Safety

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

DGX agent

arXiv:2606.20615v2 Announce Type: replace Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing

safetyarxiv-cs-ai
25 Jun 2026
Agents

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG

DGX agent

arXiv:2606.25191v1 Announce Type: cross Abstract: Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable mode

agentsarxiv-cs-cl
25 Jun 2026
Agents

Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is mor…

DGX agent

Work at OpenAI is being transformed by agents, in every department. Across our entire company, people are using Codex to do work that is more complex, longer-running, and increasingly cross-functional

agentsopenai--x
25 Jun 2026
Agents

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on worksh…

DGX agent

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on workshop on how inference engines work and what it takes to serve

agentstogether-ai--x
24 Jun 2026
Agents

AI agents are changing work — and Dell’s John Roese says it’s just beginning

DGX agent

To gain a better understanding of the longer-term impact that autonomous agents will have on the nature of work, Dell Technologies Inc. has been taking a closer look at how AI is already changing how

agentssiliconangle
24 Jun 2026
Model Releases

Are We Ready For An Agent-Native Memory System?

DGX agent

arXiv:2606.24775v1 Announce Type: new Abstract: Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a data management system that supports

model-releasesarxiv-cs-cl
24 Jun 2026
Model Releases

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

DGX agent

arXiv:2606.24429v1 Announce Type: cross Abstract: Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly unde

model-releasesarxiv-cs-ai
24 Jun 2026
Agents

Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification

DGX agent

arXiv:2606.23764v1 Announce Type: cross Abstract: Fei Xiaotong's Differential Order Pattern characterizes rural society as egocentric and relationally graded, with cooperation attenuating over social

agentsarxiv-cs-ai
24 Jun 2026
Agents

Long-horizon agent benchmarks are fragmenting: a field guide to what each one actually measures

DGX agent

A field guide to the new wave of long-horizon agent benchmarks: what each one actually measures, the realism-versus-verifiability bargain it strikes, and the seam where its score leaks. The post Long-

agentsarize-ai
24 Jun 2026
Agents

Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

DGX agent

arXiv:2606.24783v1 Announce Type: cross Abstract: Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We

agentsarxiv-cs-ai
24 Jun 2026
Safety

Red-Teaming the Agentic Red-Team

DGX agent

arXiv:2606.24496v1 Announce Type: cross Abstract: The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, wh

safetyarxiv-cs-ai
24 Jun 2026
Model Releases

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

DGX agent

arXiv:2606.24112v1 Announce Type: new Abstract: Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed proven

model-releasesarxiv-cs-ai
24 Jun 2026
Agents

AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents

DGX agent

arXiv:2606.20625v1 Announce Type: cross Abstract: LLM agents are promising for alpha mining via combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refine

agentsarxiv-cs-lg
23 Jun 2026
Agents

Causal Discovery in the Era of Agents

DGX agent

arXiv:2606.23608v1 Announce Type: cross Abstract: Recent attempts to combine large language models (LLMs) with causal discovery ask models to infer pairwise directions, propose graph structures, or in

agentsarxiv-cs-lg
23 Jun 2026
Safety

Darwin Mobile Agent: A Roadmap for Self-Evolution

DGX agent

arXiv:2606.20622v1 Announce Type: cross Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments. Guided by the 'Bitter Lesso

safetyarxiv-cs-lg
23 Jun 2026
← Previous
1…8485868788…374
Next →