AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
1 Jul 2026

1/ DSGym: A Holistic Framework for Evaluating and Training Data Science Agents Paper: https://arxiv.org/abs/2601.16344

ToolsDGX agent

DSGym is a comprehensive framework designed to evaluate and train AI agents for data science tasks, providing a structured environment for benchmarking agent performance across various data science wo

Building a serverless A2A gateway for agent discovery, routing, and access control

AgentsDGX agent

In this post, you will learn how to build a serverless A2A gateway on AWS that hosts multiple agents behind a single domain using path-based routing (/agents/{agentId}). Standard A2A clients work with

Content Independence Day, one year on: building the business model for the agentic Internet

AgentsDGX agent

One year after declaring Content Independence Day, a dynamic market for monetized content has officially emerged. In this report, we examine how the rise of autonomous AI agents is upending traditiona

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Model ReleasesDGX agent

arXiv:2606.31522v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as 'preserve

Game. Match. That's a wrap on the inaugural Agent Open 🦙🌤️🏓 Hundreds of AI builders, 64 matches, 7 partners, 2 courts. No dink shots or d…

AgentsDGX agent

Game. Match. That's a wrap on the inaugural Agent Open 🦙🌤️🏓 Hundreds of AI builders, 64 matches, 7 partners, 2 courts. No dink shots or demos, just backhands and hot takes. Thanks to our partners @Bra

@hwchase17 Love the idea of using a wiki for everything for agentic work! Recently Google created an open format for this which helps a ton …

AgentsDGX agent

@hwchase17 Love the idea of using a wiki for everything for agentic work! Recently Google created an open format for this which helps a ton creating your own skills out of it - it’s called OKF: https:

If you want frontier-level coding and agent performance but you don't want to pay closed-model prices, GLM 5.2 is the open model you've prob…

AgentsDGX agent

If you want frontier-level coding and agent performance but you don't want to pay closed-model prices, GLM 5.2 is the open model you've probably been hearing about. Here's why you should run it on Fir

Introducing Computer in WorkPods. Computer lets your agents use a browser. Every task that lives on the web, behind a login, a form, a butto…

AgentsDGX agent

Introducing Computer in WorkPods. Computer lets your agents use a browser. Every task that lives on the web, behind a login, a form, a button, now runs inside the same session you're already watching.

LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents

ResearchDGX agent

arXiv:2606.30697v1 Announce Type: cross Abstract: Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, visual groupi

OpenWiki is the easiest way to document your codebase, built specifically for agents to consume. OpenWiki can generate repo docs, automatica…

AgentsDGX agent

OpenWiki is the easiest way to document your codebase, built specifically for agents to consume. OpenWiki can generate repo docs, automatically update them as your codebase evolves, and do Q&A over yo

Robust Autonomous UAV Landing on Maritime Platforms via Multimodal Agentic AI and Active Wave Compensation

AgentsDGX agent

arXiv:2606.31613v1 Announce Type: new Abstract: Autonomous aerial inspection of marine infrastructure is frequently compromised by stochastic sea states, introducing risks of high-kinetic impacts, pos

Super proud to say that the team and I put almost all our effort into resolving every P0 and P1 issue and PR in the entire Hermes Agent repo…

AgentsDGX agent

Super proud to say that the team and I put almost all our effort into resolving every P0 and P1 issue and PR in the entire Hermes Agent repo over the last week and a half, and as of 5 minutes ago, aft

The Agent Open was a galaxy brain idea that @murtazakhomusi and I did together — couldn’t be happier we went for it. See you at the next one…

AgentsDGX agent

Jerry Liu reflects positively on collaborating with Murtaza Khomusi on 'The Agent Open,' expressing satisfaction with the decision to pursue the project. The post suggests this was a successful initia

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States

Model ReleasesDGX agent

arXiv:2606.31612v1 Announce Type: new Abstract: Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Exi

You can now run recursive language model (RLM) workflows in Deep Agents. Everything you need to know in 6 minutes from @sydneyrunkle.

AgentsDGX agent

Recursive Language Model (RLM) workflows are now available as a feature in Deep Agents, allowing AI systems to iteratively call language models within multi-step processes. This capability enables mor

30 Jun 2026

A great read from @LangChain Best framing I've seen for agent architecture. Every non-model piece like the filesystem, sandbox, compaction, …

AgentsDGX agent

A great read from @LangChain Best framing I've seen for agent architecture. Every non-model piece like the filesystem, sandbox, compaction, Ralph loops exists to patch a specific model limitation. htt

Added this for everyone else through the /usage command! Run `/usage` and get this breakdown wherever you use Hermes Agent :)

AgentsDGX agent

Added this for everyone else through the /usage command! Run `/usage` and get this breakdown wherever you use Hermes Agent :) starting today you can now see a breakdown of your context usage per sessi

AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control

AgentsDGX agent

arXiv:2603.26005v2 Announce Type: replace Abstract: Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems. Realistic studi

Build generative UI for AI agents on Amazon Bedrock AgentCore with the AG-UI protocol

AgentsDGX agent

This post walks through how AG-UI integrates into the Fullstack AgentCore Solution Template (FAST) to build interactive agent frontends on Amazon Bedrock AgentCore. We then show how CopilotKit extends

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich archite…

Model ReleasesDGX agent

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich architecture How to build a voice research agent with both: ✅ Gemini

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

AgentsDGX agent

arXiv:2602.12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove bot

Collective cooperation without individual fidelity in LLM agents

Model ReleasesDGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

Complementary RL: Towards Efficient Experience-Driven Agent Learning

SafetyDGX agent

arXiv:2603.17621v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for training LLM-based agents, yet remains limited by low sample efficiency, st

Doors are open at Agent Open, right next door to AI Engineer Two courts, blue sky, no fog (a San Francisco miracle) 🌤️ This is the perfect …

AgentsDGX agent

Doors are open at Agent Open, right next door to AI Engineer Two courts, blue sky, no fog (a San Francisco miracle) 🌤️ This is the perfect walk-over window, we are just 3 minutes from Moscone. Don't m

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

Model ReleasesDGX agent

arXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail

Giving agents the ability to write code makes them dramatically more capable. It also makes security a lot harder. At LangChain, we have spe…

AgentsDGX agent

Giving agents the ability to write code makes them dramatically more capable. It also makes security a lot harder. At LangChain, we have spent a lot of time this year figuring out how to do both. http

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2606.22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show t

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

Model ReleasesDGX agent

arXiv:2606.28514v1 Announce Type: new Abstract: Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these m

Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents

AgentsDGX agent

arXiv:2606.29459v1 Announce Type: cross Abstract: Inverse design of metal-organic frameworks (MOFs) requires searching a combinatorially vast space where property labels are expensive and most machine

I've added video support to my 'shot-scraper' browser automation tool - you (or your coding agent) can now create a storyboard YAML file and…

AgentsDGX agent

I've added video support to my 'shot-scraper' browser automation tool - you (or your coding agent) can now create a storyboard YAML file and use that to record a video demo of new web application feat

KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Search

AgentsDGX agent

arXiv:2606.29863v1 Announce Type: new Abstract: Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward spars

Memory has somehow consistently been the most exciting area of agent development over the last 3 years (imo), and it's still a largely unsol…

AgentsDGX agent

Memory has somehow consistently been the most exciting area of agent development over the last 3 years (imo), and it's still a largely unsolved problem!! Wiki's are the biggest advancement I've seen i

MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.30602v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agent communication channels introduc

ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

AgentsDGX agent

arXiv:2606.28357v1 Announce Type: cross Abstract: Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reasoning and s

Row-Bot's approach to memory is hybrid. We have taken the best of both worlds. A knowledge graph for the agent + bi-directionaly synced Wiki…

AgentsDGX agent

Row-Bot's approach to memory is hybrid. We have taken the best of both worlds. A knowledge graph for the agent + bi-directionaly synced Wiki for human readability/audit. With in-app graph visualizatio

SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents

AgentsDGX agent

arXiv:2606.28356v1 Announce Type: cross Abstract: Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recommendation age

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Model ReleasesDGX agent

arXiv:2606.29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the impr…

Model ReleasesDGX agent

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the improved version that can complete agentic tasks more reliably.

Today’s the day. The inaugural Agent Open is here and it’s gonna be awesome. Arrive early, play some pickleball, meet some people, eat good …

AgentsDGX agent

Today’s the day. The inaugural Agent Open is here and it’s gonna be awesome. Arrive early, play some pickleball, meet some people, eat good food, and enjoy the panel featuring @travers00 @ankrgyl @Sir

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

Model ReleasesDGX agent

arXiv:2606.28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader ra

We are rebuilding Hydrogen from the ground up with @Shopify. It's agent-first, runtime-agnostic, and runs anywhere JavaScript does. Learn mo…

AgentsDGX agent

We are rebuilding Hydrogen from the ground up with @Shopify. It's agent-first, runtime-agnostic, and runs anywhere JavaScript does. Learn more and try the Next.js developer preview ↓ https://vercel.co

29 Jun 2026

Build your first voice agent ↓ https://vercel.com/blog/realtime-voice-agents-on-ai-gateway

AgentsDGX agent

This guide demonstrates how to build a voice agent using Vercel's AI Gateway, enabling real-time voice interactions with AI models. It covers the technical setup and implementation steps needed to cre

COOPA: A Modular LLM Agent Architecture for Operations Research Problems

AgentsDGX agent

arXiv:2606.27611v1 Announce Type: new Abstract: Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge,

Deep Agents supports sandboxes from any provider; with first-class support for E2B, Daytona, LangSmith, Modal, Vercel, Runloop, AgentCore, a…

AgentsDGX agent

Deep Agents supports sandboxes from any provider; with first-class support for E2B, Daytona, LangSmith, Modal, Vercel, Runloop, AgentCore, and BYOS (bring-your-own-sandbox implementation) support http

DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection

Model ReleasesDGX agent

arXiv:2606.27499v1 Announce Type: cross Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when a

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

SafetyDGX agent

arXiv:2606.27483v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-hor

SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game

Model ReleasesDGX agent

arXiv:2606.27397v1 Announce Type: cross Abstract: Evaluating LLM agents requires dynamic environments that go beyond static reasoning and zero-sum games. Real-world economic interaction is often open-

Stay in the loop with Live Activities, and get notified when an agent finishes or needs your input. Review demos and diffs before merging PR…

AgentsDGX agent

Cursor's Live Activities feature enables real-time notifications to keep users informed when AI agents complete tasks or require user input during development workflows. The feature supports reviewing

ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2606.28061v1 Announce Type: cross Abstract: Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments

We are starting to role out our Trace Judge model to early partners today Designed to detect errors in agent trajectories at 1/100th of the …

AgentsDGX agent

We are starting to role out our Trace Judge model to early partners today Designed to detect errors in agent trajectories at 1/100th of the cost compared to closed models If interested in early access

We have more releases to come for Zenith & our sota open source II Agent, pushing the boundaries of novel & first principles reasoning Large…

AgentsDGX agent

We have more releases to come for Zenith & our sota open source II Agent, pushing the boundaries of novel & first principles reasoning Larger models will show strong numbers, but we've reached a point

When Does Personality Composition Matter for Multi-Agent LLM Teams?

AgentsDGX agent

arXiv:2606.27443v1 Announce Type: new Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-exp

28 Jun 2026

Fleet agents can be shared in channels where your company is already doing work. Slack, Teams, etc Video from Caspar showing how to bring th…

TutorialsDGX agent

Fleet agents can be shared in channels where your company is already doing work. Slack, Teams, etc Video from Caspar showing how to bring them into Slack! For agents to spread through a company they h

There's a big difference between a single model call and serving an agent at scale. @ZainHasan6 breaks down what actually changes. Catch our…

AgentsDGX agent

There's a big difference between a single model call and serving an agent at scale. @ZainHasan6 breaks down what actually changes. Catch our team this Monday at 9 a.m. PST for their open-source infere

26 Jun 2026

A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation

AgentsDGX agent

arXiv:2603.09448v2 Announce Type: replace-cross Abstract: Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers. W

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

AgentsDGX agent

arXiv:2606.27330v1 Announce Type: cross Abstract: Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks in

Hermes Agent + Computer Use by @trycua is pretty cool! Looking at Hermes interacting with apps and windows is mind blowing and a bit scary a…

AgentsDGX agent

Hermes Agent + Computer Use by @trycua is pretty cool! Looking at Hermes interacting with apps and windows is mind blowing and a bit scary at the same time 😂 Here on Mac with MiniMax M3 and Reachy Min

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

SafetyDGX agent

arXiv:2606.26575v1 Announce Type: cross Abstract: Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of learning-based

In a real conversation, deciding when to speak takes about as much brainpower as deciding what to say. Voice agents haven't been built that …

AgentsDGX agent

In a real conversation, deciding when to speak takes about as much brainpower as deciding what to say. Voice agents haven't been built that way. @SierraPlatform's unlock was parallelizing thinking, li

not trying to glaze but the hermes agent experience is so much better now it is running so much of our triage stuff now at ***** ai labs now…

AgentsDGX agent

not trying to glaze but the hermes agent experience is so much better now it is running so much of our triage stuff now at ***** ai labs now, still have to correct/customize stuff every now and then b

← Previous
1…7980818283…299
Next →