AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
18 May 2026

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design

Model ReleasesDGX agent

arXiv:2605.15871v1 Announce Type: new Abstract: Toward recursive self-improvement, we investigate LLM agents autonomously designing foundation models beyond standard Transformers. We introduce a dual-

announcing deepagents v0.6, our biggest release yet! it’s all about performance: at the model layer w harness profiles, agent layer w code i…

AgentsDGX agent

announcing deepagents v0.6, our biggest release yet! it’s all about performance: at the model layer w harness profiles, agent layer w code interpreter, and at scale w streaming and delta channels cont

Congrats to the @cursor_ai team on Composer 2.5 — a huge milestone for agentic coding models. Together AI, the AI Native Cloud, is proud to …

Agents
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Congrats to the @cursor_ai team on Composer 2.5 — a huge milestone for agentic coding models. Together AI, the AI Native Cloud, is proud to partner on this launch. Composer 2.5 is pushing the frontier

Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning

AgentsDGX agent

arXiv:2605.15315v1 Announce Type: new Abstract: LLM-powered coding agents spend the majority of their token budget reading repository files, yet much of the retrieved code is irrelevant to the task at

DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery

AgentsDGX agent

arXiv:2605.15461v1 Announce Type: cross Abstract: Building state-of-the-art (SOTA) predictive models for drug discovery requires expensive search over tools, architectures, and training strategies. Cu

Nebius and @LangChain have partnered to integrate Nebius Token Factory with LangChain's Deep Agents. The integration, combined with LangChai…

AgentsDGX agent

Nebius and @LangChain have partnered to integrate Nebius Token Factory with LangChain's Deep Agents. The integration, combined with LangChain's existing Tavily integration, gives teams building on Lan

paper.json: A Coordination Convention for LLM-Agent-Actionable Papers

AgentsDGX agent

arXiv:2605.16194v1 Announce Type: cross Abstract: LLM agents routinely serve as first (and sometimes only) readers of academic papers, skimming for sub-claims, extracting reproducibility steps, and ge

Runtime-Structured Task Decomposition for Agentic Coding Systems

AgentsDGX agent

arXiv:2605.15425v1 Announce Type: cross Abstract: Agentic coding systems increasingly use large language models (LLMs) for software engineering tasks such as debugging, root cause analysis, and code r

17 May 2026

okay i’m finally starting to get my mind wrapped around this whole “stateful” agent thing

AgentsDGX agent

This post likely discusses Yohei Nakajima's understanding of stateful AI agents—systems that maintain and utilize information across multiple interactions rather than treating each conversation indepe

We gave a full 90 minute workshop on how to build agentic workflows over your enterprise documents at @aiDotEngineer Singapore 🇸🇬🦙 The ma…

AgentsDGX agent

We gave a full 90 minute workshop on how to build agentic workflows over your enterprise documents at @aiDotEngineer Singapore 🇸🇬🦙 The majority of unstructured information is locked up within PDFs. @h

16 May 2026

1/ you are probably overcomplicating the environment that your agent has to work with when running evals

AgentsDGX agent

This post likely discusses how developers often create unnecessarily complex evaluation environments when testing AI agents, and suggests simplifying these setups for more effective testing. The advic

@Gavriel_Cohen @thsottiaux head of AI Govtech at Singapore estimates 1.3 billion agents in the country in the next 2 years and is building a…

AgentsDGX agent

Singapore's head of AI Govtech estimates the country will have 1.3 billion AI agents deployed within the next two years and is actively developing infrastructure to support this expansion. This sugges

Interesting interpretability paper on tool-using agents. The authors probe hidden states and find the model often recognizes it should call …

AgentsDGX agent

Interesting interpretability paper on tool-using agents. The authors probe hidden states and find the model often recognizes it should call a tool, but fails to actually call one. The mismatch ranges

the Codex app is in a category of its own. “agentic excel on mac” is an interesting description.

AgentsDGX agent

the Codex app is in a category of its own. “agentic excel on mac” is an interesting description. gotta say Codex is completely unrecognizable from 3 months ago. guys went extreme founder mode on this

15 May 2026

Agentic Design of Compositional Descriptors via Autoresearch for Materials Science Applications

Model ReleasesDGX agent

arXiv:2605.14671v1 Announce Type: cross Abstract: Autoresearch offers a flexible paradigm for automating scientific tasks, in which an AI agent proposes, implements, evaluates, and refines candidate s

// Beyond Individual Intelligence // One of the more useful multi-agent surveys I've read this year. 200+ papers mapped along three axes: co…

AgentsDGX agent

// Beyond Individual Intelligence // One of the more useful multi-agent surveys I've read this year. 200+ papers mapped along three axes: collaboration mechanisms, failure attribution, and self-evolut

Boomi CTO: Agentic engineering is how enterprise AI finally earns its keep

AgentsDGX agent

Three years of enterprise AI investment, and most companies are still waiting for the payoff. Boomi LP thinks it has found it, unveiling Boomi Companion, a collection of open-source agent skills that

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

Model ReleasesDGX agent

arXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend

GEAR: Genetic AutoResearch for Agentic Code Evolution

AgentsDGX agent

arXiv:2605.13874v1 Announce Type: cross Abstract: Autonomous research agents can already run machine learning experiments without human supervision, but many rely on a narrow search strategy: they rep

Herculean: An Agentic Benchmark for Financial Intelligence

Model ReleasesDGX agent

arXiv:2605.14355v1 Announce Type: new Abstract: As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carr

Lang2MLIP: End-to-End Language-to-Machine Learning Interatomic Potential Development with Autonomous Agentic Workflows

AgentsDGX agent

arXiv:2605.14527v1 Announce Type: new Abstract: Developing machine learning interatomic potentials (MLIPs) for complex materials systems remains challenging because it requires expertise in atomistic

MemLineage: Lineage-Guided Enforcement for LLM Agent Memory

AgentsDGX agent

arXiv:2605.14421v1 Announce Type: cross Abstract: We introduce MemLineage, a defense for LLM agent memory that attaches both cryptographic provenance and LLM-mediated derivation lineage to every entry

Polaris: A Godel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair

Model ReleasesDGX agent

arXiv:2603.23129v2 Announce Type: replace Abstract: Godel agent realize recursive self-improvement: an agent inspects its own policy and traces and then modifies that policy in a tested loop. We intro

Progent: Securing AI Agents with Privilege Control

SafetyDGX agent

arXiv:2504.11703v3 Announce Type: replace-cross Abstract: AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger una

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

AgentsDGX agent

arXiv:2605.14563v1 Announce Type: cross Abstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding ag

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

Model ReleasesDGX agent

arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory

14 May 2026

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

AgentsDGX agent

arXiv:2605.12925v1 Announce Type: cross Abstract: Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view trea

Agents feedback tip

IndustryDGX agent

This article likely provides practical advice on how to effectively give feedback to AI agents or autonomous systems, covering techniques for improving agent performance through constructive feedback

An Agentic LLM-Based Framework for Population-Scale Mental Health Screening

AgentsDGX agent

arXiv:2605.13046v1 Announce Type: new Abstract: Mental health disorders affect millions worldwide, and healthcare systems are increasingly overwhelmed by the volume of clinical data generated from ele

CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent

AgentsDGX agent

arXiv:2508.01031v5 Announce Type: replace Abstract: Computer-Aided Design (CAD) is widely used for conceptual design and parametric 3D modeling, but typically requires a high level of expertise from d

Can LLM Agents Simulate Dynamic Networks? A Case Study on Email Networks with Phishing Synthesis

AgentsDGX agent

arXiv:2605.12507v1 Announce Type: cross Abstract: While Large Language Model (LLM) multi-agent systems (MAS) offer a transformative approach to simulating human behavior in complex systems, it remains

Counterfactual Reasoning for Causal Responsibility Attribution in Probabilistic Multi-Agent Systems

SafetyDGX agent

arXiv:2605.13077v1 Announce Type: cross Abstract: Responsibility allocation -- determining the extent to which agents are accountable for outcomes -- is a fundamental challenge in the design and analy

Great move. If you’re building agents, so much time is spent engineering your way through traces. And then comes the hard to rapidly iterate…

AgentsDGX agent

Great move. If you’re building agents, so much time is spent engineering your way through traces. And then comes the hard to rapidly iterate and figure out the changes needed. 🚀Launching: LangSmith En

if you're an artist or creative wondering how to use agents to accelerate your work, or skeptical that they can: please stop what you are do…

AgentsDGX agent

if you're an artist or creative wondering how to use agents to accelerate your work, or skeptical that they can: please stop what you are doing and look at the 6 examples in this thread The Hermes Age

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning

AgentsDGX agent

arXiv:2605.13037v1 Announce Type: new Abstract: Current interactive LLM agents rely on goal-conditioned stepwise planning, where environmental understanding is acquired reactively during execution rat

Natively-agentic is just one of the perks. The overall experience is a step change from anything out there. And then you add the focus to de…

AgentsDGX agent

Natively-agentic is just one of the perks. The overall experience is a step change from anything out there. And then you add the focus to details, user experience, and how supreme the team behind this

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

SafetyDGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

Position: Assistive Agents Need Accessibility Alignment

SafetyDGX agent

arXiv:2605.13579v1 Announce Type: new Abstract: Assistive agents for Blind and Visually Impaired (BVI) users require accessibility alignment as a first-class design objective. Despite rapid progress i

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

SafetyDGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

the future intelligent agents will learn and grow with your team over ultra long timescales 🕰️ that research direction centers around Conti…

AgentsDGX agent

the future intelligent agents will learn and grow with your team over ultra long timescales 🕰️ that research direction centers around Continual Learning LangSmith is the data layer that helps us under

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

Model ReleasesDGX agent

arXiv:2602.16246v3 Announce Type: replace Abstract: Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Ben

Untrained AI agents are easy security targets — they don’t know bad people exist, says KnowBe4 CEO

AgentsDGX agent

The dual-threat landscape of enterprise AI security is coming into focus. The same autonomous agents transforming workforce productivity are also expanding the attack surface — and most organizations

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling

Model ReleasesDGX agent

arXiv:2605.13172v1 Announce Type: cross Abstract: Recent advances in agent and multi-agent systems have shown strong performance on tool use, reasoning, and collaborative tasks. However, existing benc

13 May 2026

1/5 We’re seeing 4 common agent optimization methods for hitting the right accuracy-cost or accuracy-latency tradeoff. We tried them all out…

AgentsDGX agent

AI21 Labs discusses four common agent optimization methods used to balance accuracy against cost and latency constraints. The post indicates the team evaluated all four approaches, likely covering tec

ABRA: Agent Benchmark for Radiology Applications

Model ReleasesDGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

ComfyUI Skill in Hermes Agent Deep Dive https://x.com/i/broadcasts/1NGaraWmzBdJj

AgentsDGX agent

This broadcast likely provides an in-depth technical discussion of how to implement and utilize ComfyUI skills within the Hermes Agent framework, covering integration methods, practical applications,

Coming up at Interrupt: 🎙️ How @clay scales their GTM Engineering agents with Head of AI @JeffBarg. 🎙️ @ummadisetti and @kordelfrance on h…

AgentsDGX agent

Coming up at Interrupt: 🎙️ How @clay scales their GTM Engineering agents with Head of AI @JeffBarg. 🎙️ @ummadisetti and @kordelfrance on how @Toyota Motor North America equipped 56,000 employees with

Customers like Decagon, Amplitude, BILT, and Snyk use development environments to let their agents handle tasks end-to-end. Learn more: http…

AgentsDGX agent

Cursor highlights that companies like Decagon, Amplitude, BILT, and Snyk leverage development environments within Cursor to enable AI agents to complete tasks independently from start to finish. This

Exclusive: AirOps targets emerging AI search market with autonomous content optimization agent

AgentsDGX agent

AirOps, the business name of Rivington Labs Inc., today introduced Quill, an artificial intelligence agent designed to help brands maintain visibility in generative AI search engines by continuously m

GeomHerd: A Forward-looking Herding Quantification via Ricci Flow Geometry on Agent Interactive Simulations

Model ReleasesDGX agent

arXiv:2605.11645v1 Announce Type: cross Abstract: Herding -- where agents align their behaviors and act collectively -- is a central driver of market fragility and systemic risk. Existing approaches t

Meet Runway Agent. Your new AI creative partner that helps you ideate and execute fully finished, sound designed and edited videos. All with…

AgentsDGX agent

Meet Runway Agent. Your new AI creative partner that helps you ideate and execute fully finished, sound designed and edited videos. All with just a simple conversation. From ads to shorts to content f

Multi-Agent System Identification with Nonlinear Sheaf Diffusion

AgentsDGX agent

arXiv:2605.11204v1 Announce Type: cross Abstract: Local interaction laws governing multi-agent systems can be difficult to recover from trajectory data, even when the dynamics are observed faithfully.

Observability going from 'read the traces' to 'the traces read themselves and tell you what to fix.' That's the feedback loop most agent tea…

AgentsDGX agent

Observability going from 'read the traces' to 'the traces read themselves and tell you what to fix.' That's the feedback loop most agent teams are missing. 🚀Launching: LangSmith Engine LangSmith Engin

On Problems of Implicit Context Compression for Software Engineering Agents

AgentsDGX agent

arXiv:2605.11051v1 Announce Type: cross Abstract: LLM-based Software Engineering agents face a critical bottleneck: context length limitations cause failures on complex, long-horizon tasks. One promis

PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement

AgentsDGX agent

arXiv:2605.11225v1 Announce Type: cross Abstract: Large language model (LLM)-based agents frequently generate seemingly coherent plans that fail upon execution due to infeasible actions, constraint vi

P.S. Join our Discord to view all of the submissions and get involved with the Hermes Agent community https://discord.gg/nousresearch

AgentsDGX agent

Nous Research invites community members to join their Discord server to view submissions and participate in the Hermes Agent community. The Discord link provided offers access to community discussions

Robust Multi-Agent Path Finding under Observation Attacks: A Principled Adversarial-Plus-Smoothing Training Recipe

SafetyDGX agent

arXiv:2605.11469v1 Announce Type: new Abstract: Decentralized multi-agent path finding (MAPF) routes a team of agents on a shared grid, each acting from its own local view. The standard solution train

Thank you again to all 227 submitters! We loved seeing the range of creative domains and all of the unusual uses for Hermes Agent you found.…

AgentsDGX agent

Thank you again to all 227 submitters! We loved seeing the range of creative domains and all of the unusual uses for Hermes Agent you found. The team will be highlighting more personal favorites and s

12 May 2026

A beginner-friendly guide to building a simple RAG (RETRIEVAL-AUGMENTED GENERATION) AI Agent with Pinecone + n8n 🧵 If you’re learning AI au…

AgentsDGX agent

A beginner-friendly guide to building a simple RAG (RETRIEVAL-AUGMENTED GENERATION) AI Agent with Pinecone + n8n 🧵 If you’re learning AI automation and you’re confused about Pinecone, knowledge bases,

Agentic SOC startup Exaforce closes 125M round at reported 725M valuation

AgentsDGX agent

Agentic security operations startup Exaforce Inc. today announced that it has raised 125 million in new funding to scale up its AI-driven security operations platform and expand its real-time threat r

← Previous
1…7273747576…299
Next →