AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
16 May 2026

Nectar Social, which offers an agentic OS for marketers, raised a $30M Series A led by Menlo Ventures, with GV and True Ventures among investors (Dominic-Madori Davis/TechCrunch)

AgentsDGX agent

Dominic-Madori Davis / TechCrunch: Nectar Social, which offers an agentic OS for marketers, raised a 30M Series A led by Menlo Ventures, with GV and True Ventures among investors — AI-powered marketin

Your agent also will now automatically have X Search tool available when using your Grok subscription as well! See details here: https://her…

AgentsDGX agent

Your agent also will now automatically have X Search tool available when using your Grok subscription as well! See details here: https://hermes-agent.nousresearch.com/docs/user-guide/features/x-search

15 May 2026

AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.13940v1 Announce Type: cross Abstract: Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents

Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

Model ReleasesDGX agent

arXiv:2605.14322v1 Announce Type: new Abstract: Language agents are increasingly deployed in complex professional workflows, with tutoring emerging as a particularly high-stakes capability that remain

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both

AgentsDGX agent

arXiv:2605.15198v1 Announce Type: cross Abstract: Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

AgentsDGX agent

arXiv:2510.02837v2 Announce Type: replace Abstract: Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory as

Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification

AgentsDGX agent

arXiv:2605.14495v1 Announce Type: cross Abstract: Multimedia verification requires not only accurate conclusions but also transparent and contestable reasoning. We propose a contestable multi-agent fr

Databricks brings GPT-5.5 to enterprise agent workflows

Model ReleasesDGX agent

Databricks has integrated OpenAI's GPT-5.5 model into enterprise agent workflows, enabling organizations to build and deploy AI agents with advanced language capabilities. This partnership leverages D

Fine, you all want to code like this I guess. (Runway's new Agent mode is quite impressive, doing fairly complex story building from just a …

AgentsDGX agent

Fine, you all want to code like this I guess. (Runway's new Agent mode is quite impressive, doing fairly complex story building from just a short text description of what you want. Not error free obvi

@hwchase17 Started on this and finding it awesome; also LangSmith engine sparked an idea. The 'Dependabot like for LLM agent failures'. Lang…

AgentsDGX agent

@hwchase17 Started on this and finding it awesome; also LangSmith engine sparked an idea. The 'Dependabot like for LLM agent failures'. LangSmith Engine gives you the smoke detector. The natural next

IFPV: An Integrated Multi-Agent Framework for Generative Operational Planning and High-Fidelity Plan Verification

AgentsDGX agent

arXiv:2605.14851v1 Announce Type: cross Abstract: Operational plan generation and verification are critical for modern complex and rapidly changing battlefield environments, yet traditional generation

it’s kind of awesome that continual learning has now come to the agent & harness level. i remember when online learning was the big craze in…

AgentsDGX agent

it’s kind of awesome that continual learning has now come to the agent & harness level. i remember when online learning was the big craze in traditional ml a couple of years ago, that transitioned int

@LangChain’s Interrupt 2026 conference was a blast!! such a pleasure to with @VictorMoreira16 about deep agents! ICYMI: we just dropped v0.6…

AgentsDGX agent

@LangChain’s Interrupt 2026 conference was a blast!! such a pleasure to with @VictorMoreira16 about deep agents! ICYMI: we just dropped v0.6, which is focused on performance at the model, harness, and

MediaClaw: Multimodal Intelligent-Agent Platform Technical Report

AgentsDGX agent

arXiv:2605.14771v1 Announce Type: new Abstract: MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, plu

Run @NousResearch's Hermes Agent fully locally on DGX Spark. 🚀 Our newest playbook shows you how to get set up via @Ollama step by step. 👇

Local AiDGX agent

This playbook provides step-by-step instructions for running Nous Research's Hermes Agent locally on NVIDIA DGX Spark using Ollama, enabling users to deploy an open-source AI agent entirely on local h

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents

Model ReleasesDGX agent

arXiv:2605.14033v1 Announce Type: new Abstract: Scientific theory shift in AI agents requires more than fitting equations to data. An artificial scientific agent must detect whether an existing repres

Temporal Fair Division in Multi-Agent Systems: From Precise Alternation Metrics to Scalable Coordination Proxies

SafetyDGX agent

arXiv:2605.14879v1 Announce Type: cross Abstract: A plethora real-world environments require agents to compete repeatedly for the same limited resource, calling for a temporal notion of fairness judge

14 May 2026

A Multi-Agent Orchestration Framework for Venture Capital Due Diligence

AgentsDGX agent

arXiv:2605.13110v1 Announce Type: cross Abstract: We present a fully automated multi-agent framework for corporate due diligence and market analysis in venture capital. The system runs on an event-dri

AI co-mathematician: Accelerating mathematicians with agentic AI

AgentsDGX agent

arXiv:2605.06651v2 Announce Type: replace Abstract: We introduce the AI co-mathematician, a workbench for mathematicians to interactively leverage AI agents to pursue open-ended research. The AI co-ma

Aleph, our fully autonomous AI agent system for formal verification, aced all major theorem proving benchmarks including PutnamBench, VeriSo…

AgentsDGX agent

Aleph is a fully autonomous AI agent system developed for formal verification that has achieved top performance across major theorem proving benchmarks, including PutnamBench and VeriSo. The system re

CANTANTE: Optimizing Agentic Systems via Contrastive Credit Attribution

Local AiDGX agent

arXiv:2605.13295v1 Announce Type: cross Abstract: LLM-based multi-agent systems have demonstrated strong performance across complex real-world tasks, such as software engineering, predictive modeling,

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies

SafetyDGX agent

arXiv:2508.01049v2 Announce Type: replace Abstract: Independent on-policy policy gradient algorithms are widely used for multi-agent reinforcement learning (MARL) in cooperative and no-conflict games,

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation

Model ReleasesDGX agent

arXiv:2605.12857v1 Announce Type: cross Abstract: Existing API-based agentic systems for RTL code generation are fundamentally misaligned with industrial practice: they assume a golden testbench is av

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Model ReleasesDGX agent

arXiv:2605.12673v1 Announce Type: new Abstract: Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hackin

From manual to autonomous: how AI agents are transforming electric grid operations

AgentsDGX agent

AI agents are automating traditional manual processes in electric grid operations, improving efficiency and response times for grid management tasks. These autonomous systems leverage machine learning

if you're already thinking about hacking around on agents this weekend, why not play around with @videodb_io api and potentially win a prize…

AgentsDGX agent

if you're already thinking about hacking around on agents this weekend, why not play around with @videodb_io api and potentially win a prize? videodb is an API for turning streaming video into text in

Introducing LangSmith LLM Gateway: The runtime governance layer for your agents. 💸 Enforce cost limits 🔒 Detect PII ✅ Act on violations …A…

AgentsDGX agent

Introducing LangSmith LLM Gateway: The runtime governance layer for your agents. 💸 Enforce cost limits 🔒 Detect PII ✅ Act on violations …All without leaving LangSmith. Now in Private Beta https://www.

our engineers had an awesome time attending Interrupt:2026 by @LangChain learning about observability, agent evals, and best practices from …

AgentsDGX agent

our engineers had an awesome time attending Interrupt:2026 by @LangChain learning about observability, agent evals, and best practices from the industry leaders! can't wait to see what they cook up wi

Quantitative Certification of Agentic Tool Selection

SafetyDGX agent

arXiv:2510.03992v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant extern

Such a cool hackathon submission for the Hermes Agent @Kimi_Moonshot track!

AgentsDGX agent

Such a cool hackathon submission for the Hermes Agent @Kimi_Moonshot track! I mean just look at this beauty! seriously well deserved finalist @evvaaannnn in the @NousResearch x @Kimi_Moonshot creative

suuuuper excited to be collaborating with the excellent LangChain Labs team on this effort prod agent tracing is the seed that lets you clos…

AgentsDGX agent

suuuuper excited to be collaborating with the excellent LangChain Labs team on this effort prod agent tracing is the seed that lets you close the loop for continual learning. too much data gets collec

Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2505.11556v4 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet syst

There's a ton on unexplored space in building enterprise-grade agent harnesses that continuously improve over time. Congrats to @hwchase17 a…

AgentsDGX agent

There's a ton on unexplored space in building enterprise-grade agent harnesses that continuously improve over time. Congrats to @hwchase17 and @LangChain on the launch of LangChain Labs and excited to

ToolMol: Evolutionary Agentic Framework for Multi-objective Drug Discovery

AgentsDGX agent

arXiv:2605.12784v1 Announce Type: new Abstract: Advances in large language models (LLMs) have recently opened new and promising avenues for small-molecule drug discovery. Yet existing LLM-based approa

You can now import your project from Lovable, Base44, V0 into @Replit for free. After importing, Replit Agent will build a free mobile app f…

AgentsDGX agent

You can now import your project from Lovable, Base44, V0 into @Replit for free. After importing, Replit Agent will build a free mobile app for it and get it onto the App Store in minutes. All free for

Your AI agent can create an entire Google Form for you by chatting, and it will use the browser to type and build the survey automatically.

AgentsDGX agent

An AI agent can autonomously create a complete Google Form through natural conversation, using browser automation to directly interact with Google Forms' interface by typing and building the survey wi

yst on stream codex wrote a skill for my hermes agent to make music it can take in feedback and improve its output i clipped the best music …

AgentsDGX agent

yst on stream codex wrote a skill for my hermes agent to make music it can take in feedback and improve its output i clipped the best music it made with minimax. still far from good music but im prett

13 May 2026

4/5 Best-of-N: Run one agent config N times in parallel → select the best trajectory. Leverages LLMs’ non-determinism - but hinges on a good…

AgentsDGX agent

4/5 Best-of-N: Run one agent config N times in parallel → select the best trajectory. Leverages LLMs’ non-determinism - but hinges on a good eval mechanism (we use an LLM-as-a-Judge). Can increase acc

5/5 Ensemble: Run multiple distinct agent configs in parallel → select the best trajectory. Leverages success of the portfolio vs single var…

AgentsDGX agent

5/5 Ensemble: Run multiple distinct agent configs in parallel → select the best trajectory. Leverages success of the portfolio vs single variants (see also: @/karpathy’s LLM Council). Can outperform b

An Empirical Study of Automating Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.11378v1 Announce Type: new Abstract: Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition

AgentsDGX agent

arXiv:2605.11388v1 Announce Type: new Abstract: Humans intuitively solve complex problems by flexibly shifting among reasoning modes: they plan, execute, revise intermediate goals, resolve ambiguity t

Do multi-agent systems make LLM reasoning better? Most AI devs assume that it should. But this new paper shows that this is often not the ca…

SafetyDGX agent

Do multi-agent systems make LLM reasoning better? Most AI devs assume that it should. But this new paper shows that this is often not the case. It ran 22,500 deterministic trajectories across GAIA, SW

DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers

AgentsDGX agent

arXiv:2605.11683v1 Announce Type: new Abstract: Vision Transformers (ViTs) incur significant computational overhead due to the quadratic complexity of self-attention relative to the token sequence len

Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

AgentsDGX agent

arXiv:2605.11369v1 Announce Type: new Abstract: Generating physically plausible dynamic motions of human-object interaction (HOI) remains challenging, mainly due to existing HOI datasets limited to st

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Model ReleasesDGX agent

arXiv:2605.11086v1 Announce Type: cross Abstract: AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is

Great example of why you should 1. Run your agent on a separate machine from the sandbox it uses (e.g. sandbox as a tool) 2. Never set env v…

AgentsDGX agent

Great example of why you should 1. Run your agent on a separate machine from the sandbox it uses (e.g. sandbox as a tool) 2. Never set env vars in your sandbox. Instead, use something like LangSmith’s

Introducing SWE-ZERO-12M-trajectories: the largest agentic trace dataset in the open, 5.7x larger than the previous largest. 112B tokens · 1…

AgentsDGX agent

Introducing SWE-ZERO-12M-trajectories: the largest agentic trace dataset in the open, 5.7x larger than the previous largest. 112B tokens · 12M trajectories · 122K PRs · 3K repos · 16 languages https:/

🚀Launching: LangSmith Engine LangSmith Engine is an agent that sits on top of your traces It runs in the background and automatically ident…

AgentsDGX agent

🚀Launching: LangSmith Engine LangSmith Engine is an agent that sits on top of your traces It runs in the background and automatically identifies issues It then proactively suggests action items (code

No More, No Less: Task Alignment in Terminal Agents

Model ReleasesDGX agent

arXiv:2605.12233v1 Announce Type: new Abstract: Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret

Not to mention 7 blog posts dropped today, including a new Deep Agents version with significant improvements… Especially excited about the C…

AgentsDGX agent

Not to mention 7 blog posts dropped today, including a new Deep Agents version with significant improvements… Especially excited about the Code Interpreter feature which is a sneaky powerful feature e

Securing AI agents: How AWS and Cisco AI Defense scale MCP and A2A deployments

AgentsDGX agent

The Cisco and AWS partnership addresses three challenges enterprises face when scaling AI agents: visibility gaps, security bottlenecks, and compliance risks. In this post, we explore how you can over

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

Model ReleasesDGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

The Hermes Agent Creative Hackathon sponsored by @Kimi_Moonshot has ended! Finalists were selected by Nous and Kimi staff out of 227 submiss…

AgentsDGX agent

The Hermes Agent Creative Hackathon sponsored by @Kimi_Moonshot has ended! Finalists were selected by Nous and Kimi staff out of 227 submissions on creativity, usefulness and presentation. We were abs

Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the conseq…

AgentsDGX agent

Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the consequences of their actions before taking them 'they just act, a

12 May 2026

AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning

Local AiDGX agent

arXiv:2410.13181v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have been remarkable. Users face a choice between using cloud-based LLMs for generation quality

CellDX AI Autopilot: Agent-Guided Training and Deployment of Pathology Classifiers

HardwareDGX agent

arXiv:2605.10362v1 Announce Type: new Abstract: Training AI models for computational pathology currently requires access to expensive whole-slide-image datasets, GPU infrastructure, deep expertise in

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

Model ReleasesDGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.09826v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent

Enhancing Consistency Models for Multi-Agent Trajectory Prediction

AgentsDGX agent

arXiv:2605.08572v1 Announce Type: new Abstract: Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

← Previous
1…8586878889…300
Next →