AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
22 May 2026

Great conversation on the Max Agency podcast with @cogent_security Co-Founder + CTO Geng Sng on building agents for autonomous cyber defense…

AgentsDGX agent

Great conversation on the Max Agency podcast with @cogent_security Co-Founder + CTO Geng Sng on building agents for autonomous cyber defense. Check out the full episode: ⏯️ YouTube: https://youtu.be/D

im going to be in NYC in ~1 week, and am doing a fireside chat with one of the top agent companies in new york - @_anish_agarwal of @travers…

AgentsDGX agent

im going to be in NYC in ~1 week, and am doing a fireside chat with one of the top agent companies in new york - @_anish_agarwal of @traversal_ai come join us! https://partiful.com/e/MLzglw6al1KrUSXSr

MARS: Modular Agent with Reflective Search for Automated AI Research

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.02660v3 Announce Type: replace Abstract: A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general softw

21 May 2026

APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2605.21240v1 Announce Type: new Abstract: LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision mak

Grok Build 0.1 is now available for early access in Hermes Agent

AgentsDGX agent

Nous Research has released Grok Build 0.1 as an early access feature within the Hermes Agent platform. This release likely introduces initial capabilities for building and customizing Grok-based appli

Here's more about Datasette Agent on my blog: https://simonwillison.net/2026/May/21/datasette-agent/ - and the announcement on the new Datas…

AgentsDGX agent

Here's more about Datasette Agent on my blog: https://simonwillison.net/2026/May/21/datasette-agent/ - and the announcement on the new Datasette project blog too: https://datasette.io/blog/2026/datase

I don’t know why people talk in absolutes about agentic design. Especially for coders and builders. We just invented this stuff. It’s wild t…

AgentsDGX agent

I don’t know why people talk in absolutes about agentic design. Especially for coders and builders. We just invented this stuff. It’s wild to think it won’t radically change over the next year. We’re

Replit shipped 5 new security tools in the last month: Security Agent, Auto-Protect, Security Center V2, External Access Tokens, and Private…

AgentsDGX agent

Replit shipped 5 new security tools in the last month: Security Agent, Auto-Protect, Security Center V2, External Access Tokens, and Private Publishing. 🛡️ Tomorrow we go live with @AlexandreCuoci, wh

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Model ReleasesDGX agent

arXiv:2605.21384v1 Announce Type: cross Abstract: As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Re

> These engineers can review their agent's code much faster than reviewing human code. wat

AgentsDGX agent

> These engineers can review their agent's code much faster than reviewing human code. wat Today we reduced headcount by 22%. The business is the strongest it's ever been. So I think it's important to

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

ResearchDGX agent

arXiv:2605.20291v1 Announce Type: new Abstract: Large language models (LLMs) have enabled web agents that follow natural language goals through multi-step browser interactions. However, agents fine-tu

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

AgentsDGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

Model ReleasesDGX agent

arXiv:2605.21404v1 Announce Type: new Abstract: We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was ru

working on a 'take this vibecoded slop app and make it a production-ready, e2e tested, maintainable, parallelizable agent repo' skill. this …

AgentsDGX agent

working on a 'take this vibecoded slop app and make it a production-ready, e2e tested, maintainable, parallelizable agent repo' skill. this thing ran for ~16 hours yesterday and made 103 commits all t

20 May 2026

A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation

AgentsDGX agent

arXiv:2605.19316v1 Announce Type: new Abstract: Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting

Add a Specialized Deep Research Skill to Agent Harnesses

Model ReleasesDGX agent

This article covers the NVIDIA AI-Q Blueprint for building specialized deep research agents that empower AI systems to gather context, synthesize information, and support complex decision-making acros

as agents run for longer and with more context, there's a lot of tricky things you have to consider for deployment! super excited to dive in…

ApplicationsDGX agent

as agents run for longer and with more context, there's a lot of tricky things you have to consider for deployment! super excited to dive into deploying Deep Agents for this Boston tech week meetup! .

Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers question…

AgentsDGX agent

Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers questions with exact citations highlighted on the original PDF page.

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

Model ReleasesDGX agent

arXiv:2605.18758v1 Announce Type: cross Abstract: Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction rout

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

Model ReleasesDGX agent

arXiv:2605.18805v1 Announce Type: cross Abstract: LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet ex

Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from e…

AgentsDGX agent

Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from experimentation to production. Sovereign AI. For all. Downloa

the core concept is a graph that represents everything about the agents knowledge, history, behaviors, capabilities graph is made of events …

AgentsDGX agent

the core concept is a graph that represents everything about the agents knowledge, history, behaviors, capabilities graph is made of events behaviors react to graph changes relationships can carry beh

When an agent is allowed to decompose a goal into smaller sub-tasks, it frequently suffers from goal drift. Left unchecked, it will redefine…

AgentsDGX agent

When an agent is allowed to decompose a goal into smaller sub-tasks, it frequently suffers from goal drift. Left unchecked, it will redefine the optimization metric to favor a simpler, useless sub-tas

working on some hex dashboards -- this is a bummer! if you use langchain/deepagents as your agent harness, you can swap models with fallback…

AgentsDGX agent

working on some hex dashboards -- this is a bummer! if you use langchain/deepagents as your agent harness, you can swap models with fallback middleware when your first choice provider is down https://

You can also now create automations without any attached repos, like a daily Slack digest agent that prioritizes unread threads and DMs for …

AgentsDGX agent

You can also now create automations without any attached repos, like a daily Slack digest agent that prioritizes unread threads and DMs for your response. Find more templates in our marketplace: http:

19 May 2026

Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments

AgentsDGX agent

arXiv:2605.16419v1 Announce Type: cross Abstract: Kinematic monitoring plays a critical role in long-term rehabilitation for patients with spinal cord injury (SCI), where multi-view markerless motion

Announcing Claude Managed Agents on Cloudflare

Model ReleasesDGX agent

Cloudflare has integrated with Anthropic's Claude Managed Agents to provide a fast, isolated execution environment for autonomous code delivery. This means builders can scale agent workflows globally

ClawArena: Benchmarking AI Agents in Evolving Information Environments

Model ReleasesDGX agent

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse

Model ReleasesDGX agent

arXiv:2605.17450v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to ins

Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a direct swap to o…

AgentsDGX agent

Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a direct swap to open source models can bring immediate cost savings, that’s n

Cursor is now available in Jira. Assign Cursor to work items, or mention @​Cursor in a comment to kick off a cloud agent. Cursor uses the ti…

AgentsDGX agent

Cursor is now available in Jira. Assign Cursor to work items, or mention @​Cursor in a comment to kick off a cloud agent. Cursor uses the title, description, comments, and your team's repository setti

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

AgentsDGX agent

arXiv:2605.17439v1 Announce Type: cross Abstract: Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-lev

感谢 @dotey 将他的 baoyu-comic skill 移植到 Hermes Agent! 这个 skill 可以把一段 prompt 或源文档转成多页知识漫画,支持 6 种风格、7 种语气、7 种版式和 5 个预设。 http://github.com/NousRese…

AgentsDGX agent

感谢 @dotey 将他的 baoyu-comic skill 移植到 Hermes Agent! 这个 skill 可以把一段 prompt 或源文档转成多页知识漫画,支持 6 种风格、7 种语气、7 种版式和 5 个预设。 http://github.com/NousResearch/hermes-agent/tree/main/skills/creative/baoyu-comic Medi

Evaluating Cognitive Age Alignment in Interactive AI Agents

Model ReleasesDGX agent

arXiv:2605.17894v1 Announce Type: new Abstract: While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across doma

Natural-Language Agent Harnesses

SafetyDGX agent

arXiv:2603.25723v2 Announce Type: replace-cross Abstract: Agent performance is strongly shaped by the surrounding harness: the external execution system around a model that organizes a task run. Yet t

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

SafetyDGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

SafetyDGX agent

arXiv:2605.16551v1 Announce Type: new Abstract: Evaluating LLM-based agents remains challenging because identifying meaningful failure cases often requires substantial human effort to design realistic

Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

SafetyDGX agent

arXiv:2605.17036v1 Announce Type: new Abstract: This paper studies autonomous generative AI agents in multi-echelon supply chains using the MIT Beer Game. We identify four inference-time levers that s

Responsible Agentic AI Requires Explicit Provenance

Model ReleasesDGX agent

arXiv:2605.17169v1 Announce Type: new Abstract: Agentic AI is rapidly proliferating across diverse real-world domains such as software engineering, yet public trust has not kept pace. The central reas

Scalable Environments Drive Generalizable Agents

ResearchDGX agent

arXiv:2605.18181v1 Announce Type: new Abstract: Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such gene

Starting today, use your Grok or X Premium subscription in @openclaw. Chat with your agent, generate images and videos, or search for X post…

AgentsDGX agent

Grok and X Premium subscribers can now access their subscriptions within OpenClaw, enabling them to chat with AI agents, generate images and videos, and search X posts directly through the platform. T

TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents

SafetyDGX agent

arXiv:2605.17320v1 Announce Type: cross Abstract: Computer-use agents increasingly operate inside live personal workspaces, where their actions can modify files, applications, GUI state, credentials,

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

SafetyDGX agent

arXiv:2605.17480v1 Announce Type: new Abstract: Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates ne

The @xai team has published a full setup guide on how to use the xurl skill, which allows your Hermes Agent to read and write to X on your b…

AgentsDGX agent

The @xai team has published a full setup guide on how to use the xurl skill, which allows your Hermes Agent to read and write to X on your behalf — posting, searching, pulling bookmarks, managing list

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.16909v1 Announce Type: new Abstract: Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

Model ReleasesDGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

18 May 2026

ALSO: Adversarial Online Strategy Optimization for Social Agents

Model ReleasesDGX agent

arXiv:2605.15768v1 Announce Type: new Abstract: Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving cont

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents

AgentsDGX agent

arXiv:2512.01089v2 Announce Type: replace Abstract: Automated Scientific Discovery (ASD) systems can help automatically generate and run code-based experiments, but their capabilities are limited by t

Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis

AgentsDGX agent

arXiv:2605.15569v1 Announce Type: cross Abstract: Microservices are widely adopted in modern cloud systems due to their scalability and fault tolerance. However, microservice architectures introduce s

@gabrielchua the agentic excel thing is basically what u get when u expand the side panel to be the main thing https://x.com/jxnlco/status/2…

AgentsDGX agent

@gabrielchua the agentic excel thing is basically what u get when u expand the side panel to be the main thing https://x.com/jxnlco/status/2056139571641872765?s=12 jason from the codex team here, here

I resurrected my 8-year-old phone into a Hermes Agent server⚕️ Replaced Android with postmarketOS (Alpine Linux on ARM64), runs Hermes as a …

AgentsDGX agent

I resurrected my 8-year-old phone into a Hermes Agent server⚕️ Replaced Android with postmarketOS (Alpine Linux on ARM64), runs Hermes as a systemd service. I chat with it over Matrix (E2E encrypted)

ICYMI: SmithDB is our purpose-built data layer for agent observability + eval workloads. Supporting increasingly complex query patterns at l…

AgentsDGX agent

ICYMI: SmithDB is our purpose-built data layer for agent observability + eval workloads. Supporting increasingly complex query patterns at low latency, over large traces, with self-hosting + multi-clo

Introducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in …

AgentsDGX agent

Introducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in real-time. Try our playable research preview today, with Ago

LangSmith Engine automates the full agent fix loop — detecting failures, diagnosing causes and drafting PRs. But multi-model enterprises say…

AgentsDGX agent

LangSmith Engine automates the full agent fix loop — detecting failures, diagnosing causes and drafting PRs. But multi-model enterprises say a neutral observability layer still wins. http://venturebea

NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it.

Model ReleasesDGX agent

NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it. NEW paper from Meta. (bookmark it) It's an agent system that autonomously discove

Run Claude Managed Agents with Vercel Sandbox

Model ReleasesDGX agent

This article describes how to run Claude's managed agents within Vercel's Sandbox environment, enabling developers to execute AI agent workloads on Vercel's infrastructure. The integration allows user

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

Model ReleasesDGX agent

arXiv:2605.16116v1 Announce Type: new Abstract: Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible,

STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices

Model ReleasesDGX agent

arXiv:2605.15581v1 Announce Type: new Abstract: LLM-based root cause analysis (RCA) agents have recently emerged as a promising paradigm for incident diagnosis in microservice AIOps. However, their re

We are currently hiring full stack engineers to work on managed services, nous portal, UX/UI for hermes agent and applications around it, an…

AgentsDGX agent

We are currently hiring full stack engineers to work on managed services, nous portal, UX/UI for hermes agent and applications around it, and solving technical challenges cross-domain. If you're inter

17 May 2026

Publicis agrees to acquire LiveRamp, which allows companies to share and build new data sets and models that can power agentic frameworks, for $2.2B in cash (Alison Weissbrot/Adweek)

AgentsDGX agent

Alison Weissbrot / Adweek: Publicis agrees to acquire LiveRamp, which allows companies to share and build new data sets and models that can power agentic frameworks, for 2.2B in cash — Publicis Groupe

← Previous
1…8485868788…300
Next →