AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
Agents

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

DGX agent

arXiv:2601.22984v2 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evalua

agentsarxiv-cs-ai
26 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism

DGX agent

arXiv:2605.22993v1 Announce Type: cross Abstract: Characteristic linguistic behaviors associated with Social Language Disorder (SLD) in autism spectrum disorder, including echoic repetition, pronoun d

agentsarxiv-cs-ai
25 May 2026
Agents

(I'm firmly on team red/green TDD for agent code, I like having a test suite that protects against them breaking old features when they make…

DGX agent

(I'm firmly on team red/green TDD for agent code, I like having a test suite that protects against them breaking old features when they make new changes - https://simonwillison.net/guides/agentic-engi

agentssimon-willison--x
25 May 2026
Model Releases

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

DGX agent

arXiv:2605.23657v1 Announce Type: new Abstract: Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improvin

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

DGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PhotoFlow: Agentic 3D Virtual Photography Missions

DGX agent

arXiv:2605.23771v1 Announce Type: cross Abstract: Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene in

model-releasesarxiv-cs-ai
25 May 2026
Research

SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval

DGX agent

arXiv:2601.03260v2 Announce Type: replace-cross Abstract: AI agents have seen widespread adoption in information retrieval for scientific research, giving rise to tools such as Deep Research. However,

researcharxiv-cs-cl
25 May 2026
Agents

last night i got an agent to fork itself, propose a modification to itself on the fork, run through tests (sandbox, etc), and only accept th…

DGX agent

Yohei Nakajima describes a technical demonstration where an AI agent was configured to create a fork of itself, propose modifications to the forked version, execute tests in a sandboxed environment, a

agentsyohei-nakajima--x
23 May 2026
Model Releases

Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

DGX agent

arXiv:2605.21768v1 Announce Type: new Abstract: Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions.

model-releasesarxiv-cs-lg
23 May 2026
Agents

AMD CEO Lisa Su projects the CPU market will grow over 35% annually through 2031, up from 3% to 4% historically, driven by AI inference and agentic AI demand (Cheng Ting-Fang/Nikkei Asia)

DGX agent

Cheng Ting-Fang / Nikkei Asia: AMD CEO Lisa Su projects the CPU market will grow over 35% annually through 2031, up from 3% to 4% historically, driven by AI inference and agentic AI demand — TAIPEI —

agentstechmeme
22 May 2026
Model Releases

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

DGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

DGX agent

arXiv:2605.20690v1 Announce Type: new Abstract: Agentic discovery has shown that LLM-driven search can find novel algorithms, designs, and code under benchmark conditions. Translating the paradigm to

model-releasesarxiv-cs-ai
22 May 2026
Agents

Great conversation on the Max Agency podcast with @cogent_security Co-Founder + CTO Geng Sng on building agents for autonomous cyber defense…

DGX agent

Great conversation on the Max Agency podcast with @cogent_security Co-Founder + CTO Geng Sng on building agents for autonomous cyber defense. Check out the full episode: ⏯️ YouTube: https://youtu.be/D

agentsharrison-chase--x
22 May 2026
Agents

im going to be in NYC in ~1 week, and am doing a fireside chat with one of the top agent companies in new york - @_anish_agarwal of @travers…

DGX agent

im going to be in NYC in ~1 week, and am doing a fireside chat with one of the top agent companies in new york - @_anish_agarwal of @traversal_ai come join us! https://partiful.com/e/MLzglw6al1KrUSXSr

agentsharrison-chase--x
22 May 2026
Agents

MARS: Modular Agent with Reflective Search for Automated AI Research

DGX agent

arXiv:2602.02660v3 Announce Type: replace Abstract: A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general softw

agentsarxiv-cs-ai
22 May 2026
Model Releases

APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents

DGX agent

arXiv:2605.21240v1 Announce Type: new Abstract: LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision mak

model-releasesarxiv-cs-lg
21 May 2026
Agents

Grok Build 0.1 is now available for early access in Hermes Agent

DGX agent

Nous Research has released Grok Build 0.1 as an early access feature within the Hermes Agent platform. This release likely introduces initial capabilities for building and customizing Grok-based appli

agentsnous-research--x
21 May 2026
Agents

Here's more about Datasette Agent on my blog: https://simonwillison.net/2026/May/21/datasette-agent/ - and the announcement on the new Datas…

DGX agent

Here's more about Datasette Agent on my blog: https://simonwillison.net/2026/May/21/datasette-agent/ - and the announcement on the new Datasette project blog too: https://datasette.io/blog/2026/datase

agentssimon-willison--x
21 May 2026
Agents

I don’t know why people talk in absolutes about agentic design. Especially for coders and builders. We just invented this stuff. It’s wild t…

DGX agent

I don’t know why people talk in absolutes about agentic design. Especially for coders and builders. We just invented this stuff. It’s wild to think it won’t radically change over the next year. We’re

agentsyohei-nakajima--x
21 May 2026
Agents

Replit shipped 5 new security tools in the last month: Security Agent, Auto-Protect, Security Center V2, External Access Tokens, and Private…

DGX agent

Replit shipped 5 new security tools in the last month: Security Agent, Auto-Protect, Security Center V2, External Access Tokens, and Private Publishing. 🛡️ Tomorrow we go live with @AlexandreCuoci, wh

agentsreplit--x
21 May 2026
Model Releases

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

DGX agent

arXiv:2605.21384v1 Announce Type: cross Abstract: As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Re

model-releasesarxiv-cs-cl
21 May 2026
Agents

> These engineers can review their agent's code much faster than reviewing human code. wat

DGX agent

> These engineers can review their agent's code much faster than reviewing human code. wat Today we reduced headcount by 22%. The business is the strongest it's ever been. So I think it's important to

agentsjeremy-howard--x
21 May 2026
Research

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

DGX agent

arXiv:2605.20291v1 Announce Type: new Abstract: Large language models (LLMs) have enabled web agents that follow natural language goals through multi-step browser interactions. However, agents fine-tu

researcharxiv-cs-lg
21 May 2026
Agents

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

DGX agent

arXiv:2602.11499v2 Announce Type: replace Abstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Ope

agentsarxiv-cs-cv
21 May 2026
Model Releases

What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

DGX agent

arXiv:2605.21404v1 Announce Type: new Abstract: We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was ru

model-releasesarxiv-cs-lg
21 May 2026
Agents

working on a 'take this vibecoded slop app and make it a production-ready, e2e tested, maintainable, parallelizable agent repo' skill. this …

DGX agent

working on a 'take this vibecoded slop app and make it a production-ready, e2e tested, maintainable, parallelizable agent repo' skill. this thing ran for ~16 hours yesterday and made 103 commits all t

agentsswyx--x
21 May 2026
Agents

A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation

DGX agent

arXiv:2605.19316v1 Announce Type: new Abstract: Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting

agentsarxiv-cs-cl
20 May 2026
Model Releases

Add a Specialized Deep Research Skill to Agent Harnesses

DGX agent

This article covers the NVIDIA AI-Q Blueprint for building specialized deep research agents that empower AI systems to gather context, synthesize information, and support complex decision-making acros

model-releasesnvidia-developer
20 May 2026
Applications

as agents run for longer and with more context, there's a lot of tricky things you have to consider for deployment! super excited to dive in…

DGX agent

as agents run for longer and with more context, there's a lot of tricky things you have to consider for deployment! super excited to dive into deploying Deep Agents for this Boston tech week meetup! .

applicationsharrison-chase--x
20 May 2026
Agents

Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers question…

DGX agent

Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers questions with exact citations highlighted on the original PDF page.

agentsllamaindex--x
20 May 2026
Model Releases

OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

DGX agent

arXiv:2605.18758v1 Announce Type: cross Abstract: Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction rout

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

DGX agent

arXiv:2605.18805v1 Announce Type: cross Abstract: LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet ex

model-releasesarxiv-cs-ai
20 May 2026
Agents

Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from e…

DGX agent

Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from experimentation to production. Sovereign AI. For all. Downloa

agentscohere--x
20 May 2026
Agents

the core concept is a graph that represents everything about the agents knowledge, history, behaviors, capabilities graph is made of events …

DGX agent

the core concept is a graph that represents everything about the agents knowledge, history, behaviors, capabilities graph is made of events behaviors react to graph changes relationships can carry beh

agentsyohei-nakajima--x
20 May 2026
Agents

When an agent is allowed to decompose a goal into smaller sub-tasks, it frequently suffers from goal drift. Left unchecked, it will redefine…

DGX agent

When an agent is allowed to decompose a goal into smaller sub-tasks, it frequently suffers from goal drift. Left unchecked, it will redefine the optimization metric to favor a simpler, useless sub-tas

agentsfrancois-chollet--x
20 May 2026
Agents

working on some hex dashboards -- this is a bummer! if you use langchain/deepagents as your agent harness, you can swap models with fallback…

DGX agent

working on some hex dashboards -- this is a bummer! if you use langchain/deepagents as your agent harness, you can swap models with fallback middleware when your first choice provider is down https://

agentsharrison-chase--x
20 May 2026
Agents

You can also now create automations without any attached repos, like a daily Slack digest agent that prioritizes unread threads and DMs for …

DGX agent

You can also now create automations without any attached repos, like a daily Slack digest agent that prioritizes unread threads and DMs for your response. Find more templates in our marketplace: http:

agentscursor--x
20 May 2026
Agents

Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments

DGX agent

arXiv:2605.16419v1 Announce Type: cross Abstract: Kinematic monitoring plays a critical role in long-term rehabilitation for patients with spinal cord injury (SCI), where multi-view markerless motion

agentsarxiv-cs-ai
19 May 2026
Model Releases

Announcing Claude Managed Agents on Cloudflare

DGX agent

Cloudflare has integrated with Anthropic's Claude Managed Agents to provide a fast, isolated execution environment for autonomous code delivery. This means builders can scale agent workflows globally

model-releasescloudflare-ai
19 May 2026
Model Releases

ClawArena: Benchmarking AI Agents in Evolving Information Environments

DGX agent

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse

DGX agent

arXiv:2605.17450v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to ins

model-releasesarxiv-cs-ai
19 May 2026
Agents

Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a direct swap to o…

DGX agent

Curious finding while creating evals and benchmarks for long-horizon (100+ turn) agents While it’s generally thought that a direct swap to open source models can bring immediate cost savings, that’s n

agentsharrison-chase--x
19 May 2026
Agents

Cursor is now available in Jira. Assign Cursor to work items, or mention @​Cursor in a comment to kick off a cloud agent. Cursor uses the ti…

DGX agent

Cursor is now available in Jira. Assign Cursor to work items, or mention @​Cursor in a comment to kick off a cloud agent. Cursor uses the title, description, comments, and your team's repository setti

agentscursor--x
19 May 2026
Agents

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

DGX agent

arXiv:2605.17439v1 Announce Type: cross Abstract: Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-lev

agentsarxiv-cs-ai
19 May 2026
Agents

感谢 @dotey 将他的 baoyu-comic skill 移植到 Hermes Agent! 这个 skill 可以把一段 prompt 或源文档转成多页知识漫画,支持 6 种风格、7 种语气、7 种版式和 5 个预设。 http://github.com/NousRese…

DGX agent

感谢 @dotey 将他的 baoyu-comic skill 移植到 Hermes Agent! 这个 skill 可以把一段 prompt 或源文档转成多页知识漫画,支持 6 种风格、7 种语气、7 种版式和 5 个预设。 http://github.com/NousResearch/hermes-agent/tree/main/skills/creative/baoyu-comic Medi

agentsnous-research--x
19 May 2026
Model Releases

Evaluating Cognitive Age Alignment in Interactive AI Agents

DGX agent

arXiv:2605.17894v1 Announce Type: new Abstract: While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across doma

model-releasesarxiv-cs-ai
19 May 2026
Safety

Natural-Language Agent Harnesses

DGX agent

arXiv:2603.25723v2 Announce Type: replace-cross Abstract: Agent performance is strongly shaped by the surrounding harness: the external execution system around a model that organizes a task run. Yet t

safetyarxiv-cs-ai
19 May 2026
Safety

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

DGX agent

arXiv:2605.18672v1 Announce Type: new Abstract: This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for

safetyarxiv-cs-ai
19 May 2026
← Previous
1…105106107108109…375
Next →