AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
13 May 2026

Hermes Agent now runs natively on NVIDIA RTX PCs and DGX Spark. Hermes is designed for exactly the kind of always-on workload that NVIDIA's …

Local AiDGX agent

Hermes Agent now runs natively on NVIDIA RTX PCs and DGX Spark. Hermes is designed for exactly the kind of always-on workload that NVIDIA's hardware is built for, and their blog explains in depth why

Learning Agentic Policy from Action Guidance

SafetyDGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

Microsoft’s new agentic security system MDASH uncovers four critical Windows RCE flaws

AgentsDGX agent

Microsoft Corp. today detailed a new artificial intelligence-powered vulnerability discovery system that uncovered 16 previously unknown flaws in Windows networking and authentication components, incl

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

SafetyDGX agent

arXiv:2605.12039v1 Announce Type: new Abstract: Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entr

12 May 2026

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

AgentsDGX agent

arXiv:2605.08956v1 Announce Type: new Abstract: A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they alread

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox

Model ReleasesDGX agent

arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

Model ReleasesDGX agent

arXiv:2605.09998v1 Announce Type: cross Abstract: Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agen

CrackMeBench: Binary Reverse Engineering for Agents

Model ReleasesDGX agent

arXiv:2605.10597v1 Announce Type: cross Abstract: Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-fl

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

SafetyDGX agent

arXiv:2605.08717v1 Announce Type: cross Abstract: Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

Model ReleasesDGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

Model ReleasesDGX agent

arXiv:2605.10332v1 Announce Type: new Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied enviro

Generating synthetic electronic health record data using agent-based models to evaluate machine learning robustness under mass casualty incidents

AgentsDGX agent

arXiv:2605.09951v1 Announce Type: new Abstract: ML models in healthcare are typically evaluated using curated real-world EHR data. A key limitation of such evaluations is that they may fail to assess

HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics

AgentsDGX agent

arXiv:2507.15518v5 Announce Type: replace Abstract: Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large langua

How Mobile World Model Guides GUI Agents?

ResearchDGX agent

arXiv:2605.10347v1 Announce Type: new Abstract: Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable predi

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization

SafetyDGX agent

arXiv:2605.08978v1 Announce Type: new Abstract: Recent advancements in agentic test-time scaling allow models to gather environmental feedback before committing to final actions. A key limitation of e

Log analysis is necessary for credible evaluation of AI agents

Model ReleasesDGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

Manifold scores 7,700 MCP servers in Manifest expansion aimed at agent security teams

AgentsDGX agent

Artificial intelligence detection and response platform startup Manifold Security Inc. today announced an expansion of its Manifest supply chain intelligence tool to cover Model Context Protocol serve

MDGYM: Benchmarking AI Agents on Molecular Simulations

Model ReleasesDGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

Model ReleasesDGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

HardwareDGX agent

arXiv:2605.09055v1 Announce Type: cross Abstract: Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-

Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning

Model ReleasesDGX agent

arXiv:2605.09822v1 Announce Type: cross Abstract: We define Oracle Poisoning, an attack class in which an adversary corrupts a structured knowledge graph that AI agents query at runtime via tool-use p

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

Model ReleasesDGX agent

arXiv:2605.08876v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical f

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Model ReleasesDGX agent

arXiv:2604.01527v3 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fi

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

SafetyDGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review

AgentsDGX agent

arXiv:2601.22638v2 Announce Type: replace-cross Abstract: The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for

Simulus: Combining Improvements in Sample-Efficient World Model Agents

AgentsDGX agent

arXiv:2502.11537v4 Announce Type: replace-cross Abstract: World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements

The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions

SafetyDGX agent

arXiv:2605.10698v1 Announce Type: cross Abstract: Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

Model ReleasesDGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

We integrated FrontierCS into Harbor and are releasing a preview long-horizon agent leaderboard (up to 835 turns, ~200K output tokens) with …

Model ReleasesDGX agent

We integrated FrontierCS into Harbor and are releasing a preview long-horizon agent leaderboard (up to 835 turns, ~200K output tokens) with Kimi K2.6 @Kimi_Moonshot (score 46.9) and Claude Code Opus 4

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

AgentsDGX agent

arXiv:2602.06286v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of

11 May 2026

AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation

AgentsDGX agent

arXiv:2602.04672v3 Announce Type: replace Abstract: Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

Model ReleasesDGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

AgentsDGX agent

arXiv:2512.03476v2 Announce Type: replace-cross Abstract: Bridging the gap between theoretical conceptualization and computational implementation is a major bottleneck in Scientific Computing (SciC) a

Building web search-enabled agents with Strands and Exa

TutorialsDGX agent

In this post, you will learn how to set up the Exa integration in Strands Agents, understand the two core tools it exposes, and walk through real-world use cases that show how agents use web search to

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

Local AiDGX agent

arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

Model ReleasesDGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

From Specification to Deployment: Empirical Evidence from a W3C VC + DID Trust Infrastructure for Autonomous Agents

SafetyDGX agent

arXiv:2605.06738v1 Announce Type: cross Abstract: Autonomous AI agents now transact at production scale -- 69,000 bots executing 165 million transactions across 50 million USDC in cumulative volume on

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Model ReleasesDGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

Model ReleasesDGX agent

arXiv:2605.07510v1 Announce Type: cross Abstract: Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input

Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization

SafetyDGX agent

arXiv:2605.06864v1 Announce Type: new Abstract: We study multi-objective multi-agent multi-armed bandits (MO-MA-MAB) under stochastic rewards, where agents observe heterogeneous reward vectors and com

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

SafetyDGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Model ReleasesDGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

Model ReleasesDGX agent

arXiv:2605.07001v1 Announce Type: cross Abstract: Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reaso

Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent

AgentsDGX agent

arXiv:2605.07202v1 Announce Type: new Abstract: Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limi

10 May 2026

Deep Agents Deploy! https://docs.langchain.com/oss/python/deepagents/deploy

Model ReleasesDGX agent

Deep Agents Deploy! https://docs.langchain.com/oss/python/deepagents/deploy Claude Managed Agents is really good But we need an open source solution Haven’t found any options so might need to build my

8 May 2026

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safe…

Model ReleasesDGX agent

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safety. I'm devastated to inform doomers that 'full stack open s

7 May 2026

Agent harnesses have an expiration date

Model ReleasesDGX agent

A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o, and Gemma. The post Agent harnesses have an expiration date appeared first on

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

AgentsDGX agent

arXiv:2605.03042v1 Announce Type: cross Abstract: This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance me

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

AgentsDGX agent

arXiv:2605.04304v1 Announce Type: cross Abstract: Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While ex

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundred…

Model ReleasesDGX agent

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundreds of tools, not your simple coding agent stuff) and to my su

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent

AgentsDGX agent

arXiv:2602.19837v3 Announce Type: replace-cross Abstract: Humans are highly effective at utilizing prior knowledge to adapt to novel tasks, a capability that standard machine learning models struggle

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

Model ReleasesDGX agent

arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The chal

Parloa builds service agents customers want to talk to

ApplicationsDGX agent

Parloa, an AI company, has developed service agents powered by OpenAI's technology that are designed to provide natural, conversational customer interactions. These agents aim to improve customer expe

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

SafetyDGX agent

arXiv:2604.03976v2 Announce Type: replace Abstract: Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI syste

SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment

AgentsDGX agent

arXiv:2605.04012v1 Announce Type: new Abstract: Language models excel at diagnostic assessments on currated medical case-studies and vignettes, performing on par with, or better than, clinical profess

The Hive Mind is a Single Reinforcement Learning Agent

Local AiDGX agent

arXiv:2410.17517v5 Announce Type: replace-cross Abstract: Decision-making is an essential attribute of any intelligent agent or group. Natural systems are known to converge to effective strategies thr

6 May 2026

A Compound AI Agent for Conversational Grant Discovery

AgentsDGX agent

arXiv:2605.02366v1 Announce Type: new Abstract: Research funding discovery remains fundamentally fragmented: researchers navigate disparate agency portals (e.g., in the United States, NSF, NIH, DARPA,

Adobe introduces productivity agent to transform PDF creation and sharing

AgentsDGX agent

Adobe Inc. today introduced an artificial intelligence experience embedded in its ubiquitous PDF document reader and creator, Acrobat, to transform how people create, understand and share information.

AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development

AgentsDGX agent

arXiv:2605.02741v1 Announce Type: cross Abstract: The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of

Atlassian opens Teamwork Graph and pushes Rovo into agentic execution at Team ’26

AgentsDGX agent

Atlassian Corp. today unveiled a sweeping set of artificial intelligence updates at its annual Team ’26 conference, headlined by the broad opening of its Teamwork Graph and the evolution of its Rovo A

← Previous
1…979899100101…300
Next →