AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
Agents

ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review

DGX agent

arXiv:2601.22638v2 Announce Type: replace-cross Abstract: The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for

agentsarxiv-cs-ai
12 May 2026
Agents
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Simulus: Combining Improvements in Sample-Efficient World Model Agents

DGX agent

arXiv:2502.11537v4 Announce Type: replace-cross Abstract: World models (WMs) represent the frontier of sample-efficient reinforcement learning, but their complexity leaves many promising improvements

agentsarxiv-cs-ai
12 May 2026
Safety

The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions

DGX agent

arXiv:2605.10698v1 Announce Type: cross Abstract: Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that

safetyarxiv-cs-ai
12 May 2026
Model Releases

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

DGX agent

arXiv:2605.09330v1 Announce Type: cross Abstract: Agentic memory enables LLMs to persist information beyond a single context window and reuse it in later decisions, but it also introduces a new vulner

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

We integrated FrontierCS into Harbor and are releasing a preview long-horizon agent leaderboard (up to 835 turns, ~200K output tokens) with …

DGX agent

We integrated FrontierCS into Harbor and are releasing a preview long-horizon agent leaderboard (up to 835 turns, ~200K output tokens) with Kimi K2.6 @Kimi_Moonshot (score 46.9) and Claude Code Opus 4

model-releaseskimi-moonshot--x
12 May 2026
Agents

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

DGX agent

arXiv:2602.06286v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of

agentsarxiv-cs-ai
12 May 2026
Agents

AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation

DGX agent

arXiv:2602.04672v3 Announce Type: replace Abstract: Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic

agentsarxiv-cs-cv
11 May 2026
Model Releases

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

DGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

model-releasesarxiv-cs-cl
11 May 2026
Agents

ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms

DGX agent

arXiv:2512.03476v2 Announce Type: replace-cross Abstract: Bridging the gap between theoretical conceptualization and computational implementation is a major bottleneck in Scientific Computing (SciC) a

agentsarxiv-cs-ai
11 May 2026
Tutorials

Building web search-enabled agents with Strands and Exa

DGX agent

In this post, you will learn how to set up the Exa integration in Strands Agents, understand the two core tools it exposes, and walk through real-world use cases that show how agents use web search to

tutorialsaws-ml-blog
11 May 2026
Local Ai

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

DGX agent

arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as

local-aiarxiv-cs-cv
11 May 2026
Model Releases

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

DGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

model-releasesjerry-liu--x
11 May 2026
Safety

From Specification to Deployment: Empirical Evidence from a W3C VC + DID Trust Infrastructure for Autonomous Agents

DGX agent

arXiv:2605.06738v1 Announce Type: cross Abstract: Autonomous AI agents now transact at production scale -- 69,000 bots executing 165 million transactions across 50 million USDC in cumulative volume on

safetyarxiv-cs-ai
11 May 2026
Model Releases

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

DGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

DGX agent

arXiv:2605.07510v1 Announce Type: cross Abstract: Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input

model-releasesarxiv-cs-cl
11 May 2026
Safety

Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization

DGX agent

arXiv:2605.06864v1 Announce Type: new Abstract: We study multi-objective multi-agent multi-armed bandits (MO-MA-MAB) under stochastic rewards, where agents observe heterogeneous reward vectors and com

safetyarxiv-cs-lg
11 May 2026
Safety

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

DGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

safetyarxiv-cs-ai
11 May 2026
Model Releases

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

DGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

DGX agent

arXiv:2605.07001v1 Announce Type: cross Abstract: Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reaso

model-releasesarxiv-cs-cl
11 May 2026
Agents

Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent

DGX agent

arXiv:2605.07202v1 Announce Type: new Abstract: Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limi

agentsarxiv-cs-ai
11 May 2026
Model Releases

Deep Agents Deploy! https://docs.langchain.com/oss/python/deepagents/deploy

DGX agent

Deep Agents Deploy! https://docs.langchain.com/oss/python/deepagents/deploy Claude Managed Agents is really good But we need an open source solution Haven’t found any options so might need to build my

model-releasesharrison-chase--x
10 May 2026
Model Releases

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safe…

DGX agent

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safety. I'm devastated to inform doomers that 'full stack open s

model-releasesclem-delangue--x
8 May 2026
Model Releases

Agent harnesses have an expiration date

DGX agent

A benchmark-driven look at why agent harnesses need adaptive finish logic as model behavior changes across Claude, GPT-4o, and Gemma. The post Agent harnesses have an expiration date appeared first on

model-releasesarize-ai
7 May 2026
Agents

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

DGX agent

arXiv:2605.03042v1 Announce Type: cross Abstract: This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance me

agentsarxiv-cs-ai
7 May 2026
Agents

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

DGX agent

arXiv:2605.04304v1 Announce Type: cross Abstract: Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While ex

agentsarxiv-cs-cl
7 May 2026
Model Releases

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundred…

DGX agent

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundreds of tools, not your simple coding agent stuff) and to my su

model-releasesclem-delangue--x
7 May 2026
Agents

Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent

DGX agent

arXiv:2602.19837v3 Announce Type: replace-cross Abstract: Humans are highly effective at utilizing prior knowledge to adapt to novel tasks, a capability that standard machine learning models struggle

agentsarxiv-cs-lg
7 May 2026
Model Releases

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

DGX agent

arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The chal

model-releasesarxiv-cs-ai
7 May 2026
Applications

Parloa builds service agents customers want to talk to

DGX agent

Parloa, an AI company, has developed service agents powered by OpenAI's technology that are designed to provide natural, conversational customer interactions. These agents aim to improve customer expe

applicationsopenai
7 May 2026
Safety

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

DGX agent

arXiv:2604.03976v2 Announce Type: replace Abstract: Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI syste

safetyarxiv-cs-ai
7 May 2026
Agents

SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment

DGX agent

arXiv:2605.04012v1 Announce Type: new Abstract: Language models excel at diagnostic assessments on currated medical case-studies and vignettes, performing on par with, or better than, clinical profess

agentsarxiv-cs-ai
7 May 2026
Local Ai

The Hive Mind is a Single Reinforcement Learning Agent

DGX agent

arXiv:2410.17517v5 Announce Type: replace-cross Abstract: Decision-making is an essential attribute of any intelligent agent or group. Natural systems are known to converge to effective strategies thr

local-aiarxiv-cs-ai
7 May 2026
Agents

A Compound AI Agent for Conversational Grant Discovery

DGX agent

arXiv:2605.02366v1 Announce Type: new Abstract: Research funding discovery remains fundamentally fragmented: researchers navigate disparate agency portals (e.g., in the United States, NSF, NIH, DARPA,

agentsarxiv-cs-ai
6 May 2026
Agents

Adobe introduces productivity agent to transform PDF creation and sharing

DGX agent

Adobe Inc. today introduced an artificial intelligence experience embedded in its ubiquitous PDF document reader and creator, Acrobat, to transform how people create, understand and share information.

agentssiliconangle
6 May 2026
Agents

AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development

DGX agent

arXiv:2605.02741v1 Announce Type: cross Abstract: The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of

agentsarxiv-cs-ai
6 May 2026
Agents

Atlassian opens Teamwork Graph and pushes Rovo into agentic execution at Team ’26

DGX agent

Atlassian Corp. today unveiled a sweeping set of artificial intelligence updates at its annual Team ’26 conference, headlined by the broad opening of its Teamwork Graph and the evolution of its Rovo A

agentssiliconangle
6 May 2026
Safety

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

DGX agent

arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing

safetyarxiv-cs-ai
6 May 2026
Model Releases

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files

DGX agent

arXiv:2603.00822v2 Announce Type: replace-cross Abstract: As Large Language Model (LLM) agents increasingly execute complex, autonomous software engineering tasks, developers rely on natural language

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.03310v1 Announce Type: cross Abstract: Multi-agent LLM systems fail in production at rates between 41% and 87%, mostly due to coordination defects rather than base-model capability. Existin

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

DataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis

DGX agent

arXiv:2605.02503v1 Announce Type: new Abstract: Evaluating autonomous data analysis agents requires testing their ability to perform exploratory analysis in underexplored data environments. However, m

model-releasesarxiv-cs-ai
6 May 2026
Agents

EngiAgent: Fully Connected Coordination of LLM Agents for Solving Open-ended Engineering Problems with Feasible Solutions

DGX agent

arXiv:2605.02289v1 Announce Type: new Abstract: Engineering problem solving is central to real-world decision-making, requiring mathematical formulations that not only represent complex problems but a

agentsarxiv-cs-ai
6 May 2026
Agents

Explainable AI for Blind and Low-Vision Users: Navigating Trust, Modality, and Interpretability in the Agentic Era

DGX agent

arXiv:2604.00187v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) is critical for ensuring trust and accountability, yet its development remains predominantly visual.

agentsarxiv-cs-ai
6 May 2026
Agents

Extreme Connect 2026: Agentic AI, Platform ONE and the next phase of enterprise networking

DGX agent

Extreme Networks Inc. used its Extreme Connect 2026 user conference this week to make a strong case that artificial intelligence-driven networking has finally arrived. Building on Platform ONE, the co

agentssiliconangle
6 May 2026
Agents

From Experimental Limits to Physical Insight: A Retrieval-Augmented Multi-Agent Framework for Interpreting Searches Beyond the Standard Model

DGX agent

arXiv:2605.02491v1 Announce Type: cross Abstract: Modern searches for physics beyond the Standard Model produce rapidly expanding literature containing heterogeneous information, including textual ana

agentsarxiv-cs-ai
6 May 2026
Safety

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents

DGX agent

arXiv:2603.00977v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents have recently demonstrated strong capabilities in interactive decision-making, yet they remain fundamentally

safetyarxiv-cs-lg
6 May 2026
Safety

NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science

DGX agent

arXiv:2605.02092v1 Announce Type: new Abstract: The automation of scientific research workflows has emerged as a transformative frontier in artificial intelligence, yet existing autonomous research ag

safetyarxiv-cs-ai
6 May 2026
Model Releases

ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling

DGX agent

arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike

model-releasesarxiv-cs-ai
6 May 2026
Agents

Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-Drones

DGX agent

arXiv:2605.03788v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly explored as high-level reasoning engines for cyber-physical systems, yet their application to real-time

agentsarxiv-cs-ro
6 May 2026
← Previous
1…122123124125126…375
Next →