AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
Model Releases

Parallel Context Compaction for Long-Horizon LLM Agent Serving

DGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

model-releasesarxiv-cs-ai
25 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Safety

Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good

DGX agent

arXiv:2605.22995v1 Announce Type: cross Abstract: Agentic AI systems are increasingly proposed for social-good domains, often invoking the United Nations Sustainable Development Goals (SDGs) as a voca

safetyarxiv-cs-ai
25 May 2026
Model Releases

GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving

DGX agent

arXiv:2605.22566v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents demonstrate strong reasoning and execution capabilities on complex tasks when guided by structured instructions,

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

DGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

model-releasesarxiv-cs-ai
22 May 2026
Safety

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

DGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

safetyarxiv-cs-cv
22 May 2026
Model Releases

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

DGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

DGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

model-releasesarxiv-cs-ai
22 May 2026
Agents

Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration

DGX agent

arXiv:2605.20190v1 Announce Type: new Abstract: Iterative industrial design-simulation optimization is bottlenecked by the CAD-CAE semantic gap: translating simulation feedback into valid geometric ed

agentsarxiv-cs-ai
22 May 2026
Applications

1/ At @LangChain’s Interrupt conference last week, one question kept coming up: What happens when AI agents need to spend money? Enterprise …

DGX agent

1/ At @LangChain’s Interrupt conference last week, one question kept coming up: What happens when AI agents need to spend money? Enterprise agents have moved from prototype to production, but payments

applicationsharrison-chase--x
21 May 2026
Model Releases

Build AI agents for business intelligence with Amazon Bedrock AgentCore

DGX agent

In this post, we show you how OPLOG developed three AI agents using the Strands Agents SDK, deployed them to Amazon Bedrock AgentCore, and integrated Amazon Bedrock with Anthropic’s Claude Sonnet and

model-releasesaws-ml-blog
21 May 2026
Industry

Build AI-powered dashboard automation agents with NLP on Amazon Bedrock AgentCore

DGX agent

This solution combines the power of Amazon Bedrock AgentCore, Strands Agents, and Amazon Quick transforms to deliver a secure, scalable, and intelligent system for building and operating AI agents whi

industryaws-ml-blog
21 May 2026
Model Releases

FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents

DGX agent

arXiv:2603.01712v2 Announce Type: replace-cross Abstract: Fine-tuning large language models for vertical domains remains labor-intensive, requiring practitioners to curate data, configure training, an

model-releasesarxiv-cs-lg
21 May 2026
Agents

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

DGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

agentsarxiv-cs-cv
21 May 2026
Model Releases

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

DGX agent

arXiv:2605.19743v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address mul

model-releasesarxiv-cs-ai
20 May 2026
Agents

From Intent to AI Pipelines: A Controlled Agentic Framework for Non-AI Expert Scientists

DGX agent

arXiv:2605.18764v1 Announce Type: cross Abstract: Artificial Intelligence (AI) pipelines have become integral to modern research, supporting fields such as Medical Sciences, Agriculture, and Social Sc

agentsarxiv-cs-ai
20 May 2026
Model Releases

Measuring Safety Alignment Effects in Autonomous Security Agents

DGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + …

DGX agent

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + 1,400+ pre-built MCP tools plug Claude into agent → give it

model-releasesboris-cherny--x
20 May 2026
Safety

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

DGX agent

arXiv:2605.19151v1 Announce Type: new Abstract: We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human ap

safetyarxiv-cs-ai
20 May 2026
Model Releases

Sequential Consensus for Multi-Agent LLM Debates: A Wald-SPRT compute governor with calibration-based failure detection

DGX agent

arXiv:2605.19193v1 Announce Type: new Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on h

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

DGX agent

arXiv:2603.05910v2 Announce Type: replace Abstract: LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Y

model-releasesarxiv-cs-ai
20 May 2026
Agents

Tribal AI lands $10M in seed funding to bring metadata-native agents to the enterprise

DGX agent

Enterprise software veterans have become all too familiar with the gap between flashy artificial intelligence demos and their performance in real-world production environments. The reality is that a l

agentssiliconangle
20 May 2026
Model Releases

Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation

DGX agent

arXiv:2510.03879v3 Announce Type: replace-cross Abstract: Translating C to memory-safe languages, like Rust, prevents critical memory safety vulnerabilities that are prevalent in legacy C software. Ev

model-releasesarxiv-cs-ai
19 May 2026
Safety

AI Agents May Always Fall for Prompt Injections

DGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

safetyarxiv-cs-cl
19 May 2026
Local Ai

Aurora: Unified Video Editing with a Tool-Using Agent

DGX agent

arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref

local-aiarxiv-cs-cv
19 May 2026
Model Releases

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

DGX agent

arXiv:2602.03012v2 Announce Type: replace-cross Abstract: Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing wor

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

DGX agent

arXiv:2605.18421v1 Announce Type: cross Abstract: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agen

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

I/O 2026: Welcome to the agentic Gemini era

DGX agent

At I/O 2026, Google announced that AI is transitioning from something users actively open to a background service that completes tasks automatically. The company introduced Gemini Spark, a new agentic

model-releasesgoogle-ai
19 May 2026
Agents

Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate

DGX agent

arXiv:2601.22297v2 Announce Type: replace Abstract: The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). A

agentsarxiv-cs-cl
19 May 2026
Agents

MA^{2}P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion

DGX agent

arXiv:2605.18572v1 Announce Type: new Abstract: Persuasive dialogue generation plays a vital role in decision-making, negotiation, counseling, and behavior change, yet it remains a challenging problem

agentsarxiv-cs-cl
19 May 2026
Agents

RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation

DGX agent

arXiv:2605.17072v1 Announce Type: new Abstract: Existing LLM-driven knowledge graph (KG) construction methods predominantly employ stateless batch processing pipelines, exhibiting structural deficienc

agentsarxiv-cs-ai
19 May 2026
Model Releases

SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors

DGX agent

arXiv:2605.16626v1 Announce Type: cross Abstract: Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dan

model-releasesarxiv-cs-ai
19 May 2026
Safety

State Contamination in Memory-Augmented LLM Agents

DGX agent

arXiv:2605.16746v1 Announce Type: new Abstract: LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon inter

safetyarxiv-cs-ai
19 May 2026
Model Releases

Supervising the search process produces reliable and generalizable information-seeking agents

DGX agent

arXiv:2502.13957v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming web search by shifting from document ranking to synthesizing answers, and are increasingly deplo

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

DGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

model-releasesarxiv-cs-ai
19 May 2026
Agents

The End of Trust: How Agentic AI Breaks Security Assumptions

DGX agent

arXiv:2605.16436v1 Announce Type: cross Abstract: For decades, the security of digital interaction has rested on an unacknowledged economic constraint. Attackers faced a tradeoff between the fidelity

agentsarxiv-cs-ai
19 May 2026
Model Releases

[video] why we need a new continuity layer for long-running agents (claude did this video! all except the voice which was @elevenlabs)

DGX agent

This video discusses the architectural need for a continuity layer in long-running AI agents, explaining how agents require persistent memory and state management mechanisms to maintain coherence acro

model-releasesyohei-nakajima--x
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

A3D: Agentic AI flow for autonomous Accelerator Design

DGX agent

arXiv:2605.15237v1 Announce Type: cross Abstract: Accelerating applications through the design of hardware accelerators can significantly enhance system performance and energy efficiency. Despite adva

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Argus: Evidence Assembly for Scalable Deep Research Agents

DGX agent

arXiv:2605.16217v1 Announce Type: cross Abstract: Deep research agents have achieved remarkable progress on complex information seeking tasks. Even long ReAct style rollouts explore only a single traj

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Coding agent tracing and evaluation: An open source tool to improve AI coding workflows

DGX agent

Announcing coding harness tracing for observing, evaluating, and improving coding agent workflows across Claude Code, Cursor, Codex, GitHub Copilot, and Gemini CLI. The post Coding agent tracing and e

model-releasesarize-ai
18 May 2026
Model Releases

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

DGX agent

arXiv:2512.00417v5 Announce Type: replace Abstract: This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

FormulaCode: Evaluating Agentic Optimization on Large Codebases

DGX agent

arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to op

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

How do you know your document parser is ready for production? 🤔 Existing benchmarks miss what AI agents actually need. That's the gap Parse…

DGX agent

How do you know your document parser is ready for production? 🤔 Existing benchmarks miss what AI agents actually need. That's the gap ParseBench, the first doc OCR benchmark for AI agents, fills. We'l

model-releasesjerry-liu--x
18 May 2026
Model Releases

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

DGX agent

arXiv:2605.15777v1 Announce Type: new Abstract: Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex envi

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

DGX agent

arXiv:2605.15710v1 Announce Type: new Abstract: Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use ev

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Today in AI Engineering (May 17) • Nous Research ships Hermes Agent v0.14.0: Grok subs, Codex runtime, Windows beta • LangSmith Engine relea…

DGX agent

Today in AI Engineering (May 17) • Nous Research ships Hermes Agent v0.14.0: Grok subs, Codex runtime, Windows beta • LangSmith Engine releases trace issue clustering, drafts PRs and evals from prod t

model-releasesharrison-chase--x
18 May 2026
Safety

Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents

DGX agent

arXiv:2505.11708v3 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) agents are increasingly used to simulate sophisticated cyberattacks, but their decision-making processes remain op

safetyarxiv-cs-lg
18 May 2026
← Previous
1…120121122123124…375
Next →