AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,975 results
Safety

Secure AI agents with Policy and Lambda interceptors in Amazon Bedrock AgentCore gateway

DGX agent

In this post, we use a lakehouse data agent to demonstrate how you can use Policy for deterministic access control and Lambda interceptors for dynamic validation. We then show how to combine Lambda in

safetyaws-ml-blog
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study

DGX agent

arXiv:2605.31408v1 Announce Type: cross Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granulari

model-releasesarxiv-cs-ai
1 Jun 2026
Applications

The gains are most obvious in coding and operational areas. Even conservative organizations are adopting coding agents fast. That doesn't me…

DGX agent

The gains are most obvious in coding and operational areas. Even conservative organizations are adopting coding agents fast. That doesn't mean AI adoption doesn't come with huge challenges and cost co

applicationsethan-mollick--x
1 Jun 2026
Hardware

We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with t…

DGX agent

We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with the new OpenShell runtime, which connects Hermes to @Microsof

hardwarenous-research--x
1 Jun 2026
Applications

/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to kn…

DGX agent

/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to know when to ask you GOOD questions, maybe because it is stuck

applicationsethan-mollick--x
31 May 2026
Model Releases

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

DGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

model-releasesclem-delangue--x
31 May 2026
Safety

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

DGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

safetyarxiv-cs-ai
29 May 2026
Applications

Deep Agents v0.6 makes harness profiles a first-class abstraction. Now, you can get production-grade performance from models like @Kimi_Moon…

DGX agent

Deep Agents v0.6 makes harness profiles a first-class abstraction. Now, you can get production-grade performance from models like @Kimi_Moonshot, @Alibaba_Qwen, and @DeepSeek_ai at 20x+ lower cost tha

applicationsharrison-chase--x
29 May 2026
Safety

GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering

DGX agent

arXiv:2605.29584v1 Announce Type: new Abstract: Reinforcement learning (RL) is a natural fit for agentic knowledge base question answering (KBQA), where a model must issue executable actions, observe

safetyarxiv-cs-cl
29 May 2026
Model Releases

GroundAct: Can LLM Agents Ground Actions in Environmental States?

DGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

model-releasesarxiv-cs-ai
29 May 2026
Applications

I watched the LangChain keynote expecting product releases. Got something better — the clearest breakdown I've heard of why most agents neve…

DGX agent

I watched the LangChain keynote expecting product releases. Got something better — the clearest breakdown I've heard of why most agents never make it past demo. One thesis lands hard. Here's what sepa

applicationsharrison-chase--x
29 May 2026
Model Releases

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

DGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

model-releasesarxiv-cs-ai
29 May 2026
Local Ai

Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents

DGX agent

arXiv:2605.30335v1 Announce Type: new Abstract: Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic p

local-aiarxiv-cs-ai
29 May 2026
Safety

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

DGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

safetyarxiv-cs-cl
29 May 2026
Model Releases

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days…

DGX agent

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days shipped in 13. One PR delivered 21 endpoints at 100% test c

model-releasesboris-cherny--x
29 May 2026
Model Releases

Scaling Laws for Agent Harnesses via Effective Feedback Compute

DGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow

DGX agent

arXiv:2605.29368v1 Announce Type: cross Abstract: The intricate nature of modern surgical care necessitates intelligent systems that can synthesize extensive patient records, support collaborative dec

model-releasesarxiv-cs-ai
29 May 2026
Tools

As agents use more context, input tokens have become the majority of price-equivalent token costs.

DGX agent

As AI agents handle increasingly complex tasks with extended context windows, the cost structure of LLM usage has shifted such that input tokens now represent the majority of price-equivalent expenses

toolscursor--x
28 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Safety

CPPO: Contrastive Perception Policy Optimization for VLM Agents

DGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

safetyarxiv-cs-cv
28 May 2026
Safety

Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles

DGX agent

arXiv:2605.27784v1 Announce Type: new Abstract: LLM agents are governed by long-lived natural-language prompt policies, but individually reasonable standing rules can interact in uninspected ways. We

safetyarxiv-cs-ai
28 May 2026
Safety

Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

DGX agent

arXiv:2605.28598v1 Announce Type: cross Abstract: LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has

safetyarxiv-cs-ai
28 May 2026
Model Releases

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills

DGX agent

arXiv:2604.05333v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on reusable skills, and as they interact with personal applications, web browsers, and other interfaces, skill l

model-releasesarxiv-cs-ai
28 May 2026
Safety

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

DGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

safetyarxiv-cs-ro
28 May 2026
Safety

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

DGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

safetyarxiv-cs-cl
28 May 2026
Model Releases

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

DGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

DGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

SHIPPED. Mistral Vibe is now the AI agent for long-horizon productivity and coding, and the home for Work mode, Code mode, the CLI, and a br…

DGX agent

Mistral AI has released Mistral Vibe, an AI agent designed for long-horizon productivity and coding tasks, featuring Work mode, Code mode, a CLI, and additional capabilities. The product consolidates

model-releasesmistral-ai--x
28 May 2026
Safety

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

DGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

safetyarxiv-cs-ai
28 May 2026
Model Releases

A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing (Steven Levy/Wired)

DGX agent

Steven Levy / Wired: A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing — The definitive stor

model-releasestechmeme
27 May 2026
Model Releases

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

DGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

model-releasesarxiv-cs-ai
27 May 2026
Safety

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

DGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

safetyarxiv-cs-cl
27 May 2026
Industry

Building AI agents for business support using Amazon Bedrock AgentCore

DGX agent

In this post, we share how the AWS Generative AI Innovation Center (GenAIIC) collaborated with Works Human Intelligence (WHI) to build two AI agents using Amazon Bedrock AgentCore. We discuss the chal

industryaws-ml-blog
27 May 2026
Safety

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

DGX agent

arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counte

safetyarxiv-cs-ai
27 May 2026
Safety

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

DGX agent

arXiv:2605.26926v1 Announce Type: new Abstract: Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the

safetyarxiv-cs-ai
27 May 2026
Industry

How AI is starting to dismantle the hegemony of the Big Four consultancies and other large firms, as AI agents help smaller consultancies handle big workloads (Financial Times)

DGX agent

Financial Times: How AI is starting to dismantle the hegemony of the Big Four consultancies and other large firms, as AI agents help smaller consultancies handle big workloads — The technology opens t

industrytechmeme
27 May 2026
Tools

In many cases, I stopped answering my agents. Instead, I give them my criteria and let them answer

DGX agent

A manager describes shifting their leadership approach by providing agents (team members) with decision-making criteria rather than directly answering their questions, enabling them to develop problem

toolsitamar-friedman--x
27 May 2026
Model Releases

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

DGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Multi-Agent Causal Discovery Using Large Language Models

DGX agent

arXiv:2407.15073v4 Announce Type: replace Abstract: Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical c

model-releasesarxiv-cs-ai
27 May 2026
Research

Natural Language Query to Configuration for Retrieval Agents

DGX agent

arXiv:2605.27361v1 Announce Type: new Abstract: Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping

researcharxiv-cs-ai
27 May 2026
Local Ai

New skill from K-Dense: LiteParse in Scientific Agent Skills — built for fast, local research paper ingestion. Your AI co-scientist can now:…

DGX agent

New skill from K-Dense: LiteParse in Scientific Agent Skills — built for fast, local research paper ingestion. Your AI co-scientist can now: * Parse PDFs and supplementary files on your machine (no do

local-aijerry-liu--x
27 May 2026
Hardware

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

DGX agent

arXiv:2605.26720v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planni

hardwarearxiv-cs-ai
27 May 2026
Model Releases

Automated Benchmark Auditing for AI Agents and Large Language Models

DGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Code2UML: Agentic LLMs with context engineering for scalable software visualization

DGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

DGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

model-releasesarxiv-cs-cl
26 May 2026
Safety

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

DGX agent

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings

safetyarxiv-cs-ai
26 May 2026
Tutorials

everyone is talking about self-optimizing loops in software & agents. but what does that actually mean? in my mind, it's a system that obser…

DGX agent

everyone is talking about self-optimizing loops in software & agents. but what does that actually mean? in my mind, it's a system that observes it's own outputs, evaluates them, and uses that signal t

tutorialsharrison-chase--x
26 May 2026
← Previous
1…164165166167168…375
Next →