AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
19 May 2026

PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments

Model ReleasesDGX agent

arXiv:2603.23231v2 Announce Type: replace Abstract: Empowering large language models with long-term memory is crucial for building agents that adapt to users' evolving needs. Existing evaluations of t

The official Pinecone plugin for @cursor_ai has been released! To install, run /add-plugin pinecone within Cursor. What you get: Agent Skill…

Model ReleasesDGX agent

The official Pinecone plugin for @cursor_ai has been released! To install, run /add-plugin pinecone within Cursor. What you get: Agent Skills and scripts to help you build with our vector database and

The Scaling Laws of Skills in LLM Agent Systems

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.16508v1 Announce Type: cross Abstract: As agent systems scale, skills accumulate into large reusable libraries, yet their scaling laws remain poorly understood. Across 15 frontier LLMs, 1,1

TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks

HardwareDGX agent

arXiv:2605.17170v1 Announce Type: new Abstract: Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context process

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

What’s new in Unity AI Gateway: service policies, guardrails, observability, and cost controls for AI agents and MCPs

ApplicationsDGX agent

Unity AI Gateway introduces new enterprise features including service policies, guardrails, observability tools, and cost controls designed to manage and monitor AI agents and Model Context Protocols

18 May 2026

Belief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation

Model ReleasesDGX agent

arXiv:2605.15343v1 Announce Type: new Abstract: LLM-based agents are increasingly used to simulate deliberative interactions such as negotiation, conflict resolution, and multi-turn opinion exchange.

continuing my HTML era, I had so much fun talking with Claire at Code w/ Claude about staying in the loop with long running agents

Model ReleasesDGX agent

continuing my HTML era, I had so much fun talking with Claire at Code w/ Claude about staying in the loop with long running agents Soooo @trq212 has straight up changed my life with these 5 words: 'HT

DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory

Model ReleasesDGX agent

arXiv:2605.15759v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a

From I/O to Code with Discovery Agent

SafetyDGX agent

arXiv:2605.15334v1 Announce Type: cross Abstract: The automatic synthesis of a program from any form of specification is regarded as a holy grail of computer science. Fueled by LLMs, NL2Code has achie

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Model ReleasesDGX agent

arXiv:2605.15229v1 Announce Type: cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a des

RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

Model ReleasesDGX agent

arXiv:2605.15846v1 Announce Type: cross Abstract: Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision

Model ReleasesDGX agent

arXiv:2605.15537v1 Announce Type: new Abstract: This paper introduces RTL-BenchMT, an agentic framework for dynamically maintaining RTL generation benchmarks. Large Language Models (LLMs) assisted aut

SDOF: Taming the Alignment Tax in Multi-Agent Orchestration with State-Constrained Dispatch

Model ReleasesDGX agent

arXiv:2605.15204v1 Announce Type: new Abstract: Multi-agent orchestration frameworks such as LangChain, LangGraph, and CrewAI route tasks through graph-based pipelines but do not enforce the stage con

17 May 2026

🚨Breaking new study: memory in LLM agents still can’t really be trusted, even after over trillion dollars has gone into the development of …

SafetyDGX agent

🚨Breaking new study: memory in LLM agents still can’t really be trusted, even after over trillion dollars has gone into the development of the field. Excited to share our new paper: “Useful Memories B

stateful agents, decision traces, context graphs… talked about a lot, but has anyone seen an elegant primitive around how to actually implem…

TutorialsDGX agent

Yohei Nakajima discusses the lack of elegant primitive implementations for stateful agents, decision traces, and context graphs, which are frequently discussed concepts in AI but rarely demonstrated i

15 May 2026

40 Grok Build agents tearing through C code in parallel. All supervised by DAD. DAD is a lightweight autonomous tmux supervisor for long-run…

Model ReleasesDGX agent

40 Grok Build agents tearing through C code in parallel. All supervised by DAD. DAD is a lightweight autonomous tmux supervisor for long-running Grok tasks. Built entirely in Grok Build. /dad 'your ob

A Minimal Agent for Automated Theorem Proving

Model ReleasesDGX agent

arXiv:2602.24273v3 Announce Type: replace Abstract: We propose a minimal agentic baseline that enables systematic comparison across different AI-based theorem prover architectures. This design impleme

A Security Analysis of the OpenClaw AI Agent Framework

SafetyDGX agent

arXiv:2603.27517v3 Announce Type: replace-cross Abstract: AI agent frameworks connecting large language model (LLM) reasoning to host execution surfaces -- shell, filesystem, containers, and messaging

AI agents are turning SaaS applications into headless, deterministic engines

ApplicationsDGX agent

SaaS applications are being transformed into deterministic engines as AI agents replace traditional user interfaces and drive the rise of the headless enterprise — where AI, not humans, serves as the

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

SafetyDGX agent

arXiv:2605.11611v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG)

Do Coding Agents Understand Least-Privilege Authorization?

Model ReleasesDGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents

SafetyDGX agent

arXiv:2605.14057v1 Announce Type: new Abstract: Most existing dialogue systems are user-driven, primarily designed to fulfill user requests. However, in many critical real-world scenarios, a conversat

From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents

Model ReleasesDGX agent

arXiv:2605.14034v1 Announce Type: new Abstract: Wide applications of LLM-based agents require strong alignment with human social values. However, current works still exhibit deficiencies in self-cogni

GraphFlow: An Architecture for Formally Verifiable Visual Workflows Enabling Reliable Agentic AI Automation

Local AiDGX agent

arXiv:2605.14968v1 Announce Type: new Abstract: GraphFlow is a visual workflow system designed to improve the reliability of agentic AI automation in multi-step, mission-critical processes. In these w

Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2511.16964v2 Announce Type: replace-cross Abstract: Maximizing performance on available GPU hardware is an ongoing challenge for modern AI inference systems. Traditional approaches include writi

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

AgentsDGX agent

arXiv:2605.14113v1 Announce Type: cross Abstract: While interpretable prototype networks offer compelling case-based reasoning for clinical diagnostics, their raw continuous outputs lack the semantic

SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration

SafetyDGX agent

arXiv:2605.14089v1 Announce Type: new Abstract: In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existi

xAI launches Grok Build, an agentic CLI for coding, building apps, and automating workflows, in beta for SuperGrok Heavy subscribers (Carmen Arroyo/Bloomberg)

Model ReleasesDGX agent

Carmen Arroyo / Bloomberg: xAI launches Grok Build, an agentic CLI for coding, building apps, and automating workflows, in beta for SuperGrok Heavy subscribers — Elon Musk's xAI is rolling out its fir

14 May 2026

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

Model ReleasesDGX agent

arXiv:2605.13841v1 Announce Type: cross Abstract: Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applica

GAAMA: Graph Augmented Associative Memory for Agents

ResearchDGX agent

arXiv:2603.27910v2 Announce Type: replace Abstract: AI agents that interact with users across multiple sessions require persistent long-term memory to maintain coherent, personalized behavior. Current

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

SafetyDGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

SafetyDGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

Macro-Action Based Multi-Agent Instruction Following through Value Cancellation

SafetyDGX agent

arXiv:2605.12655v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) in real-world use cases may need to adapt to external natural language instructions that interrupt ongoing beh

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

Model ReleasesDGX agent

arXiv:2605.13228v1 Announce Type: cross Abstract: Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and co

Will be giving a talk titled “You should do RL for long-running agents (and use RLMs)” at 4pm on Sat at AI Engineer Singapore. Excited to se…

ToolsDGX agent

Swyx is giving a talk at AI Engineer Singapore on Saturday at 4pm about using reinforcement learning for long-running agents and reinforcement learning models (RLMs), exploring why this approach shoul

xAI just released Grok Build CLI and it’s a game changer for developers Grok Build is a powerful AI coding agent and CLI built for professio…

Model ReleasesDGX agent

xAI just released Grok Build CLI and it’s a game changer for developers Grok Build is a powerful AI coding agent and CLI built for professional software engineering and complex coding workflows runnin

13 May 2026

Beyond Manual Curation: Augmenting Targeted Protein Degradation Databases via Agentic Literature Extraction Workflows

AgentsDGX agent

arXiv:2605.11221v1 Announce Type: cross Abstract: Predictive models in biomedicine depend on structured assay data locked in the text, tables, and supplements of primary publications. This bottleneck

Blumira launches Kindling pilot, an agentic SIEM investigation engine that cuts alert volume up to 50x

Model ReleasesDGX agent

Security operations platform startup Blumira Inc. today launched the pilot of Kindling, an agentic security information and event management investigation engine that the company says can reduce alert

Correcting Selection Bias in Sparse User Feedback for Large Language Model Quality Estimation: A Multi-Agent Hierarchical Bayesian Approach

Model ReleasesDGX agent

arXiv:2605.12177v1 Announce Type: new Abstract: [Abridged] Production LLM deployments receive feedback from a non-random fraction of users: thumbs sit mostly in the tails of the satisfaction distribut

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

SafetyDGX agent

arXiv:2605.11853v1 Announce Type: cross Abstract: Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that pr

JACoP: Joint Alignment for Compliant Multi-Agent Prediction

SafetyDGX agent

arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e

Poolside is hosting a 2-day model research hackathon in London. Join us to push an open-weight agent model as far as you can. RL and fine-tu…

Local AiDGX agent

Poolside is hosting a 2-day model research hackathon in London. Join us to push an open-weight agent model as far as you can. RL and fine-tune Laguna XS.2, our latest-generation model, on Prime Intell

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

ResearchDGX agent

arXiv:2605.11212v1 Announce Type: new Abstract: Computer-use agents~(CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual toke

Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: - Claude Agent …

Model ReleasesDGX agent

Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: - Claude Agent SDK - claude -p - Claude Code GitHub Actions - Third-party a

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

Model ReleasesDGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

12 May 2026

Agentic MIP Research: Accelerated Constraint Handler Generation

Model ReleasesDGX agent

arXiv:2605.09186v1 Announce Type: new Abstract: Mixed-integer programming (MIP) research is both mathematically sophisticated and engineering-intensive: testing an algorithmic hypothesis within a bran

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

Model ReleasesDGX agent

arXiv:2604.08178v2 Announce Type: replace Abstract: In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2605.08271v1 Announce Type: cross Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For

CARL: Criticality-Aware Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2512.04949v3 Announce Type: replace-cross Abstract: Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction.

CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents

AgentsDGX agent

arXiv:2605.08399v1 Announce Type: new Abstract: Tool-augmented language models can extend small language models with external executable skills, but scaling the tool library creates a coupled challeng

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

Model ReleasesDGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

Model ReleasesDGX agent

arXiv:2605.08112v1 Announce Type: cross Abstract: AI coding agents powered by large language models can read codebases and produce functional code, but they routinely violate team-specific product dec

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

SafetyDGX agent

arXiv:2605.10366v1 Announce Type: new Abstract: Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text,

Idira launches as Palo Alto Networks extends CyberArk tech to machine and agentic identities

Model ReleasesDGX agent

Palo Alto Networks Inc. today launched Idira, a new identity security platform designed to manage human, machine and artificial intelligence agent identities across the enterprise under a single privi

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

Model ReleasesDGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

SafetyDGX agent

arXiv:2511.01008v2 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current method

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

Model ReleasesDGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA o…

Model ReleasesDGX agent

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics

← Previous
1…119120121122123…300
Next →