Provably Secure Agent Guardrail
arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates
Knowledge catalogue
arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates
arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most
arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck
arXiv:2602.01869v3 Announce Type: replace Abstract: LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. Thi
arXiv:2605.30256v1 Announce Type: cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonve
arXiv:2605.27643v1 Announce Type: new Abstract: Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations a
arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain
arXiv:2601.04505v3 Announce Type: replace Abstract: Generating accurate circuit schematics from high-level natural language descriptions remains a persistent challenge in electronic design automation
arXiv:2605.27470v1 Announce Type: cross Abstract: Graph anomaly detection aims to identify anomaly nodes in attributed graphs and plays an important role in real-world applications. However, existing
arXiv:2605.28371v1 Announce Type: new Abstract: Industrial Prognostics and Health Management (PHM) provides a representative case study for a broader challenge in applied machine learning: translating
arXiv:2605.27433v1 Announce Type: cross Abstract: With the increasing complexity of collaboration among various social entities and user demands, the factors affecting the stable development of the da
Enterprise leaders are implementing AI agents across their organizations through strategies that address deployment, governance, and integration challenges. The article from Databricks likely covers b
arXiv:2605.28114v1 Announce Type: new Abstract: As autonomous AI agents are deployed in persistent, interacting networks -- coordinating tasks, routing resources, and accumulating reputational histori
arXiv:2605.27628v1 Announce Type: new Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remai
arXiv:2605.28120v1 Announce Type: cross Abstract: Graph-based Retrieval-Augmented Generation (GraphRAG) advances flat document retrieval by structuring knowledge as relational graphs, enabling more co
arXiv:2605.28046v1 Announce Type: new Abstract: Existing agent memory systems universally follow what we term a Memory-as-Tool paradigm where a single query triggers one-shot retrieval of flat passage
arXiv:2605.27393v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligne
Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context engineering. Together we build an intuition on the strengths a
arXiv:2406.03474v2 Announce Type: replace Abstract: Language-guided autonomous driving requires bridging a large abstraction gap between high-level natural-language instructions and low-level vehicle
arXiv:2605.27240v1 Announce Type: new Abstract: Memory-augmented language agents are increasingly deployed in affective applications such as emotional support, where understanding and responding to us
arXiv:2605.26305v1 Announce Type: new Abstract: This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows. Both systems leverage a hybrid Local Body, Remote
arXiv:2602.22190v2 Announce Type: replace-cross Abstract: Open-source native GUI agents still lag behind closed-source systems on long-horizon navigation tasks. This gap stems from two limitations: a
arXiv:2605.26835v1 Announce Type: new Abstract: LLM-based multi-agent systems have been widely adopted for knowledge retrieval and report generation, synthesizing known information through web search
Conductor migrated its parallel coding agents infrastructure from local laptops to the cloud using Vercel Sandbox, enabling improved scalability and distributed execution of AI-driven code generation
Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staf
arXiv:2602.00959v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far t
arXiv:2605.26177v1 Announce Type: cross Abstract: Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on e
arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin
Agentic artificial intelligence security startup 7AI Inc. today announced the launch of PLAID ELITE, a fully managed AI-native security operations service. The new service combines autonomous investig
arXiv:2605.25440v1 Announce Type: cross Abstract: Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, asse
In this post, we demonstrate the capabilities of AgentWatch through practical implementation. You will see how the solution performs infrastructure checks every 15 minutes, summarizing CloudWatch metr
Agents are turning up the heat for enterprises, turning air-cooled AI infrastructure into a boardroom priority. While GPUs have largely dominated the AI conversation, CPUs are increasingly coming into
arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc
In this post you'll learn how to build a multi-agent campaign review system that demonstrates parallel reasoning, context persistence, and traceable execution paths using an integrated architecture th
arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint
arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its exte
arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With
arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuari
arXiv:2602.02474v2 Announce Type: replace-cross Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixe
arXiv:2605.25030v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (L
arXiv:2605.25958v1 Announce Type: new Abstract: This paper introduces PolyGnosis 2.0, a pioneering multi-agent architecture designed to extract predictive intelligence by synthesizing Polymarket anoma
Nous Research has added support for Qwen 3.7 Max, a large language model, within their Hermes Agent framework. This integration enables users to leverage Qwen 3.7 Max's capabilities when building or d
arXiv:2605.23949v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into interactive agents, understanding their behavioral alignment within human social dynamics becomes essentia
arXiv:2605.23950v1 Announce Type: new Abstract: This position paper argues that, for long-horizon tasks evaluated across models with comparable frontier capability, the agent execution harness, namely
Amazon Bedrock AgentCore payments is now available in preview, it provides instant payments to paid external services with no manual billing setup per provider, stablecoin support for cost-effective m
arXiv:2605.24138v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to software engineering (SE), yet their potential for autonomous, role-oriented collaboration re
arXiv:2605.22896v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representa
Simon Willison discusses whether coding agents should automatically generate tests alongside the code they produce, addressing a key quality assurance consideration in AI-assisted development. This li
arXiv:2605.22883v1 Announce Type: new Abstract: Current AI energy benchmarks measure consumption at the granularity of a single model invocation or training run. For classical single-turn workloads th
arXiv:2602.11243v2 Announce Type: replace-cross Abstract: Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augm
arXiv:2604.05157v2 Announce Type: replace Abstract: Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evalu
arXiv:2605.22825v1 Announce Type: cross Abstract: Key Value Indicators (KVIs) provide a decision oriented view of a service by summarizing how operational performance translates into stakeholder value
arXiv:2602.13473v2 Announce Type: replace Abstract: Although foundation models have demonstrated remarkable success in general domains, the application of these models to electroencephalography (EEG)
arXiv:2605.22855v1 Announce Type: cross Abstract: Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision makin
arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and
arXiv:2602.11210v4 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a key paradigm for training software engineering (SWE) agents, but existing pipelines typically rely on
As enterprises move from AI experimentation to full-scale production, the absence of a unified enterprise AI operating system is emerging as the single most consequential bottleneck in realizing retur
arXiv:2605.20210v1 Announce Type: cross Abstract: Agentic AI systems - systems that can pursue goals through multi-step planning and tool-mediated action with limited direct supervision - are moving f
arXiv:2410.04753v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been used to generate formal proofs of mathematical theorems in proofs assistants such as Lean. However, we
OpenAI has been recognized as a Leader in Gartner's 2026 evaluation of enterprise coding agents, highlighting its competitive position in AI-powered code generation and software development tools. Thi