AI Research Agents Narrow Scientific Exploration
arXiv:2605.27905v1 Announce Type: new Abstract: AI research agents can now generate research ideas, design experiments, run code, and draft papers, raising the possibility of large-scale AI-assisted s
Knowledge catalogue
arXiv:2605.27905v1 Announce Type: new Abstract: AI research agents can now generate research ideas, design experiments, run code, and draft papers, raising the possibility of large-scale AI-assisted s
arXiv:2605.28359v1 Announce Type: new Abstract: Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a his
arXiv:2605.27489v1 Announce Type: cross Abstract: Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves interpretabilit
Hermes Agent v0.15.0, released by Nous Research as 'The Velocity Release,' represents an update to their Hermes Agent framework that likely focuses on performance improvements and speed optimizations.
arXiv:2605.27999v1 Announce Type: cross Abstract: We address the problem of learning to assign prediction tasks to one agent from a set of available human or AI agents. In particular, we focus on the
arXiv:2605.27825v1 Announce Type: cross Abstract: Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure p
Nous Research has announced support for Opus 4.8 in the Hermes Agent framework. This update enables the Hermes Agent to utilize Anthropic's Opus 4.8 model, expanding the available model options for us
arXiv:2605.28424v1 Announce Type: new Abstract: Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks. Agent sk
arXiv:2605.27850v1 Announce Type: new Abstract: Effective multi-agent systems cannot be designed by selecting prompts or communication graphs in isolation. Agent behavior depends on the information an
arXiv:2502.14321v3 Announce Type: replace-cross Abstract: Large language model-based multi-agent systems have recently gained significant attention due to their potential for complex, collaborative, a
cognition is now the largest independent agent lab in the world. take the 200% utilization that everyone is hitting from this chart and run out the sales growth from this, i encourage you to go thru t
Fleet agents now come with a computer! They can write and execute code, which is helpful for general purpose tasks beyond coding Fleet agents can now securely write and run code. With computer use in
Great to see a GBrain x ActiveGraph crossover here Great retrieval pairs well with almost anything you want to do with agents Gbrain is what an agent *knows* — a durable markdown/git knowledge substra
arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents
arXiv:2605.26256v1 Announce Type: new Abstract: Multimodal large language model (MLLM)-based embodied agents have shown strong potential for solving complex tasks in physical environments. However, pe
As agent adoption scaled, we saw a common pattern emerge across enterprises, including our own sales organization: specialized agents deliver value, but without orchestration, users carry the cognitiv
Robinhood is opening its trading platform to AI agents. In an announcement on Wednesday, Robinhood says traders can now create a separate account for an AI agent and add a specific amount of money, al
Yohei Nakajima discusses the emergence of self-improving AI agents, likely exploring how autonomous agents can iteratively enhance their capabilities and performance through feedback loops and learnin
“agent debt” is a new term, but it was inevitable, an instantiation of the AI-driven technical debt I keep warning about. what happens when you build it fast, but you don’t really know how to fix it.
arXiv:2605.24110v1 Announce Type: new Abstract: Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assess
arXiv:2605.25641v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems in complex B2B (business-to-business) settings may often receive free-form response feedback. Rathe
LangChain Academy Course: LangSmith Fleet Essentials Learn how to build your own agents with LangSmith Fleet. Anyone can now build, use, and manage an agent fleet for complex daily tasks, without writ
arXiv:2605.24468v1 Announce Type: new Abstract: Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and
New paper from Microsoft on Self-Evolving Agent Skills New research from Microsoft Research I see a lot of AI engineers handwriting agent skill docs and hope they generalize. Probably not optimal. Thi
nice write up from the HuggingFace folks aggregating works on defining agents, harnesses, environments, RL, etc. The more we can roughly have a shared vocabulary the better…I still find it confusing (
Latent Space reports on a rebranding or reorganization where Model Labs have been renamed or converted into Agent Labs, reflecting a shift in focus toward AI agent development and capabilities. This c
arXiv:2605.22794v1 Announce Type: cross Abstract: Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the nex
arXiv:2605.21850v1 Announce Type: new Abstract: Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly lo
arXiv:2605.20425v1 Announce Type: new Abstract: Designing multi-agent workflows is especially difficult in open-ended scientific settings where tasks lack curated training sets, reliable scalar evalua
arXiv:2605.20704v1 Announce Type: cross Abstract: Autonomous AI agents that spawn sub-agent swarms create a safety gap: existing credential revocation mechanisms, OAuth~2.0 introspection, OCSP, and W3
Hermes Agent has been updated to support integration with Bitwarden Secrets Manager, enabling secure credential and secret management capabilities within the Hermes agent framework. This integration a
Kafka, but for agents! i'm excited to open source Active Graph: an event-sourced reactive graph runtime for long-running, agents 🔄🧠 events/logs projects a graph. reactive behaviors react and affect th
MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to support efficient agent
arXiv:2603.06007v2 Announce Type: replace Abstract: Large language model-based (LLM-based) multi-agent systems (MAS) are increasingly used to extend agentic problem solving via role specialization and
arXiv:2605.20173v1 Announce Type: new Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a firs
The innovation race has entered a new phase with enterprise agentic AI spreading across every part of the business. The potential of agents — once limited to technical teams — is quickly becoming acce
Bot and agent trust management company DataDome SAS today launched Priority Protect, a virtual waiting room product designed to sort human shoppers, authorized artificial intelligence agents and malic
arXiv:2605.20049v1 Announce Type: cross Abstract: As autonomous coding agents see rapid adoption, their evaluation has primarily focused on task completion rates holding the target codebase fixed. Thi
arXiv:2603.03140v3 Announce Type: replace-cross Abstract: AI agents are increasingly active on social media platforms, generating content and interacting with one another at scale. Yet the behavioral
arXiv:2605.19330v1 Announce Type: new Abstract: LLM agents organize behavior through skills - structured natural-language specifications governing how an agent reasons, retrieves, and responds. Unlike
arXiv:2605.05974v2 Announce Type: replace-cross Abstract: LLM agents rely on prompts to implement task-specific capabilities based on foundation LLMs, making agent prompts valuable intellectual proper
you can condense long horizon evals with agents into smaller subsets that still let you test intended behavior. i'm currently evaluating an agent that runs for 30+ minutes, and analyzes thousands of t
arXiv:2605.16819v1 Announce Type: cross Abstract: GPU kernel optimization is increasingly critical for efficient deep learning systems, but writing high-performance kernels still requires substantial
arXiv:2605.18535v1 Announce Type: new Abstract: The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This p
Engine is a sophisticated production agent system developed by a team that Harrison Chase recommends for understanding advanced agent architecture and implementation. The post suggests it represents n
arXiv:2605.18703v1 Announce Type: new Abstract: Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robus
arXiv:2605.17734v1 Announce Type: new Abstract: Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizo
arXiv:2605.18024v1 Announce Type: cross Abstract: Cooperation is central to multi-agent reinforcement learning (MARL), yet learned coordination can be fragile when external perturbations disrupt inter
arXiv:2605.18597v1 Announce Type: new Abstract: Large language model (LLM) agents often rely on long sequences of low-level textual actions, resulting in large effective decision horizons and high inf
arXiv:2605.18077v1 Announce Type: new Abstract: Communication is a key component in multi-agent reinforcement learning (MARL) for mitigating partial observability, yet prior approaches often rely on i
arXiv:2601.00360v3 Announce Type: replace-cross Abstract: As multi-agent AI systems become increasingly autonomous, evidence shows they can develop collusive strategies similar to those long observed
arXiv:2605.16821v1 Announce Type: new Abstract: The rapid evolution of Large Language Model (LLM) agents has produced diverse interaction paradigms, yet few production systems integrate multiple parad
arXiv:2602.02262v3 Announce Type: replace-cross Abstract: LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require
arXiv:2601.01155v3 Announce Type: replace Abstract: Existing methods for multi-agent navigation typically assume fully known environments, offering limited support for partially known scenarios with o
arXiv:2605.17830v1 Announce Type: new Abstract: Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under
arXiv:2605.18332v1 Announce Type: cross Abstract: Behavioral studies of LLM-based software engineering agents extract operational rules about which trajectory shapes correlate with higher resolution r
arXiv:2605.16872v1 Announce Type: cross Abstract: AI agents increasingly act consequentially in the real world. This creates a problem we call consequence reception: harm occurs, the producing system
two years ago we started building agents to automate work. turns out these are really useful, so there’s a LOT of runs and long traces that are hard to reason about now, use an ambient agent (engine)
arXiv:2605.15815v1 Announce Type: cross Abstract: Code agents increasingly help developers work with unfamiliar repositories, but every such task depends on a costly prerequisite: bootstrapping the re
arXiv:2605.16205v1 Announce Type: new Abstract: Deploying compound LLM agents in adversarial, partially observable sequential environments requires navigating several design dimensions: (1) what the a