AI Agents and Hard Choices
arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t
Knowledge catalogue
arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t
arXiv:2510.07043v2 Announce Type: replace Abstract: Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shop
arXiv:2504.13541v5 Announce Type: replace-cross Abstract: Training resource-constrained autonomous agents on multiple tasks simultaneously is crucial for adapting to diverse real-world environments. R
arXiv:2604.15075v1 Announce Type: cross Abstract: Open-weight Small Language Models(SLMs) can provide faster local inference at lower financial cost, but may not achieve the same performance level as
arXiv:2604.13151v1 Announce Type: new Abstract: Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these
Simon Willison discusses the perception that LLMs and coding agents are primarily useful for greenfield development (starting new projects from scratch) rather than for maintaining and modifying large
One portal, unlimited possibilities. You can now access Modal via Tool Gateway by @NousResearch, makers of Hermes Agent. Check it out 👇 Tool Gateway is now live in Nous Portal. No separate accounts, n
arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg
arXiv:2510.27420v3 Announce Type: replace Abstract: Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often l
AI Search is the search primitive for your agents. Create instances dynamically, upload files, and search across instances with hybrid retrieval and relevance boosting. Just create a search instance,
arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.
LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core is extremely fast text and accurate parsing from any document
arXiv:2512.20798v4 Announce Type: replace Abstract: As autonomous AI agents are deployed in high-stakes environments, ensuring their safety has become a paramount concern. Existing safety benchmarks p
arXiv:2604.12762v1 Announce Type: cross Abstract: We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an ag
arXiv:2604.12147v1 Announce Type: cross Abstract: Agents aspire to eliminate the need for task-specific prompt crafting through autonomous reason-act-observe loops. Still, they are commonly instructed
I'm going all in on Hermes (@NousResearch, @Teknium1) as my entire agent and coding stack. Six profiles. One shared self-hosted memory store. Zero hosted-coder dependencies. The fleet: - pmax-mousa —
An open-source agent self-reflection harness built on top of Ollama, shared in the r/ollama community, that enables locally run LLMs to evaluate and iteratively refine their own outputs. The project p
arXiv:2604.12282v1 Announce Type: new Abstract: Spreadsheets are central to real-world applications such as enterprise reporting, auditing, and scientific data management. Despite their ubiquity, exis
arXiv:2604.09621v1 Announce Type: new Abstract: We present an agent-driven approach to the construction of parameter inference pipelines for scientific data analysis. Our method leverages a multi-agen
arXiv:2604.09813v1 Announce Type: new Abstract: Existing synthetic tool-use corpora are primarily designed for offline supervised fine-tuning, yet reinforcement learning (RL) requires executable envir
arXiv:2604.01236v2 Announce Type: replace-cross Abstract: Traditional network architectures suffer from severe protocol ossification and structural fragility due to their reliance on static, human-def
arXiv:2604.09666v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) and its graph-based extensions (GraphRAG) are effective paradigms for improving large language model (LLM) reason
arXiv:2604.09568v1 Announce Type: cross Abstract: High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant chall
arXiv:2604.10989v1 Announce Type: new Abstract: Emergency situations in scheduling systems often trigger local functional failures that undermine system stability and even cause system collapse. Exist
Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real t
arXiv:2604.10261v1 Announce Type: new Abstract: Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of
arXiv:2604.11465v1 Announce Type: new Abstract: Large language model (LLM) agents show promise on realistic tool-use tasks, but deploying capable agents on modest hardware remains challenging. We stud
arXiv:2506.02387v3 Announce Type: replace Abstract: Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain lim
We've been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to apply it to optimizing CUDA kernels. In 3 weeks, it delivered
arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com
a big part of agent harnesses is how they interact with context memory is just context its therefor impossible to separate harness from memory - as @sarahwooders says, 'memory isn't a plugin (it's a h
LiteParse is the best document parsing library for coding agents. It's free, fast, integrates natively with the LLM's native visual understanding capabilities, and comes with support for 50+ formats a
Databricks Research introduced **MemAlign**, a memory framework for AI agents that stores past interactions as episodic memories and uses an LLM to distill them into generalized semantic rules, whi...
arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs
arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra
GLM-5.1 is Z.ai's post-training upgrade to GLM-5, now available on Together AI, delivering a 28% coding performance improvement through a refined reinforcement learning pipeline while retaining the...
How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: - Parse files and take page-level screenshots with LiteParse,
arXiv:2608.12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per t
arXiv:2608.13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We mode
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting
arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social med
arXiv:2608.13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, viola
arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is struc
arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and
our outbound harness is focused on delivering the best economics we can, we're continuing to push on how our agents work to get another 90%+ cost savings for customers great talk by @hwchase17 and @He
arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills,
arXiv:2608.11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequent
arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical im
arXiv:2608.11879v1 Announce Type: new Abstract: Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to
arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework
arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti
arXiv:2608.10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. Howeve
arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essenti
arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh
arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl
arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven
arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long
arXiv:2608.08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible clai
arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length i
arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly availab