Learning Agent Routing From Early Experience
arXiv:2605.07180v1 Announce Type: new Abstract: LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capa
Knowledge catalogue
arXiv:2605.07180v1 Announce Type: new Abstract: LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capa
arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H
arXiv:2605.07038v1 Announce Type: new Abstract: Risk-aware navigation should be selective: a policy should expose evasive degrees of freedom only when the local scene admits a lower-risk feasible mane
arXiv:2605.07640v1 Announce Type: cross Abstract: Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general lan
arXiv:2508.16571v4 Announce Type: replace Abstract: In this paper, we describe and benchmark a competitor-discovery component used within an agentic AI system for fast drug asset due diligence. A comp
Local open-weight AI on a laptop has been improving more than twice as fast as Moore's Law! Between May 2024 and May 2026, the most expensive MacBook Pro you could buy stayed at 128 GB of unified memo
arXiv:2605.07342v1 Announce Type: cross Abstract: Compile-pass rate is the dominant evaluation signal for LLM code generation, yet for multi-component domain-specific artifacts it can be actively misl
arXiv:2605.05949v2 Announce Type: replace Abstract: Algorithmic problem solving serves as a rigorous testbed for evaluating structured reasoning in AI coding systems, as it directly reflects a model's
arXiv:2605.07280v1 Announce Type: cross Abstract: Leveraging deep learning for causal discovery in time series remains challenging because existing neural methods predominantly rely on component-wise
arXiv:2605.07600v1 Announce Type: cross Abstract: Recent methods for improving LLM mathematical reasoning, whether through MCTS-based test-time search or causal graph-guided knowledge injection, canno
arXiv:2605.07147v1 Announce Type: cross Abstract: The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes
arXiv:2605.07850v1 Announce Type: cross Abstract: With the rise in scale for deep learning models to billions of parameters, the computational cost of fine-tuning remains a significant barrier to depl
arXiv:2605.07646v1 Announce Type: cross Abstract: While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verifi
arXiv:2605.06894v1 Announce Type: cross Abstract: Machine learning (ML) in real-world systems must contend with concept drift, adversarial actors, and a spectrum of potential features with varying cos
arXiv:2605.07345v1 Announce Type: new Abstract: Mean-pooled cosine similarity is the default metric for comparing neural representations across languages, modalities, and tasks. We establish that this
arXiv:2605.07305v1 Announce Type: cross Abstract: Most existing LLM diagnoses are evaluated on static, single-turn settings where complete patient information is provided upfront, an oversimplificatio
arXiv:2605.07919v1 Announce Type: new Abstract: Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property:
Managing a modern database fleet is both a scale and cognitive problem. As database estates grow, the effort required to monitor, troubleshoot, and optimize them often outpaces teams’ capacity, who fi
arXiv:2605.06903v1 Announce Type: cross Abstract: Large language models are now embedded in everyday writing workflows, making reliable AI-generated text detection important for academic integrity, co
arXiv:2602.06523v2 Announce Type: replace Abstract: Human Activity Recognition (HAR) on resource constrained wearables requires models that balance accuracy against strict memory and computational bud
arXiv:2605.06797v1 Announce Type: new Abstract: We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Frechet I
arXiv:2603.09652v3 Announce Type: replace Abstract: With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynami
arXiv:2605.07269v1 Announce Type: new Abstract: Indirect prompt injection remains a persistent weakness in retrieval-augmented and tool-using LLM systems, and the problem becomes harder to characteris
arXiv:2605.07363v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) sets the state of the art for fine-grained inference-time sparse attention by introducing a learned token-wise indexer
arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu
arXiv:2603.24946v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance on automated software engineering tasks, yet existing benchmarks focus primarily on
arXiv:2605.07520v1 Announce Type: new Abstract: Differentiable planning enables gradient-based optimization of decision-making problems by leveraging differentiable models of system dynamics. However,
arXiv:2605.07075v1 Announce Type: new Abstract: The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly in
arXiv:2605.06709v1 Announce Type: new Abstract: This paper addresses PDE-based control for flexible multibody robotic systems, presenting a subsystem-based framework for serial manipulators with arbit
arXiv:2605.06672v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning and reasoning-tuned models such as DeepSeek-R1 are commonly assumed to reduce shallow heuristic biases by thinking care
arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin
arXiv:2605.06940v1 Announce Type: new Abstract: Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set i
arXiv:2604.04891v2 Announce Type: replace-cross Abstract: Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter b
My Mac had less available memory than I expected, turned out the 'claude' Claude Code processes on this machine (running in various terminal windows) were consuming ~30GB on their own! The largest one
arXiv:2605.06846v1 Announce Type: cross Abstract: Recent work identifies secret loyalties as a distinct threat from standard backdoors. A secret loyalty causes a model to covertly advance the interest
arXiv:2603.08256v2 Announce Type: replace Abstract: Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short nar
arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo
arXiv:2605.07792v1 Announce Type: cross Abstract: Neural operators (NOs) are designed to learn maps between infinite-dimensional function spaces. We propose a novel reframing of their use. By introduc
Simon Willison discovered how to use an LLM command-line interface tool in Unix shebang lines, enabling the creation of executable scripts written in English or natural language. This technique allows
arXiv:2605.07476v1 Announce Type: new Abstract: Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variab
arXiv:2508.01248v4 Announce Type: replace Abstract: The rapid progress of generative models, such as GANs and diffusion models, has facilitated the creation of highly realistic images, raising growing
arXiv:2605.07051v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown good performance on various science educational benchmarks, demonstrating their potential for use in science and
arXiv:2605.06728v1 Announce Type: cross Abstract: Interpreting transcriptomic data is one of the most common analytical tasks in modern biology. Yet most current models either consume expression profi
arXiv:2605.07546v1 Announce Type: new Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocatio
One of the most important properties of LLMs that we take for granted is that newer, bigger models are just better at everything. The AI Labs are pouring effort into economically valuable fields like
The OpenAI Campus Network is a program that facilitates student engagement with OpenAI's technology and research on college campuses. This interest form allows students to express interest in starting
OpenAI is launching Daybreak, an AI initiative focused on detecting and patching vulnerabilities before attackers find them. Daybreak uses the Codex Security AI agent that launched in March to create
Alexey Shabanov / TestingCatalog AI News: OpenAI launches Daybreak, a cybersecurity initiative integrating AI models and Codex Security to help organizations patch vulnerabilities — OpenAI launches Da
OpenAI launched DeployCo, a new service designed to assist businesses in building and deploying applications leveraging OpenAI's AI models and intelligence capabilities. The offering appears to focus
OpenAI Group PBC today unveiled a new business unit, The OpenAI Deployment Company, that will help companies adopt its artificial intelligence models. The subsidiary is launching with 4 billion in fun
Reuters: OpenAI launches the OpenAI Deployment Company with a 4B+ investment to help organizations build and deploy AI systems, and acquires AI consulting firm Tomoro — OpenAI said on Monday it is set
arXiv:2605.06993v1 Announce Type: new Abstract: Causal queries are often only partially identifiable from observational data, and experiments that could tighten the resulting bounds are typically cost
arXiv:2603.04678v2 Announce Type: replace-cross Abstract: Large language models are known to often exhibit inconsistent knowledge. This is particularly problematic in multilingual scenarios, where mod
arXiv:2605.07815v1 Announce Type: cross Abstract: Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a glo
arXiv:2511.22316v2 Announce Type: replace Abstract: Large Language Models (LLMs) quantization facilitates deploying LLMs in resource-limited settings, but existing methods that combine incompatible gr
arXiv:2602.00465v3 Announce Type: replace-cross Abstract: Functional miRNA--mRNA targeting is a large-bag prediction problem where each transcript yields a heavy-tailed pool of candidate target sites
arXiv:2605.07496v1 Announce Type: new Abstract: Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided b
arXiv:2605.07267v1 Announce Type: new Abstract: Personalized healthcare decisions require reasoning about how physiological and behavioral variables influence an individual patient over time. Existing
arXiv:2512.14018v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable progress in automatic code generation, yet their ability to produce high-performance cod
arXiv:2605.07687v1 Announce Type: new Abstract: Physics-based digital twins aim to predict the dynamics of real-world objects under interaction, enabling real-to-sim-to-real applications in robotics.