Workspace Optimization: How to Train Your Agent
arXiv:2605.09650v1 Announce Type: new Abstract: Modern agents built on frontier language models often cannot adapt their weights. What, then, remains trainable? We argue it is the agent's workspace, t
Knowledge catalogue
arXiv:2605.09650v1 Announce Type: new Abstract: Modern agents built on frontier language models often cannot adapt their weights. What, then, remains trainable? We argue it is the agent's workspace, t
arXiv:2605.09360v1 Announce Type: cross Abstract: Execution-based evaluation of LLM-generated code implicitly treats successful execution as a proxy for correctness. In scientific simulation, this pro
arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne
arXiv:2605.08814v1 Announce Type: new Abstract: Chinese character categories are extremely large, and unseen characters frequently arise in open-world scenarios, making zero-shot Chinese character rec
arXiv:2605.07892v1 Announce Type: new Abstract: Sparse training reduces the memory and computational costs of deep neural networks. However, sparse optimization methods, e.g., those adding an ell_1 pe
arXiv:2605.06607v2 Announce Type: replace-cross Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemi
Nous Research announced free access to Qwen 3.6 Plus and additional features available through their subscription portal at portal.nousresearch.com. This promotion likely provides users with complimen
arXiv:2602.12162v3 Announce Type: replace Abstract: In structurally constrained molecular optimization, state-of-the-art methods restart an expensive oracle-driven search from scratch for every new in
arXiv:2605.07191v1 Announce Type: new Abstract: A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a random
arXiv:2605.06834v1 Announce Type: new Abstract: Continual learning research attempts to conserve two fundamental capabilities: new knowledge acquisition and the preservation of previously acquired kno
Bare Metal: Z-Image Turbo - Flux.2 Klein 9b - Wan 2.2 discusses comparisons between these lightweight AI image generation models, with Flux.2 Klein 9B showing good prompt adherence and natural composi
arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l
arXiv:2605.07306v1 Announce Type: cross Abstract: Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab environment
arXiv:2605.08057v1 Announce Type: cross Abstract: While recent advancements in inference-time learning have improved LLM reasoning on Text-to-SQL tasks, current solutions still struggle to perform wel
arXiv:2605.07230v1 Announce Type: cross Abstract: Autoregressive generation is a powerful approach for high-fidelity image synthesis, but it remains computationally demanding and slow even on the most
arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab
Congrats to @thinkymachines on the release of TML-Interaction-Small and tying for the top spot on our Audio MC S2S leaderboard! 🥇 Their interaction model scores a 43.4% APR, demonstrating an impressiv
arXiv:2605.07959v1 Announce Type: new Abstract: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications.
arXiv:2605.07724v1 Announce Type: cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal
DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is the difference between: GPT 5.4 Mini: ~394/mo Gemini 3.1 Flash
arXiv:2605.07503v1 Announce Type: new Abstract: Efficiently aligning large-scale video diffusion models with human intent requires a scalable and trajectory-aware pathway that bridges the inherent dis
arXiv:2605.07551v1 Announce Type: new Abstract: Standard Importance Sampling (IS) collapses under label corruption because high-norm examples, prioritized for variance reduction, are often adversarial
arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres
arXiv:2605.07699v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently a
arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as
arXiv:2605.07888v1 Announce Type: cross Abstract: Federated Learning (FL) enables decentralised model training across distributed clients without requiring data centralisation. However, the generalisa
arXiv:2605.07247v1 Announce Type: new Abstract: Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments
arXiv:2512.12602v4 Announce Type: replace Abstract: In this paper, we introduce Exact Flow Linear Attention~(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rul
arXiv:2603.06859v2 Announce Type: replace-cross Abstract: Removing an agent from a cooperative team to measure its contribution seems natural, yet in multi-agent LLM systems this evaluation distorts t
arXiv:2605.07501v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve strong performance via extended chain-of-thought (CoT) reasoning, yet suffer from excessive token consumption an
arXiv:2605.07273v1 Announce Type: cross Abstract: Multimodal RAG systems increasingly rely on vision-language retrievers to ground visual queries in external textual evidence. Existing adversarial stu
arXiv:2605.08022v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) have been proposed as biologically plausible and energy-efficient alternatives to conventional Artificial Neural Networ
arXiv:2410.06347v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) in robotics faces significant hurdles regarding sample efficiency and generalization across varying goals. While O
arXiv:2605.07999v1 Announce Type: cross Abstract: Multiphoton photoreduction enables high-fidelity fabrication of complex 3D microstructures, yet reliable process-structure-property (PSP) prediction r
arXiv:2605.07053v1 Announce Type: cross Abstract: Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed
arXiv:2605.07707v1 Announce Type: new Abstract: HTN planning is a variation of classical planning where, instead of searching for a linear sequence of actions, an algorithm decomposes higher-level tas
arXiv:2605.07545v1 Announce Type: cross Abstract: Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their hi
arXiv:2605.07456v1 Announce Type: new Abstract: Inference-time controllable generation is essential for real-world applications of unconditional diffusion models. However, most existing techniques foc
arXiv:2605.07284v1 Announce Type: new Abstract: Recent interpretability work has identified model-internal handles on post-trained behavior, including refusal directions, assistant/persona axes, and s
arXiv:2605.07613v1 Announce Type: new Abstract: Conversational news recommendation requires grounding each suggestion in a rapidly evolving article corpus while addressing implicit user intents that l
arXiv:2605.07510v1 Announce Type: cross Abstract: Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input
arXiv:2605.06859v1 Announce Type: cross Abstract: Vision foundation models are increasingly moving beyond 2D to volumetric domains such as 3D medical imaging, where unified pretraining across differen
arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H
arXiv:2605.07460v1 Announce Type: new Abstract: Accurate Monte Carlo (MC) modelling in high-energy physics is challenging, particularly in complex scenarios where simulations fail to reproduce observe
arXiv:2511.09907v5 Announce Type: replace Abstract: Data synthesis for training large reasoning models offers a scalable alternative to limited, human-curated datasets, enabling the creation of high-q
arXiv:2605.07505v1 Announce Type: new Abstract: Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device
Local open-weight AI on a laptop has been improving more than twice as fast as Moore's Law! Between May 2024 and May 2026, the most expensive MacBook Pro you could buy stayed at 128 GB of unified memo
arXiv:2507.21183v5 Announce Type: replace-cross Abstract: As the era of large language models (LLMs) unfolds, Preference Optimization (PO) methods have become a central approach to aligning LLMs with
arXiv:2605.07280v1 Announce Type: cross Abstract: Leveraging deep learning for causal discovery in time series remains challenging because existing neural methods predominantly rely on component-wise
arXiv:2605.07600v1 Announce Type: cross Abstract: Recent methods for improving LLM mathematical reasoning, whether through MCTS-based test-time search or causal graph-guided knowledge injection, canno
arXiv:2605.06894v1 Announce Type: cross Abstract: Machine learning (ML) in real-world systems must contend with concept drift, adversarial actors, and a spectrum of potential features with varying cos
Managing a modern database fleet is both a scale and cognitive problem. As database estates grow, the effort required to monitor, troubleshoot, and optimize them often outpaces teams’ capacity, who fi
arXiv:2605.07269v1 Announce Type: new Abstract: Indirect prompt injection remains a persistent weakness in retrieval-augmented and tool-using LLM systems, and the problem becomes harder to characteris
arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin
arXiv:2605.07792v1 Announce Type: cross Abstract: Neural operators (NOs) are designed to learn maps between infinite-dimensional function spaces. We propose a novel reframing of their use. By introduc
Simon Willison discovered how to use an LLM command-line interface tool in Unix shebang lines, enabling the creation of executable scripts written in English or natural language. This technique allows
arXiv:2605.07476v1 Announce Type: new Abstract: Multivariate time series forecasting remains a challenge due to the complexity of local temporal dynamics and global dependencies across multiple variab
arXiv:2605.07220v1 Announce Type: new Abstract: Diffusion guidance is a powerful technique that enables controllable and high-fidelity sample generation with diffusion models. At a high level, it modi
arXiv:2605.06110v2 Announce Type: replace Abstract: Agentic systems increasingly solve complex user requests by executing orchestrated workflows, where subtasks are assigned to specialized models or t
The OpenAI Campus Network is a program that facilitates student engagement with OpenAI's technology and research on college campuses. This interest form allows students to express interest in starting