Golden Handcuffs make safer AI agents
arXiv:2604.13609v1 Announce Type: new Abstract: Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand th
Knowledge catalogue
arXiv:2604.13609v1 Announce Type: new Abstract: Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand th
arXiv:2604.13870v1 Announce Type: cross Abstract: We consider the well-studied setting of minimizing a convex Lipschitz function using either gradient descent (GD) or its stochastic variant (SGD), and
arXiv:2604.13722v1 Announce Type: new Abstract: We address the challenge of synthetic-to-real transfer in forestry perception where real data have only coarse Tree labels while synthetic data provide
arXiv:2603.12725v3 Announce Type: replace Abstract: In-context operator learning enables neural networks to infer solution operators from contextual examples without weight updates. While prior work h
arXiv:2604.13127v1 Announce Type: new Abstract: The need to selectively and efficiently erase learned information from deep neural networks is becoming increasingly important for privacy, regulatory c
arXiv:2510.00573v2 Announce Type: replace Abstract: Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especia
arXiv:2512.10877v4 Announce Type: replace Abstract: Discrete diffusion models (DMs) have achieved strong performance in language and other discrete domains, offering a compelling alternative to autore
arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu
arXiv:2604.13871v1 Announce Type: new Abstract: Deep neural networks (DNNs) deliver state-of-the-art accuracy on regression and classification tasks, yet two structural deficits persistently obstruct
arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of
arXiv:2604.13258v1 Announce Type: new Abstract: Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most exist
arXiv:2604.13947v1 Announce Type: new Abstract: We present lightweight and efficient architectures to detect weather conditions from RGB images, predicting the weather type (sunny, rain, snow, fog) an
arXiv:2510.19268v2 Announce Type: replace-cross Abstract: Long-horizon routing tasks of deformable linear objects (DLOs), such as cables and ropes, are common in industrial assembly lines and everyday
arXiv:2604.05808v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents have demonstrated strong capabilities in complex interactive decision-making tasks. However, existing LLM ag
arXiv:2604.14032v1 Announce Type: cross Abstract: Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its de
arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.
arXiv:2604.13981v1 Announce Type: new Abstract: Interpretability is essential for deploying object detection systems in critical applications, especially under low-quality imaging conditions that degr
arXiv:2604.14125v1 Announce Type: new Abstract: While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often
arXiv:2604.13977v1 Announce Type: new Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing stra
arXiv:2604.13627v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to follow instructions and shaping its behavior as a hel
arXiv:2604.13179v1 Announce Type: cross Abstract: This paper presents HUANet, a constrained deep neural network architecture that unrolls the iterations of the Alternating Direction Method of Multipli
arXiv:2509.25549v2 Announce Type: replace Abstract: Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to impro
arXiv:2411.10703v3 Announce Type: replace Abstract: The availability of continuous glucose monitors as over-the-counter commodities have created a unique opportunity to monitor a person's blood glucos
arXiv:2604.13728v1 Announce Type: cross Abstract: We present a hybrid retrieval system for COVID-19 scientific literature, evaluated on the TREC-COVID benchmark (171,332 papers, 50 expert queries). Th
arXiv:2603.28554v2 Announce Type: replace Abstract: Visual document understanding typically requires separate retrieval and generation models, doubling memory and system complexity. We present Hydra,
arXiv:2604.14114v1 Announce Type: cross Abstract: Sequential recommendation has become increasingly prominent in both academia and industry, particularly in e-commerce. The primary goal is to extract
arXiv:2604.13218v1 Announce Type: cross Abstract: Causal representation learning (CRL) aims to identify the underlying latent variables from high-dimensional observations, even when variables are depe
arXiv:2512.01773v2 Announce Type: replace Abstract: The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-
arXiv:2604.13268v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remain
arXiv:2604.13686v1 Announce Type: new Abstract: While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and
arXiv:2604.13201v1 Announce Type: new Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks
arXiv:2604.14111v1 Announce Type: new Abstract: Large Language Models (LLMs) are now capable of generating highly fluent, human-like text. They enable many applications, but also raise concerns such a
arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam
arXiv:2604.13078v1 Announce Type: new Abstract: The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts
arXiv:2604.13484v1 Announce Type: cross Abstract: Clustering and dimensionality reduction have been crucial topics in machine learning and computer vision. Clustering high-dimensional data has been ch
arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper
arXiv:2603.12021v2 Announce Type: replace Abstract: Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from a high-resource language to low-resour
arXiv:2604.13058v1 Announce Type: new Abstract: We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,46
arXiv:2604.13226v1 Announce Type: new Abstract: Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: re
arXiv:2603.29159v2 Announce Type: replace Abstract: Providing timely and accurate learning support in large-scale online coding courses is challenging, particularly in resource-constrained contexts. W
arXiv:2604.13285v1 Announce Type: new Abstract: Clinical text classification requires choosing between specialized fine-tuned models (BERT variants) and general-purpose large language models (LLMs), y
arXiv:2510.13849v3 Announce Type: replace Abstract: Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks.
arXiv:2511.11334v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian
arXiv:2512.17654v3 Announce Type: replace Abstract: We present three variants of a lightweight, fully connected artificial neural network, suited for interactive estimation of three-dimensional, spati
arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly
arXiv:2511.05330v2 Announce Type: replace Abstract: Embedding non-restrictive prior knowledge, such as energy conservation laws, into learning methods is a key motive to construct physically consisten
arXiv:2604.13462v1 Announce Type: cross Abstract: Effective IT change management is important for businesses that depend on software and services, particularly in highly regulated sectors such as fina
arXiv:2604.13546v1 Announce Type: new Abstract: Conventional neural networks strictly separate learning and inference because if parameters are updated during inference, outputs become unstable and ev
arXiv:2604.13128v1 Announce Type: cross Abstract: Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Underst
arXiv:2601.17740v2 Announce Type: replace Abstract: Sewing patterns define the structural foundation of garments and are essential for applications such as fashion design, fabrication, and physical si
arXiv:2604.13713v1 Announce Type: new Abstract: Metaphor detection models achieve strong benchmark performance, yet it remains unclear whether this reflects transferable generalization or lexical memo
arXiv:2604.13520v1 Announce Type: new Abstract: Metal-organic frameworks (MOFs) are highly promising for carbon capture, yet navigating their vast design space remains challenging. Recent deep generat
arXiv:2512.10605v2 Announce Type: replace Abstract: We propose LEO-RobotAgent, a general-purpose language-driven intelligent agent framework for robots. Under this framework, LLMs can operate differen
arXiv:2604.13979v1 Announce Type: new Abstract: Open-world Question Answering (OW-QA) over knowledge graphs (KGs) aims to answer questions over incomplete or evolving KGs. Traditional KGQA assumes a c
arXiv:2604.13386v1 Announce Type: new Abstract: Linear probes can detect when language models produce outputs they 'know' are wrong, a capability relevant to both deception and reward hacking. However
arXiv:2511.16555v3 Announce Type: replace Abstract: Recent advances in stereo matching have focused on accuracy, often at the cost of significantly increased model size. Traditionally, the community h
arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources
arXiv:2601.02902v2 Announce Type: replace-cross Abstract: Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decis
arXiv:2604.14140v1 Announce Type: new Abstract: As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An
arXiv:2602.20913v2 Announce Type: replace Abstract: This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, a