ReDAct: Uncertainty-Aware Deferral for LLM Agents
arXiv:2604.07036v1 Announce Type: cross Abstract: Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However,
Knowledge catalogue
arXiv:2604.07036v1 Announce Type: cross Abstract: Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However,
arXiv:2510.12710v3 Announce Type: replace Abstract: Pre-trained Vision-Language-Action (VLA) models represent a major leap towards general-purpose robots, yet efficiently adapting them to novel, speci
arXiv:2604.07506v1 Announce Type: cross Abstract: Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment qual
arXiv:2604.07298v1 Announce Type: cross Abstract: Multiple Instance Learning (MIL) is the dominant framework for gigapixel whole-slide image (WSI) classification in computational pathology. However, c
arXiv:2604.05268v2 Announce Type: replace-cross Abstract: Multi-modal retrieval-augmented generation (MM-RAG) relies heavily on re-rankers to surface the most relevant evidence for image-question quer
arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an
arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra
arXiv:2601.04268v2 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that
arXiv:2604.07672v1 Announce Type: new Abstract: This paper presents an empirical study of reset-free reinforcement learning (RL) for real-world agile driving, in which a physical 1/10-scale vehicle le
arXiv:2404.15261v4 Announce Type: replace-cross Abstract: We study the linearization of a discrete transportation distance between probability distributions on finite weighted graphs originally due to
arXiv:2603.10512v2 Announce Type: replace Abstract: Artificial intelligence has advanced significantly through the development of intelligent game-playing systems, providing rigorous testbeds for deci
arXiv:2604.06663v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable 'silicon samples' that can approximat
arXiv:2604.07963v1 Announce Type: new Abstract: Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly redu
arXiv:2604.08003v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models
arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t
arXiv:2511.23158v2 Announce Type: replace-cross Abstract: The rapid progress of visual generative models has made AI-generated images increasingly difficult to distinguish from authentic ones, posing
arXiv:2604.06378v1 Announce Type: cross Abstract: In many real-world settings, institutions can and do adjust the consequences attached to algorithmic classification decisions, such as the size of fin
arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp
arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La
arXiv:2604.06802v1 Announce Type: new Abstract: Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competi
arXiv:2510.19225v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for unlocking advanced reasoning capabilities in large language models (LLMs). RL workflows i
arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli
arXiv:2604.07575v1 Announce Type: new Abstract: Autonomous multi-agent target tracking in GPS-denied and communication-restricted environments (e.g., underwater exploration, subterranean search and re
arXiv:2603.06257v2 Announce Type: replace-cross Abstract: In this paper, we propose a novel bounded asymmetric elastic net (L_{baen}) loss function and combine it with the support vector machine (SV
arXiv:2604.06176v1 Announce Type: cross Abstract: We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly
arXiv:2604.07331v1 Announce Type: cross Abstract: Scaling up robot learning will likely require human data containing rich and long-horizon interactions in the wild. Existing approaches for collecting
arXiv:2604.08034v1 Announce Type: new Abstract: Image registration is a fundamental task that aligns anatomical structures between images. While CNNs perform well, they lack rotation equivariance - a
arXiv:2604.06638v1 Announce Type: cross Abstract: Effective detection of unknown network security threats in multi-class imbalanced environments is critical for maintaining cyberspace security. Curren
arXiv:2505.17732v2 Announce Type: replace Abstract: Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have eme
arXiv:2604.06260v1 Announce Type: cross Abstract: Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without a
arXiv:2604.07644v1 Announce Type: new Abstract: We present GPU-SLS, a GPU-parallelized framework for safe, robust nonlinear model predictive control (MPC) that scales to high-dimensional uncertain rob
arXiv:2604.06247v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) remain highly vulnerable to textual and visual jailbreaks, as well as prompt injections
arXiv:2604.07890v1 Announce Type: new Abstract: Highly multiplexed microscopy enables rich spatial characterization of tissues at single-cell resolution, yet most analyses rely on two-dimensional sect
arXiv:2604.07599v1 Announce Type: new Abstract: SANDO is a safe trajectory planner for 3D dynamic unknown environments, where obstacle locations and motions are unknown a priori and a collision-free p
arXiv:2604.07922v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have revolutionized complex problem-solving, yet they exhibit a pervasive 'overthinking', generating unnecessarily long
arXiv:2604.07994v1 Announce Type: new Abstract: Transformer-based approaches have revolutionized image super-resolution by modeling long-range dependencies. However, the quadratic computational comple
arXiv:2604.06409v1 Announce Type: cross Abstract: LLM agents increasingly draft messages on behalf of users, yet users routinely overshare sensitive information and disagree on what counts as private.
arXiv:2604.07159v1 Announce Type: new Abstract: We study the problem of generating synthetic time series that reproduce both marginal distributions and temporal dynamics, a central challenge in financ
arXiv:2604.08542v1 Announce Type: new Abstract: This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown pro
arXiv:2604.08366v1 Announce Type: cross Abstract: Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, th
arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman
arXiv:2604.06603v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong knowledge reserves and task-solving capabilities, but still face the challenge of severe hallucination,
arXiv:2604.08211v1 Announce Type: new Abstract: Modern multimodal generators can now produce scientific figures at near-publishable quality, creating a new challenge for visual forensics and research
arXiv:2604.08501v1 Announce Type: cross Abstract: Science currently offers two options for quality assurance, both inadequate. Journal gatekeeping claims to verify both integrity and contribution, but
arXiv:2604.03134v2 Announce Type: replace Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing
arXiv:2604.06254v1 Announce Type: cross Abstract: With the rapid growth of interconnected devices in Industrial and Medical Internet of Things (IIoT and MIoT) ecosystems, ensuring timely and accurate
arXiv:2506.01062v4 Announce Type: replace Abstract: We introduce SealQA, a new challenge benchmark for evaluating SEarch-Augmented Language models on fact-seeking questions where web search yields con
arXiv:2510.07048v2 Announce Type: replace Abstract: Despite their remarkable natural language understanding capabilities, Large Language Models (LLMs) have been underutilized for retrieval tasks. We p
arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat
arXiv:2604.05650v2 Announce Type: replace Abstract: Video Large Language Models (Video-LLMs) excel in video understanding but suffer from high inference latency during autoregressive generation. Specu
arXiv:2604.08541v1 Announce Type: cross Abstract: Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, we identify a puzzling phenomenon t
arXiv:2407.04183v4 Announce Type: replace Abstract: Large language models (LLMs) are trained on broad corpora and then used in communities with specialized norms. Is providing LLMs with community rule
arXiv:2604.08299v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressivenes
arXiv:2604.07098v1 Announce Type: new Abstract: Large language models often fail on tasks they seem to already understand. In our experiments, this appears to be less about missing knowledge and more
arXiv:2604.08243v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, inherent social biases often cascade throughout the Chain-of-Though
arXiv:2604.07126v1 Announce Type: cross Abstract: Predicting vehicle trajectories plays an important role in autonomous driving and ITS applications. Although multiple deep learning algorithms are dev
arXiv:2604.03128v2 Announce Type: replace Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to pro
arXiv:2604.08532v1 Announce Type: new Abstract: Large-scale multi-view reconstruction models have made remarkable progress, but most existing approaches still rely on fully supervised training with gr
arXiv:2604.06996v1 Announce Type: cross Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend
arXiv:2604.04958v2 Announce Type: replace-cross Abstract: Recent work suggests that large-scale, multi-animal modeling can significantly improve neural recording analysis. However, for functional calc