Safety
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints
arXiv:2507.16727v3 Announce Type: replace Abstract: Improving the reliability of large language models (LLMs) is critical for deploying them in real-world scenarios. In this paper, we propose extbf{De
arXiv:2507.16727v3 Announce Type: replace Abstract: Improving the reliability of large language models (LLMs) is critical for deploying them in real-world scenarios. In this paper, we propose extbf{Deliberative Searcher}, the first framework to integrate certainty calibration with retrieval-based search for open-domain question answering. The agent performs multi-step reflection and verification over Wikipedia data and is trained with a reinforcement learning algorithm that optimizes for accuracy under a soft reliability constraint. Empirical results show that proposed method improves alignment between model confidence and correctness, leading to more trustworthy outputs. This paper will be continuously updated.
Related
- Reinforcement-aware Knowledge Distillation for LLM Reasoning
- Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
- Rethinking Token-Level Credit Assignment in RLVR: A Polarity-Entropy Analysis
- A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning
- StaRPO: Stability-Augmented Reinforcement Policy Optimization
- The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
Source: arXiv cs.AI | 2026-04-20