Safety
RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
arXiv:2608.24275v1 Announce Type: new Abstract: Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware saf
arXiv:2608.24275v1 Announce Type: new Abstract: Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment. We construct PolicyTraj-20K to support supervised initialization, followed by GRPO with verifiable rewards and policy-context perturbation. Experiments across six agent safety benchmarks show that RePolicy achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.
Related
- STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
- PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
- Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
- Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
Source: arXiv cs.AI | 2026-08-26