SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy
DGX agentarXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop