Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed
arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe