When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p