ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula