PAM: Training Policy-Aligned Moderation Filters at Scale
DGX agentarXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi