Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
DGX agentarXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech