Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories
DGX agentarXiv:2606.04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent