CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models
DGX agentarXiv:2606.01695v1 Announce Type: new Abstract: Adversaries can implant latent harmful behavior by poisoning as few as 1% of fine-tuning examples. The contamination is invisible to every output-level