Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
DGX agentarXiv:2605.00994v1 Announce Type: new Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, rese