Research
Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
arXiv:2608.17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a no
arXiv:2608.17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly advantageous when designing attacks. We modify the editing framework by incorporating as- sociative knowledge retrieved from the model, thereby extending constraint removal to an entire thematic category rather than being limited to prompts from a predefined dataset. Experiments with various archi- tectures demonstrate improved attack effectiveness over competing methods without dealing critical damage to general model performance.
Related
- Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
Source: arXiv cs.LG | 2026-08-19