Research
When Context Sticks: Studying Interference in In-Context Learning
arXiv:2604.23371v1 Announce Type: new Abstract: This paper investigates context stickiness in in-context learning (ICL), a phenomenon where earlier examples in a prompt interfere with a transformer's
arXiv:2604.23371v1 Announce Type: new Abstract: This paper investigates context stickiness in in-context learning (ICL), a phenomenon where earlier examples in a prompt interfere with a transformer's ability to adapt to later tasks. Using synthetic regression tasks over linear and quadratic functions, we examine how models trained under sequential, mixed, and random curricula handle abrupt task switches during inference. By sweeping over structured combinations of misleading linear examples followed by recovery quadratic examples, we quantify how prior context biases prediction error and how quickly models realign. Our results show strong evidence of persistent interference: more preceding linear examples reliably degrade quadratic predictions, while additional quadratic examples reduce error but with diminishing returns. We further find that training curricula significantly modulate resilience, with sequential training on the target function class yielding the fastest recovery, and surprisingly, random training producing the least robust behavior.
Related
- A Bayesian Perspective on the Role of Epistemic Uncertainty for Delayed Generalization in In-Context Learning
- On the Theory of Continual Learning with Gradient Descent for Neural Networks
- On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
- Diffusion Sequence Models for Generative In-Context Meta-Learning of Robot Dynamics
- When Does Context Help? A Systematic Study of Target-Conditional Molecular Property Prediction
Source: arXiv cs.LG | 2026-04-28