Spectral Rewiring for Exploration, Purification, and Model Merging
arXiv:2607.03065v1 Announce Type: cross Abstract: Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-re