Model Releases
Dynamical phase selection controls compute scaling in looped transformers
arXiv:2608.26556v1 Announce Type: cross Abstract: A looped transformer performs inference by iterating a weight-tied map, making its computation a dynamical process whose cost is set by the resulting
arXiv:2608.26556v1 Announce Type: cross Abstract: A looped transformer performs inference by iterating a weight-tied map, making its computation a dynamical process whose cost is set by the resulting inference dynamics. Here we show that networks with identical architecture and objective, trained to identical accuracy, nevertheless realize distinct dynamical phases depending strongly on initialization, and that the bifurcation defining each phase determines how test-time compute scales. The phases are distinguished by their bifurcation mechanisms, including a saddle-node fold and a Neimark-Sacker-type transition to bounded nonstationary motion. In the fold phase, a one-dimensional normal-form reduction predicts both the relaxation-time and spectral-gap amplitudes from local derivatives of the trained map, yielding the parameter-free relation au(arepsilon)[1-lambda_{max}(-arepsilon)]opi. Composed with a regular distribution of problem difficulty, the same critical slowing down produces the workload-level tail P(au>N)sim N^{-2}. In the Neimark--Sacker phase, the fold scaling law disappears rather than merely changing its prefactor. Thus, test-time compute is not determined by architecture alone. It is governed by the dynamical phase of the solution found by training.
Related
- Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers
- When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers
- Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
Source: arXiv cs.LG | 2026-08-28