Model Releases
Grounding latent algorithm routing in transformer reasoning
arXiv:2607.24471v1 Announce Type: new Abstract: A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias f
arXiv:2607.24471v1 Announce Type: new Abstract: A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the latent data-generating regime while prompt form is held fixed, remains stable under nuisance perturbations, and is selectively influenced by targeted activation interventions without large losses in answer quality. We introduce ROUTEBENCH, a diagnostic benchmark whose regimes differentially favor global shrinkage, sparsity, robustness, and locality, operationalized by ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across dense decoder-only transformers trained from scratch at 44M-612M parameters, a 306M model closes 80.9 percent of the oracle-routing gap and achieves route F1 of 84.1. The effect remains substantial under natural-language renderings, shuffled supports, lexical paraphrases, and a unified four-way routing setting. Stronger adaptive alternatives, including an input-conditioned soft mixture and an unsupervised Gumbel router, narrow the gap but remain below the 306M and 612M models on route F1 and OOD performance. Probe controls and matched activation-patching controls further show that route-relevant internal directions are decodable and functionally involved in solver-family-consistent output behavior. These results provide controlled evidence that dense transformers trained on ROUTEBENCH can develop route-like internal variables, but they do not establish universal routing in pretrained language models or unrestricted natural-language reasoning.
Related
- Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
- Learning Agent Routing From Early Experience
- TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization
- Learning When to Attend: Conditional Memory Access for Long-Context LLMs
Source: arXiv cs.CL | 2026-07-28