Research
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
arXiv:2607.18553v2 Announce Type: replace-cross Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We
arXiv:2607.18553v2 Announce Type: replace-cross Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts success: hidden states plus surface features (length, log-probability) reach AUROC 0.797, versus 0.731 for surface features alone (increment +0.066; task-clustered 95% CI [+0.021, +0.112]; 170 tasks, 680 candidates). On Horizon Logic, the same within-domain comparison gives +0.141 (CI [+0.004, +0.290]; 56 tasks), positive but imprecise. Low-capacity taps also read task-disjoint branch survival and generated-branch correctness, while recurrence moves candidate-quality readability to progressively earlier physical depth. A non-looped control replicates a same-class readout, so recurrence is not required for every readable signal. To test control, we build branch/carry/prune machinery over Ouro's 192-slot recurrent cache, including a bit-exact residual-capture splice that saves up to 88% of per-branch layer passes. No tested frozen intervention produces a validated substantive capability gain. Directional steering is an established negative, and a four-task branch comparison is only a bounded screen. Forced terminal selection beats an exact matched-random null on 39 informative groups, but a selector using no hidden state does as well or better, and malformed-output avoidance explains most of the margin: selection on output form is established, while content-sensitive commitment remains unresolved. A bounded LoRA changes surface behavior without improving net reachability, and a rank-corrected two-null audit does not support the simplest one-dimensional span-misalignment explanation. We call this readable-but-not-yet-usable property operational proto-introspection.
Related
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Dr.LLM: Dynamic Layer Routing in LLMs
- Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
Source: arXiv cs.AI | 2026-07-28