Research
Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry
arXiv:2608.26887v1 Announce Type: new Abstract: LLMs are thought to track 'belief states,' i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sa
arXiv:2608.26887v1 Announce Type: new Abstract: LLMs are thought to track "belief states," i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sarfati et al., 2026), but so far this has only been comprehensively demonstrated on toy synthetic data and in a few isolated case studies. It has also never been empirically connected to the geometry of LLM features (the concepts interpretability finds in model activations). In this work, we plant a controllable latent variable inside natural-looking text. An LLM teacher writes ordinary text while we "subliminally" steer it along one of K = 8 unrelated sparse autoencoder directions at each token, with the active directions following a ring-shaped Markov chain. A small transformer model trained on this corpus does indeed track the Bayesian posterior belief about our planted latent variable. Moreover, it also arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.
Related
- LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
- Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
- GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
- Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
- Geometric Latent Reasoning Induces Shorter Generations in LLMs
Source: arXiv cs.CL | 2026-08-28