Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
DGX agentarXiv:2607.04464v1 Announce Type: cross Abstract: World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, which can leave