Post-Training with Policy Gradients: Optimality and the Base Model Barrier
arXiv:2603.06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards. Given a context oldsymbol{x}, the model must predict the