Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function