HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
arXiv:2607.26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and ba