Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]
I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o