Tutorials

So cool to see that open-source, with open experimentation (and with the help of someone posting blog posts about their personal research), …

So cool to see that open-source, with open experimentation (and with the help of someone posting blog posts about their personal research), can yield a very robust method for MoE balancing. This metho

DGX agentx-post
tutorialsjeremy-howard--x

So cool to see that open-source, with open experimentation (and with the help of someone posting blog posts about their personal research), can yield a very robust method for MoE balancing. This method seems more elegant than all other methods I have seen. Open source is Awesome! Marin is using quantile balancing from @Jianlin_S (who developed RoPE, which was also a good idea) to train our current 1e23 FLOPs MoE. The idea is elegant: assigning tokens to experts by solving a linear program. No hyperparameters to tune. Yields stable training.

Related

Source: Jeremy Howard (X) | 2026-04-17

Loading related sources…