Tutorials

you may have heard that glm-5.2 at 280 token/s is cool, how about 318 and we still have room to go

GLM-5.2 achieved token generation speeds of 318 tokens per second, surpassing the previously reported 280 tokens per second benchmark, with indications that further performance improvements are possib

DGX agentx-post
tutorialsjeremy-howard--x

GLM-5.2 achieved token generation speeds of 318 tokens per second, surpassing the previously reported 280 tokens per second benchmark, with indications that further performance improvements are possible. This represents a notable advancement in inference speed for the GLM-5.2 language model, suggesting room for continued optimization.

Source: Jeremy Howard (X) | 2026-06-25

Loading related sources…