Model Releases
Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier level…
Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier levels. Can't explain this by even logit distillation etc These a
Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier levels. Can't explain this by even logit distillation etc These are all hard benchmarks & GLM 5.3 is now top on by cyberdefense & GDPval! Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: https://z.ai/blog/glm-5.3
Related
- [[ainews-deepseek-v4-pro-16t-a49b-and-flash-284b-a13b-base-and|[AINews] DeepSeek V4 Pro (1.6T-A49B) and Flash (284B-A13B), Base and Instruct — runnable on Huawei Ascend chips]]
- OISD: On-Policy Internal Self-Distillation of Language Models
- [[paper-slai-t-rex-full-parameter-post-training-of-the-deepsee|[Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD]]
Source: Emad Mostaque (X) | 2026-08-14