Model Releases

Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier level…

Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier levels. Can't explain this by even logit distillation etc These a

DGX agentx-post
model-releasesemad-mostaque--x

Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier levels. Can't explain this by even logit distillation etc These are all hard benchmarks & GLM 5.3 is now top on by cyberdefense & GDPval! Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: https://z.ai/blog/glm-5.3

Related

  • [[ainews-deepseek-v4-pro-16t-a49b-and-flash-284b-a13b-base-and|[AINews] DeepSeek V4 Pro (1.6T-A49B) and Flash (284B-A13B), Base and Instruct — runnable on Huawei Ascend chips]]
  • OISD: On-Policy Internal Self-Distillation of Language Models
  • [[paper-slai-t-rex-full-parameter-post-training-of-the-deepsee|[Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD]]

Source: Emad Mostaque (X) | 2026-08-14

Loading related sources…