Model Releases
I spent the night testing open-source coding models against Claude Opus in production. Same codebase. Same tasks. Real API calls, real file …
I spent the night testing open-source coding models against Claude Opus in production. Same codebase. Same tasks. Real API calls, real file edits, real bugs. Tested: Arcee Trinity-Large-Thinking, http
I spent the night testing open-source coding models against Claude Opus in production. Same codebase. Same tasks. Real API calls, real file edits, real bugs. Tested: Arcee Trinity-Large-Thinking, http://Z.AI GLM-5.1, and Claude Opus 4.6. Results were eye-opening. Here's what actually happened. Introducing GLM-5.1: The Next Level of Open Source - Top-Tier Performance: #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo. - Built for Long-Horizon Tasks: Runs autonomously for 8 hours, refining strategies through thousands of iterations. Blog…
Related
- The chart says GLM-5.1 scored 54.9 on coding benchmarks. Three points behind Claude Opus 4.6. Interesting but not the story. The story is wh…
- http://Z.ai releases GLM-5.1, a 754B-parameter model that it says outperforms GPT-5.4 and Claude Opus 4.6 on SWE-bench Pro, available under …
- GLM-5.1 by @Zai_org just launched in the Text Arena, and is now the #1 open model. It outperforms the next best open model, its predecessor,…
- GLM-5.1 by @Zai_org is now #3 in Code Arena - surpassing Gemini 3.1 and GPT-5.4, and now on par with Claude Sonnet 4.6. The first frontier l…
- INCREDIBLE GLM-5.1 weights are now opensource > i’ve had early access to the weights for the past few days > and yeah… this one matters a lo…
- Check out the GLM-5.1 first impressions with Peter on our YouTube https://www.youtube.com/watch?v=f11tVBXWr2g
Source: Zhipu AI (X) | 2026-04-07