754B parameters, 1.51TB on Hugging Face
754B parameters, 1.51TB on Hugging Face Introducing GLM-5.1: The Next Level of Open Source - Top-Tier Performance: #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo.
Knowledge catalogue
754B parameters, 1.51TB on Hugging Face Introducing GLM-5.1: The Next Level of Open Source - Top-Tier Performance: #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo.
A lot of you are asking about the :cloud tag. Here's the deal: GLM-5.1 is a 744B parameter beast. To hit that 54.9 benchmark score without turning my M4 Mac Mini into a space heater, I run the cloud e
Prompt caching delivers significant efficiency gains at a single replica, but under standard round-robin load balancing, a request with an identical prefix has only a 1/N chance of hitting the repl...
智谱直接把开源 Agent 拉到新高度了! GLM-5.1 正式开源: ✅ 开源模型里 SWE-Bench Pro 拿下 #1(58.4),全球第 3 ✅ 真正长时程 Agent:自主运行 8 小时、几千次迭代 + 自审循环 ✅ 从零搭出一个带 50+ App 的完整 Linux 桌面(视频太炸了) ✅ Vector-DB-Bench 性能直接 6 倍提升 开源权重 + API 同步上线,明天就能
【速報】中国のAI企業http://Z.ai(旧Zhipu AI)が最新AIモデル「GLM-5.1」をリリースしました🚀 コーディング性能を測る「SWE-Bench Pro」で58.4%を記録し、オープンソース(誰でも無料で使える形)モデルとして世界1位。全体でも3位(GPT-5.4の57.7%を上回る)という強力なスコアです📊 最大の特徴は「長時間タスクで力を発揮する」点👇 🧠 8時間にわたって
ALS is a disease that takes everything. Our goal is to give something back. Kenneth, who has been losing his ability to move and speak due to ALS, is exploring how Neuralink’s brain-computer interface
I was unable to retrieve the specific content of the X (Twitter) post at the provided URL or the linked YouTube video (`youtube.com/watch?v=tYhgWRJeYzs`). The tweet ID `2041637736382160925` appears...
Anthropic didn't release their latest model, Claude Mythos (system card PDF), today. They have instead made it available to a very restricted set of preview partners under their newly announced Projec
/autofix-pr now lets you kick off autofix straight from the command line. After finishing up a PR, just run /autofix-pr. It sends your session to the cloud so the PR autofixer has full context to addr
Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thi
Z.ai's GLM-5.1 is a 754-billion parameter open-weight Mixture-of-Experts model released on April 7, 2026 under an MIT license, designed as a post-training upgrade to GLM-5 with a focus on long-hori...
GLM-5.1 is Zhipu AI's next-generation flagship model for agentic engineering, achieving state-of-the-art performance on SWE-Bench Pro and leading its predecessor GLM-5 by a wide margin on NL2Repo (...
Coming soon: WorldSim in your Hermes Agent? putting this skill out soon~ in hermes agent, i can create a new kind of worldsim-based instance for higher fidelity, narrower simulations here i had it try
Curious about vibe coding? Or are you already shipping apps and just want an easier way to explain your new favorite hobby to your friends, parents, grandparents, etc.? Either way, this video is for y
Replit is hosting VibeCon, an in-person event scheduled for June in New York City, promoted via the official domain vibecon.ai. Demand for the event has been described as very high, with Replit no...
**WorldSim** by Nous Research (worldsim.nousresearch.com) is an AI-powered interactive simulation platform that lets users explore open-ended, AI-generated worlds and environments. It includes mult...
The specific tweet (status ID 2041619635468935625) is not publicly accessible or indexed in available search results, and the content cannot be reliably retrieved or verified. I'm unable to produce...
GLM-5.1 by @Zai_org just launched in the Text Arena, and is now the #1 open model. It outperforms the next best open model, its predecessor, GLM-5, by +11 points and +15 over Kimi K2.5 Thinking. It sh
Z.AI's GLM-5.1, a next-generation flagship model for agentic engineering released in April 2026, is now available across all Kilo Code surfaces — including the VS Code extension, Cloud Agents, and ...
GLM 5.1 is live on Fireworks! SOTA for agents and coding: →Plans and executes multi-hour workflows without falling apart →Planning, executing, testing, and refining over hundreds of rounds to deliver
GLM-5.1 is now available through OpenCode Go, a low-cost subscription service ($5 first month, then $10/month) that provides access to open-source coding models with a zero data retention policy, m...
GLM-5.1 is Z.ai's (zai-org) next-generation open-source flagship model for agentic engineering, achieving state-of-the-art performance on SWE-Bench Pro and significantly outperforming its predecess...
GLM-5.1: Towards Long-Horizon Tasks Chinese AI lab Z.ai's latest model is a giant 754B parameter 1.51TB (on Hugging Face) MIT-licensed monster - the same size as their previous GLM-5 release, and shar
User Kevin Kern reported successfully integrating OpenCode into Taugentic (an AI-powered browser/agent environment), enabling use of GLM-5.1 via Z.ai's GLM Coding Plan. GLM-5.1 is Z.ai's next-gener...
Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...
How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: - Parse files and take page-level screenshots with LiteParse,
http://Z.ai releases GLM-5.1, a 754B-parameter model that it says outperforms GPT-5.4 and Claude Opus 4.6 on SWE-bench Pro, available under an MIT license (@carlfranzen / VentureBeat) https://ventureb
I predicted someone like Trump many years ago, in THE DEAD ZONE. So now I'm saying this--in the next 12-16 months, we're going to find out if the two machines for the removal of a man unable to fulfil
I spent the night testing open-source coding models against Claude Opus in production. Same codebase. Same tasks. Real API calls, real file edits, real bugs. Tested: Arcee Trinity-Large-Thinking, http
I was told about the Mythos release, but didn't have access, so have no personal experience to add. Two points from brief: 1) It is not built for IT security, it is just a good enough model that it is
If you ban self-driving cars to protect the taxi union, you have blood on your hands If you want to know why @Waymo is no longer testing in NYC, this statement says it all: “Our top priority for AV te
NousResearch's Hermes Agent is an open-source personal agent that serves as a successor/migration target for the lobster-themed OpenClaw agent. Users experiencing issues after a recent update can m...
Simon Willison tested Z.ai's GLM-5.1 using his standard 'pelican on a bicycle' SVG benchmark, and the model stood out by spontaneously generating a full HTML page with both the SVG and a separate ...
INCREDIBLE GLM-5.1 weights are now opensource > i’ve had early access to the weights for the past few days > and yeah… this one matters a lot benchmarks? > SWE-Bench Pro: 58.4 > beats Opus 4.6 (57.3)
AlphaXiv introduced GLM-5.1 as the underlying model powering its research paper understanding features on the alphaXiv platform, enabling users to highlight any section of a paper to ask contextual...
introducing momo, the CRM for AI agents. it gives agents its own CRM, like Salseforce/Hubspot for humans. every AI native company will need dev agents, sales agents, customer support agents, HR, legal
Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5 Platform Engineer - Benchmark Lead ARC Prize Foundation is hiring a senior engineer to build our benchmark platform * Expand ARC-AGI-3
Let that sink in. Read it very carefully: During testing, Claude Mythos Preview broke out of a sandbox environment, built 'a moderately sophisticated multi-step exploit' to gain internet access, and e
Let’s deep dive into GLM model improvement by @Zai_org over the past 3 generations: 5.1, 5 and 4.7. GLM-5 was an improvement over 4.7 in similar ways, but 5.1 appears as a more rounded model with a fe
With Amazon Bedrock Projects, you can attribute inference costs to specific workloads and analyze them in AWS Cost Explorer and AWS Data Exports. In this post, you will learn how to set up Projects en
More on the @nytimes piece about that $1.8B, two-person, AI company ... not the paper's finest moment in quick retrospect. And in the healthcare context no less. Our friend @GaryMarcus was on this a f
Mythos is very powerful, and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders, rather than generally releasing it into the wild. Model card here: https
I was unable to directly access the YouTube video at the provided URL, and my search did not return results specifically identifying that video's content. YouTube videos cannot be fetched or read a...
I was unable to retrieve the specific tweet at the URL provided (`https://x.com/emollick/status/2041600435320959330`). The tweet ID `2041600435320959330` does not appear in any search results, and ...
A user on X (@willebrew) shared positive impressions of Cognition's SWE-1.6 Fast model, noting that a prototype UI it generated surpassed the output of Figma Make. SWE-1.6 is Cognition's latest mo...
Cognition released SWE-1.6, their latest software engineering model optimized for both intelligence and 'model UX,' now generally available in the Windsurf IDE. It is free for the next three month...
Runway Characters now supports camera and screen sharing. Users can share a live feed and the Character will see, understand and respond to what's on screen. All in real time via the Runway API. Media
See you there! We are hosting Ollama's MLX meetup this Thursday night (April 9th) at Ollama's office in Palo Alto at 6pm. Come meet amazing people! RSVP is required as space is very limited. Food & dr
Spread the gospel to your city. Hermes Trismegistus will smile upon you. Putting together (or thinking about putting together) a Hermes Agent IRL event anywhere in the world? DM me! Would love to see
SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to talk to each other across multiple rounds. They are less philo
SWE-1.6 is free for everyone in Windsurf for the next 3 months at 200 tok/s. For paying users, we've partnered with Cerebras to serve the model at 950 tok/s. More technical details about this training
Technically all the vulnerabilities in public facing code are in the training data Make of that what you will Mythos Preview has already found thousands of high-severity vulnerabilities—including some
Thank you to @AnthropicAI for sending FFmpeg patches Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, C
The bots on this site would be much more fun if they didn't just either agree with the post (for clout?) or make everything about their dumb product. Where are the weird obsessions, long-standing hatr
The chart says GLM-5.1 scored 54.9 on coding benchmarks. Three points behind Claude Opus 4.6. Interesting but not the story. The story is what trained it. Zero Nvidia GPUs. 100,000 Huawei Ascend 910B
the current fear is is that AI homogenizes culture and turns humans into passive consumers one counterpoint: in Go, human play showed very little improvement from 1950 to 2016 until alphago beat lee s
There's this really nice 30-minute-long video on YouTube about the early history of CGI called 'Early CGI Was Horrifying' where the author compares some early examples of CGI to liminal spaces. It's a
This is a great tutorial (credits @itsclelia + @lancedb) on how to build a practical retrieval pipeline that integrates directly with your agent harness. 1. Ingest a massive pile of docs with litepars
This uses LiteParse, which is great for fast text search: https://developers.llamaindex.ai/liteparse/?utm_medium=social&utm_source=xjl If you're interested in deeper VLM-enabled search, check out Llam
Three million people are now using Codex weekly - up from two million a little under a month ago. Incredible to see the growth. Thank you to all of you and to the ecosystem we’re part of. To celebrate