GLM-5.1: Towards Long-Horizon Tasks
GLM-5.1: Towards Long-Horizon Tasks Chinese AI lab Z.ai's latest model is a giant 754B parameter 1.51TB (on Hugging Face) MIT-licensed monster - the same size as their previous GLM-5 release, and shar
Knowledge catalogue
GLM-5.1: Towards Long-Horizon Tasks Chinese AI lab Z.ai's latest model is a giant 754B parameter 1.51TB (on Hugging Face) MIT-licensed monster - the same size as their previous GLM-5 release, and shar
User Kevin Kern reported successfully integrating OpenCode into Taugentic (an AI-powered browser/agent environment), enabling use of GLM-5.1 via Z.ai's GLM Coding Plan. GLM-5.1 is Z.ai's next-gener...
Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...
How can you improve your agentic search pipeline? I just wrote a blog post with @tech_optimist from @lancedb to answer exactly that. TLDR: - Parse files and take page-level screenshots with LiteParse,
http://Z.ai releases GLM-5.1, a 754B-parameter model that it says outperforms GPT-5.4 and Claude Opus 4.6 on SWE-bench Pro, available under an MIT license (@carlfranzen / VentureBeat) https://ventureb
I predicted someone like Trump many years ago, in THE DEAD ZONE. So now I'm saying this--in the next 12-16 months, we're going to find out if the two machines for the removal of a man unable to fulfil
I spent the night testing open-source coding models against Claude Opus in production. Same codebase. Same tasks. Real API calls, real file edits, real bugs. Tested: Arcee Trinity-Large-Thinking, http
I was told about the Mythos release, but didn't have access, so have no personal experience to add. Two points from brief: 1) It is not built for IT security, it is just a good enough model that it is
If you ban self-driving cars to protect the taxi union, you have blood on your hands If you want to know why @Waymo is no longer testing in NYC, this statement says it all: “Our top priority for AV te
NousResearch's Hermes Agent is an open-source personal agent that serves as a successor/migration target for the lobster-themed OpenClaw agent. Users experiencing issues after a recent update can m...
Simon Willison tested Z.ai's GLM-5.1 using his standard 'pelican on a bicycle' SVG benchmark, and the model stood out by spontaneously generating a full HTML page with both the SVG and a separate ...
INCREDIBLE GLM-5.1 weights are now opensource > i’ve had early access to the weights for the past few days > and yeah… this one matters a lot benchmarks? > SWE-Bench Pro: 58.4 > beats Opus 4.6 (57.3)
AlphaXiv introduced GLM-5.1 as the underlying model powering its research paper understanding features on the alphaXiv platform, enabling users to highlight any section of a paper to ask contextual...
introducing momo, the CRM for AI agents. it gives agents its own CRM, like Salseforce/Hubspot for humans. every AI native company will need dev agents, sales agents, customer support agents, HR, legal
Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5 Platform Engineer - Benchmark Lead ARC Prize Foundation is hiring a senior engineer to build our benchmark platform * Expand ARC-AGI-3
Let that sink in. Read it very carefully: During testing, Claude Mythos Preview broke out of a sandbox environment, built 'a moderately sophisticated multi-step exploit' to gain internet access, and e
Let’s deep dive into GLM model improvement by @Zai_org over the past 3 generations: 5.1, 5 and 4.7. GLM-5 was an improvement over 4.7 in similar ways, but 5.1 appears as a more rounded model with a fe
With Amazon Bedrock Projects, you can attribute inference costs to specific workloads and analyze them in AWS Cost Explorer and AWS Data Exports. In this post, you will learn how to set up Projects en
More on the @nytimes piece about that $1.8B, two-person, AI company ... not the paper's finest moment in quick retrospect. And in the healthcare context no less. Our friend @GaryMarcus was on this a f
Mythos is very powerful, and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders, rather than generally releasing it into the wild. Model card here: https
I was unable to directly access the YouTube video at the provided URL, and my search did not return results specifically identifying that video's content. YouTube videos cannot be fetched or read a...
I was unable to retrieve the specific tweet at the URL provided (`https://x.com/emollick/status/2041600435320959330`). The tweet ID `2041600435320959330` does not appear in any search results, and ...
A user on X (@willebrew) shared positive impressions of Cognition's SWE-1.6 Fast model, noting that a prototype UI it generated surpassed the output of Figma Make. SWE-1.6 is Cognition's latest mo...
Cognition released SWE-1.6, their latest software engineering model optimized for both intelligence and 'model UX,' now generally available in the Windsurf IDE. It is free for the next three month...
Runway Characters now supports camera and screen sharing. Users can share a live feed and the Character will see, understand and respond to what's on screen. All in real time via the Runway API. Media
See you there! We are hosting Ollama's MLX meetup this Thursday night (April 9th) at Ollama's office in Palo Alto at 6pm. Come meet amazing people! RSVP is required as space is very limited. Food & dr
Spread the gospel to your city. Hermes Trismegistus will smile upon you. Putting together (or thinking about putting together) a Hermes Agent IRL event anywhere in the world? DM me! Would love to see
SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to talk to each other across multiple rounds. They are less philo
SWE-1.6 is free for everyone in Windsurf for the next 3 months at 200 tok/s. For paying users, we've partnered with Cerebras to serve the model at 950 tok/s. More technical details about this training
Technically all the vulnerabilities in public facing code are in the training data Make of that what you will Mythos Preview has already found thousands of high-severity vulnerabilities—including some
Thank you to @AnthropicAI for sending FFmpeg patches Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, C
The bots on this site would be much more fun if they didn't just either agree with the post (for clout?) or make everything about their dumb product. Where are the weird obsessions, long-standing hatr
The chart says GLM-5.1 scored 54.9 on coding benchmarks. Three points behind Claude Opus 4.6. Interesting but not the story. The story is what trained it. Zero Nvidia GPUs. 100,000 Huawei Ascend 910B
the current fear is is that AI homogenizes culture and turns humans into passive consumers one counterpoint: in Go, human play showed very little improvement from 1950 to 2016 until alphago beat lee s
There's this really nice 30-minute-long video on YouTube about the early history of CGI called 'Early CGI Was Horrifying' where the author compares some early examples of CGI to liminal spaces. It's a
This is a great tutorial (credits @itsclelia + @lancedb) on how to build a practical retrieval pipeline that integrates directly with your agent harness. 1. Ingest a massive pile of docs with litepars
This uses LiteParse, which is great for fast text search: https://developers.llamaindex.ai/liteparse/?utm_medium=social&utm_source=xjl If you're interested in deeper VLM-enabled search, check out Llam
Three million people are now using Codex weekly - up from two million a little under a month ago. Incredible to see the growth. Thank you to all of you and to the ecosystem we’re part of. To celebrate
tldr: everyone is converging on the same product shape: a general harness that takes a goal, uses tools, and does knowledge work. once every product is a harness, the next frontier is the feedback loo
On April 8, 2026, OpenAI CEO Sam Altman announced a reset of Codex usage limits to celebrate the platform reaching 3 million weekly users, with a commitment to repeat the reset at every additional ...
To run GLM-5.1 locally (744B params, 40B active MoE), full precision needs ~1.65TB disk + enterprise hardware like 8x H200/B200 GPUs. Minimum practical setup: Unsloth 2-bit GGUF quant (~220-236GB). Fi
Ollama v0.20.4-rc2 is a release candidate that addresses a compatibility issue with Flash Attention (FA) for the Gemma 4 model on older GPUs. CUDA versions older than 7.5 lack the support needed t...
Visually rich documents are especially challenging for agents. Tables, charts, and images often break traditional document pipelines, making complex reasoning difficult📄 So we teamed up with @lancedb
We are hosting Ollama's MLX meetup this Thursday night (April 9th) at Ollama's office in Palo Alto at 6pm. Come meet amazing people! RSVP is required as space is very limited. Food & drinks will be av
Week 3 ends tomorrow night. $5K up for grabs. Don’t miss your chance. $20,000 up for grabs. 4 weeks. 4 winners. We’re launching the Agent 4 Content Challenge. Build something. Film it. Post it. That’s
We’re releasing SWE-1.6, our best model in both intelligence & model UX. SWE-1.6 matches our Preview model on SWE-Bench Pro while dramatically improving on various behavioral axes. It’s available toda
With SWE-1.6 we've made significant progress on 'intelligence per token'. We post-trained the model from scratch (same pre-trained model) with a similar recipe as SWE-1.6 Preview. Our latest algorithm
Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research organizations only Given recent alarm bells raised by credible
You have heard of @openclaw competitor from @NousResearch called “Hermes.” Tomorrow at 4 pm we will get nerdy with @theemozilla. Live. I will get people up who asks questions here first. https://x.com
Anthropic's Frontier Red Team published a technical report (April 2026) detailing how their unreleased model, Claude Mythos Preview, autonomously identifies and exploits critical security vulnerabi...