🔥Let's talk about Qwen! #AMA
The post announces an “Ask Me Anything” (AMA) about Qwen, the Alibaba‑developed foundation model. It introduces the Qwen Foundation Model Team and directs readers to their GitHub repository @QwenDevs,
Knowledge catalogue
The post announces an “Ask Me Anything” (AMA) about Qwen, the Alibaba‑developed foundation model. It introduces the Qwen Foundation Model Team and directs readers to their GitHub repository @QwenDevs,
Live in Command Code!🎉 Qwen 3.8 Max is now live in Command Code Go. … and it's going open weight!! most importantly, next week open-weights of Qwen3.8-Max and Qwen3.8-27B will be released! 🔹 2.4T para
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for
// Model or Harness // Great paper if you are building with agents in production. (bookmark it) It organizes 41 agent failure modes by the interaction they originate in. Each mode gets assigned to an
Most companies still rent AI by the token, build on someone else’s roadmap, and hope the next model release does not disrupt their systems. At ODSC AI West 2026, @Prof_OZ, Head of AI Developer Educati
Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as they have bad self-knowledge A good approach is to compact t
Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up with 27.3% of the money a human makes. MerchantBench is a 36
Gary Marcus commented on Twitter that OpenAI “distilled” Google’s brand name, prompting a humorous response from Tony Carden. Carden’s reply referenced Google DeepMind’s Project Astra and included a s
Gary Marcus tweeted “OpenAI: No Anthropic: Unlikely,” suggesting that the company does not expect Anthropic to satisfy particular requirements. The tweet references a discussion about the “RPO obligat
Part of the challenge with the AI math advances is the current air of secrecy around them. We don't know how Astra/Fable/Mythos etc are trained. We know many mathematicians have been paid to create tr
“poorly evidenced sensationalism makes for bad policymaking, politics, philanthropy and much else.” beautiful coda for a weekend in which people went completely nuts without even asking for a control
Simon Willison noted that Qwen 3.8 Max and MiniMax‑H3 were released within hours of each other. The MiniMax team announced that MiniMax‑H3 is now publicly available on Hugging Face (https://huggingfac
Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic tools. 開発者の皆様、大変お待たせしました!Sakana Chatのモデルが遂にAPIとして公開です。ぜひお試しください
.@ssankar says Palantir was able to make Nvidia's Nemotron Ultra model 'better than frontier': 'I literally almost felt gaslit when, within 24 hours of getting Nemotron up with no post-training, this
Sudden silence on metrics that used to be shared regularly is a bad sign. If this post below is correct, it doesn’t look great for enthusiastic projections around Anthropic’s Q3. (Author below leaves
Thanks for the recognition. We'll keep building! 🚀 Big news: Qwen3.8-Max by @Alibaba_Qwen just landed at #4 on the Frontend Code Arena leaderboard with a score of 1,668! With 1,668 points, Qwen3.8-Max
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng @Baseten @philipkiely and @waterloo_i
The math still ain’t mathing. Recently we estimated global (ex China) AI revenues of around 200bn annualised, based on four different sources. Just come across a fifth, based on Nvidia inference sales
OpenAI has redesigned the ChatGPT Voice stack—from client to model—to enable continuous audio streaming, allowing GPT‑Live to listen while speaking without interruption. The new architecture supports
The results span sphere packing, coding theory, group theory, quantum complexity, lattice cryptography, extremal combinatorics, and more. Among them: establishing the existence of non-sofic groups and
There should be a public investigation into the AI hacking incidents by OpenAI and Anthropic. We deserve to know whether these labs are genuinely world-class security organizations facing a novel thre
This is a wild result. Locus, the automated research system from @intology, post-trained Qwen3 base models that beat the official human-tuned Qwen3 1.7B Instruct release. SoTA on PostTrainBench! The m
this is the real reason people from OpenAI etc are desperate to shut me up. Astra (which didn’t even have a control group and is maybe not that much better than Fable and certainly not ASI) was perhap
Try Qwen3.8-Max on Hermes Agent and you will have to doubt on how much these open frontier models have caught up with frontier closed models. These new open models are insanely good. Meet Qwen3.8-Max:
two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a
Using optimization at inference time is a foundational concept of Energy-Based Models (EBM) and Objective-Driven AI architectures (ODAI). When the variables to be inferred are continuous, it makes sen
We’re releasing the manuscripts, formal Lean certificates, and reasoning walkthroughs so mathematicians can examine these results and build on their ideas. https://openai.com/index/ten-advances-in-mat
what does the 'last human code review' look like? to @itamar_mar, ceo and founder of @QodoAI, it looks like two agents talking to each other, backed by a context engine specific to your organization.
When an agent starts with shared context, more people can get reliable answers. When truth is shared, AI becomes infrastucture. Read about how our internal truth layer drives our teams at Replit: http
Who’s legally to blame for Anthropic and OpenAI’s autonomous AI hacks? It’s complicated https://techcrunch.com/2026/08/03/whos-legally-to-blame-for-anthropic-and-openais-autonomous-ai-hacks-its-compli
You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... but the hard part is tailoring them so they perform best on you
You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form field values, checkbox states, annotations, embedded images, vec
All other models on the Portal remain 20% discounted, aside from GPT-5.6 Terra and Luna which are 50% off. https://x.com/NousResearch/status/2080039066771337475?s=20 All models are now 20% off for a l
Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3
Mia AI Lab noted on Aug 2 2026 that running DeepSeek v4 Flash via the Hermes agent produces output files superior to those from any other harness tested. The observation was publicly shared and has ga
Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. HUGE victory for neurosymbolic AI, straight from @AnthropicA
Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to
DeepSeek V4 Flash 0731 is impressive. It feels like a capabilities leap for a model in this size class. It is also incredibly cost efficient. The model is available in Bionic, both for running locally
Don’t think I have seen more misconceptions around one model (I count 8) since the fantasies people had before GPT-5 was released. That turned out to be a letdown; so will Astra, for people who are dr
Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion parameter model trained on nothing past 1911 got to general relat
exactly. math isn’t done. not at all. I don’t think being critical of the amazing work AI is doing in pure math is fair to @OpenAI until I can start to say why I feel it’s not yet at the level of our
Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits of open AI models, which helped him defend against a rogue
Hermes Agent is now dramatically more efficient, especially for smaller/weaker/local models! With the help of @nvidia's Nemo Relay and several other strategies Hermes was able to identify a ton of opt
Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai
I wish I had found this sooner. Nous Research launched a FREE Hermes agent Skills Hub. 90,000+ community skills across 200+ categories. Skills from OpenAI, Anthropic, HuggingFace & more. Thank me late
Instead of limiting the progress of AI companies or preventing them from releasing models as proposed in the bipartisan AI Kill Switch Act, Hugging Face CEO Clément Delague says he would rather see Co
It's not time to slow down but to accelerate! The recent AI-powered cyberattacks have everyone talking about the risks of AI. We should. But let's not lose sight of the bigger picture! If we work hard
LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea
More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ I uploaded the source here so it's playable in the browser, forkable etc. https://karpath
New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augmented model ingests existing prefix weights alongside rich te
Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, 'Active Graph Agent Runtime (BabyAGI 4)', on
OpenAI guy lies about my intent. I would absolutely love to see progress in AI for science and medicine. I have said that here, in my books, on countless podcasts, in multiple NYT opeds, in the US Sen
The $0.01 meeting assistant Records meetings on his phone, sends the audio to Telegram, and his locally-hosted AI agent transcribes, identifies speakers, extracts action items, and files tasks -- for
The egg test for AI 'I hereby offer @elonmusk a million dollar bet against his prediction that Optimus will be better than the best humans in surgery by the end of the decade.' ~ @GaryMarcus Haha. I'm
The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress
The insurance negotiator Had Claude Cowork pull 20+ house insurance quotes from every major provider, read the policy fine print for gotchas, and pick the best deal. Credits to Linden Jensen-Page: htt
The QuickBooks killer Replaced his $38/month QuickBooks subscription with a Claude-built workflow that parses his bank statements, categorizes every charge, and feeds a spending dashboard. Credits to
The spam hunter Turned tedious Google Business Profile spam tracking into a Claude Skill that investigates suspicious listings and compiles evidence into a ready-to-submit report for Google. Credits t
This wins the prize for sleazy misrepresentation. @mattShumer took my 2023 argument for hybridizing LLMs with symbolic tools – which is *exactly* what everyone does nowadays – and made it sound like I
To address the limits of deep learning and avoid stalling, the field of AI started by applying patch (1), which started being demoed 9 months later in December 2024 and has now become completely ubiqu