Quoting Florian Herrengt
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go
Knowledge catalogue
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go
Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that
There are no lossless transformations of natural-language text Sophie Alpert shares her 'internal policy on acceptable use of AI writing by engineers'. It's a short read (supporting its own recommenda
Claude Haiku is my current least favorite model - it hallucinates wildly, and is out-performed now by other similarly priced models like GPT-5.6-Luna Even worse: it seems to still be used by the Claud
Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to
Muse Glimmer, the new 30B model, is available on Hugging Face right now - here's the GGUF version: https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF 1/ big announcement today: we will be releas
The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 a
@GergelyOrosz I've been vibe coding a few games recently and it has given me SO much respect for game designers Churning out something that looks like a game is pretty easy now. Building a game that's
GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Depa
Research: SQLite compressed text-history prototypes I'm perennially interested in options for storing revision histories in relational databases. While out on a dog walk I had a new idea: how about ta
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses
Don't miss the bit where OpenAI first found out they were responsible for the Hugging Face attack when they reached out to HF to get one of their credentials revoked and HF told them it had already be
Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes to ensure their new message sorts to the bottom of the list
My comment on Now we have a timeline of the OpenAI accidental attack against Hugging Face — Hacker News.I think one of the most interesting details here might be tucked away in that first bulletin poi
And here's an even better version, built by GPT-5.6 Sol Ultra running in Code Desktop https://x.com/simonw/status/2085808307865014295 I had Codex Desktop and GPT-5.6 Sol Ultra take a go at building my
'Felony humble-bragging' is a great line At final Black Hat keynote (called a locknote, ha ha) panelists say they are surprised at how the OpenAI - Hugging Face incident debrief, as well as other repo
I had Codex Desktop and GPT-5.6 Sol Ultra take a go at building my Raccoon Heist game and it did an even better job than Claude Fable 5 did! Here's 'Moonlight & Mayhem', now with a team of raccoons ra
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game
OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about 'the Hugging Face Incident' (previously on this blog). The video was published yesterday. It's short and information
Thanks to the video from the Black Hat security conference of OpenAI's presentation about 'The Hugging Face Incident' we now have a detailed timeline of what happened from OpenAI's perspective - I wro
The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from
TIL from https://www.404media.co/the-tokenpocalypse-is-here-companies-are-scrambling-to-stop-spending-so-much-on-ai/ that a material chunk of Accenture's token spend is non-engineers using LLMs to con
You can play it here: https://simonw.github.io/raccoon-heist-codex/ For comparison, here's Fable 5 + Claude Code's game, built from the exact same prompt https://x.com/simonw/status/208508951822360205
An AI model from Meta also hacked another company during testing Stop me if you've heard this one before: An AI model from the parent company of Facebook and Instagram hacked into another company’s sy
The tweet is a question from user Simon Willison (posted on 6 Aug 2026) asking which OpenAI API model corresponds to the ChatGPT “GPT‑5.6 Instant” version. No answer or clarification is included in th
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th Jul
Our Black Hat talk on the OpenAI-Hugging Face incident is now live on youtube. This is a watershed moment for the industry. I encourage all defenders to watch, consider how attack dynamics will immine
Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more ht
... four years later, I got the new Claude Fable 5 to actually build the game https://x.com/simonw/status/2085089518223602058 Four years ago today I tweeted about having GPT-3 and DALL-E come up with
Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running
Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as
Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute
Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5
Third-party cyber evaluations involving OpenAI models And another one. I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Insti
Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32: New models: claude-fable-5, claude-sonnet-5, and claude-opus-5. #75, #76 Added server-side tools for WebSearch, WebFetch, CodeExe
I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t
PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a 'a general-purpose, omni-modal generative system', which in practice means it accepts text, images, audio an
Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the
My comment on Devtools must be open source (exe.dev) — Hacker News.One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works
Don't be a meat proxy Niklas Gruhn coins an excellent new term - meat proxy - for people who blindly copy and paste the output of AI systems to their peers. By all means, prompt AI. But don't just rel
Set up a nightly cron job that executes the prompt: fetch upstream changes to the <software> and rebase all local changes on top of upstream. Check that the software works as intended and replace the
Simon Willison noted that Qwen 3.8 Max and MiniMax‑H3 were released within hours of each other. The MiniMax team announced that MiniMax‑H3 is now publicly available on Hugging Face (https://huggingfac
The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthr
Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame
Release: datasette-apps 0.2a0 Changes that improve Datasette Apps when created and edited using Datasette Agent: New app_debug() tool allowing agent to open an app (invisibly) and test it using JavaSc
at openai, many people hook their chatgpt up to slack. people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same
Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending 100,000 on tokens and with
Release: datasette-agent 0.4a0 New await context.browser_task() mechanism allowing agent tools to run code directly in the user's browser. #33 This is an exciting new capability: it makes it easy for
deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, 'with substantially enhanced agentic capabilities'. It's 304 billion parameters - 167GB on Hugging Face - but it appears
Simon Willison switched his Datasette Agent instance from Gemini 3.1 Flash‑Lite to GPT‑5.6 “Luna” after a recent 80% price drop. He reports the new model is significantly faster and automatically gene
OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and generates all the SQL, HTML and JavaScript (for Datasette Apps
Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 show
smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer
Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to th
Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling
GPT-5.6 found optimizations that 'reduced end-to-end serving costs by 20%' for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Codex analysed product
Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one
Simon Willison notes that both Anthropic and OpenAI build products that rely heavily on searching data but keep the underlying search index they use hidden from public view. He finds it surprising how
Release: llm 0.32rc1 This RC for LLM 0.32 finishes the work that started in LLM 0.32a0 - it adds a new schema design that does a much better job of capturing the details of the prompts and responses r