Model Releases
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model f…
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th Jul
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (today, 5th August) https://simonwillison.net/2026/Aug/5/muse-code-and-muse-spark-12/
Related
- Here are Fable's pelicans for the different thinking effort levels, plus how much each one cost to generate via the Claude API
- I printed a custom t-shirt that's an ode to @simonw's Pelican benchmark. My partner says he doesn't get it. But y'all get it, right? RIGHT!?
- Meta says its new AI model is ready to compete on coding
Source: Simon Willison (X) | 2026-08-06