Model Releases
The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything …
The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlate
The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is unlikely to be successful without first figuring that out.
Related
- After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a …
- LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
- NEW: Meta announces Muse Spark. All you need to know: * It's their new multi-modal reasoning model. * Strong at multi-agent orchestration an…
- what he said 🗣️ the very best agents today obsessively tailor the harness layer around the model I’m looking at you “5 things I learned fro…
Source: Francois Chollet (X) | 2026-04-08