Model Releases
Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…
On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen
On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The comment stresses that many leading benchmarks are likely understated because the evaluated models may not have been run under optimal conditions.
Related
- And now a new DeepSeek model, and appears to be fully open weights. Good benchmarks, but with open models, that isn't always as meaningful. …
- Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even withou…
- The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of m…
- I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and nee…
Source: Ethan Mollick (X) | 2026-08-07