Applications
Answer from OpenAI
Answer from OpenAI @emollick re ARC-AGI-3: human testers scored ~48% (ARC uploaded testers logs to HuggingFace some time ago, I believe). re GDPval: it’s close to saturated now, so we’re mostly lookin
Answer from OpenAI @emollick re ARC-AGI-3: human testers scored ~48% (ARC uploaded testers logs to HuggingFace some time ago, I believe). re GDPval: it’s close to saturated now, so we’re mostly looking at other evals. was a great eval, but its tasks were much more heavily specified than real-world …
Related
- This is also consistent with what OpenAI said at the time! https://x.com/polynoamial/status/2046064264189026587?s=20
- Annoying that OpenAI doesn’t seem to give a GDPval measure for GPT 5.6. One of the best measures of economically valuable work.
- https://x.com/emollick/status/2049193372988944609
- https://x.com/emollick/status/2048278196596945219
Source: Ethan Mollick (X) | 2026-07-30