Model Releases

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and nee…

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other mo

DGX agentx-post
model-releasesethan-mollick--x

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other models on the output of the public questions from GDPval, which tells us nothing. Claude Opus 4.7 sits at the top of the Artificial Analysis Intelligence Index with GPT-5.4 and Gemini 3.1 Pro, and leads GDPval-AA, our primary benchmark for general agentic capability Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over …

Source: Ethan Mollick (X) | 2026-04-18

Loading related sources…