Model Releases

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences b…

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect

DGX agentx-post
model-releasesethan-mollick--x

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect benchmarks understate progress, they are built for models, not harnessed agents

Source: Ethan Mollick (X) | 2026-05-03

Loading related sources…