Applications

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to…

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally mul

DGX agentx-post
applicationsethan-mollick--x

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard & pricey to do, but important

Related

Source: Ethan Mollick (X) | 2026-07-30

Loading related sources…