Model Releases

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing…

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders,

DGX agentx-post
model-releasesethan-mollick--x

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders, etc) would certainly help, but the jagged frontier is very much still real. Does an LLM keep the same judgment when you swap the answer order? New LLM Position Bias Benchmark! Judge models compare two lightly edited versions of the same story twice, with the order swapped. The median model flips in 45% of decisive case pairs. GPT-5.4 is worst at 66%!

Source: Ethan Mollick (X) | 2026-04-21

Loading related sources…