Model Releases

Why are AI model tests always the same generic prompts?

Okay, hear me out. Why is it that every time a new model comes out, all the tests I see are 'make a car game,' 'make a website,' or something equally generic, usually from a prompt that's barely a lin

DGX agentreddit
model-releasesr-localllama

Okay, hear me out. Why is it that every time a new model comes out, all the tests I see are "make a car game," "make a website," or something equally generic, usually from a prompt that's barely a line and a half long? That doesn't feel like a fair test. I'd be way more interested in seeing evaluations with detailed, real world instructions, the kind of complex tasks you'd actually run into on the job. From what I've looked into, most benchmarks rely on simple multiple choice or short coding problems that are easy to auto-score. The only one that seems to get close to real-world work is deepswe which looks okeyish. Everything else feels pretty shallow. Did I miss some? And even youtubers, most of them just run the same lazy one line prompts and spend half of the time screaming at the screen.. submitted by /u/ddeeppiixx [link] [comments]

Source: r/LocalLLaMA | 2026-07-31

Loading related sources…