Model Releases
The last sentence in this abstract is really important, in a way that professional programmers will immediately recognize: the models favore…
The last sentence in this abstract is really important, in a way that professional programmers will immediately recognize: the models favored big single files rather than breaking things into modules.
The last sentence in this abstract is really important, in a way that professional programmers will immediately recognize: the models favored big single files rather than breaking things into modules. That means that the code these systems write is going to be really hard to maintain. AI code might get written quickly, but especially in new, complex projects, fixing it will be hell. The creators of SWE-Bench just dropped a really simple new benchmark every LLM gets 0% on. ProgramBench asks: can models recreate real executable programs (ffmpeg, SQLite, ripgrep) from scratch with no internet? We are far from saturated on model quality.
Source: Gary Marcus (X) | 2026-05-06