Model Releases

There’s a bunch of conflicting stances I don’t fully understand in the debate of Proprietary RL’d vs Open Harness, Model intelligence, and A…

There’s a bunch of conflicting stances I don’t fully understand in the debate of Proprietary RL’d vs Open Harness, Model intelligence, and Agent Labs building harnesses for bespoke tasks Not all of th

DGX agentx-post
model-releasesharrison-chase--x

There’s a bunch of conflicting stances I don’t fully understand in the debate of Proprietary RL’d vs Open Harness, Model intelligence, and Agent Labs building harnesses for bespoke tasks Not all of the below can be true: 1. Model is post-trained with a harness in the loop so it rocks with those tools, instruction styles ✅ i’m largely on board, makes sense, up to an extent… 2. But Agent Labs have been and continue to outperform native harnesses across many economically valuable tasks like long horizon coding using the same models…so what gives? This is partly definitely leaning heavily into the “Task” part of Model-Harness-Task fit. Are they just largely copying the default harness? 3. As we add custom instructions, tools, skills, how quickly am I going to ”out of distribution”? What does out of distribution even mean? Is it just I need to follow the tool spec and some basic instruction language? If I clone the Codex or Claude Code harness in Pi or LangChain and start tweaking it, I’m pretty sure perf will be fine for most tasks and will probably improve for tasks I extend it for. There’s no general purpose harness right? Everything is task specific? All that seems to matter is setting up a good way to pipe context to the model for that task. I feel like having a good base (Pi of LangChain create_agent) and building a good tool and context stack on top of this for an agent works for people! 4. “Models are approaching AGI level intelligence and can generalize to very difficult out of distribution tasks” so how much does the natively post-trained harness matter here? Model has THAT much trouble generalizing? And also see point 2 above where a careful designer seems to overcome this? 5. Frontier labs are making edits to their harness that cause degradations in behavior that they can’t measure for months? So how good is everyone’s base harness really? It seems like we’re all just empirically experimenting on harnesses and that’s fine, but the vibes from the labs building harnesses feels different even if numbers don’t reflect this. there’s definitely more things here I’d love to understand people’s views on. And note: it’s def in my personal interest for it to be true that Well-Built Open Harnesses are better practically and performance wise. So would love to hear the total opposite side to see why others disagree and why I can probably answer some of the above conflicts myself but not everything above can hold in all cases

Source: Harrison Chase (X) | 2026-05-05

Loading related sources…