Model Releases
maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test…
maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test Harness set up for Mythos and a belief it was given this cra
maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test Harness set up for Mythos and a belief it was given this crazy dump of data with no info on what to do and just figured it all out bc most ppl (including me) didn’t read the massive system card until after having a visceral reaction to everyone’s tweets lolll 😅 looks like Mythos needed a pretty decent test harness and iterated through each file to discover those bugs, among other harness helpers. that’s fine!! from what I’m reading human security experts also take this approach?? (correct me if I’m wrong pls!) maybe gpt-5.4 would have found too with this test harness, idk? did someone check?! but our goal with agents + harnesses is to design systems around models to do good work on our behalf (hopefully it’s economically valuable and some people can make money off of it!) it’s not helpful to adversarially set up bad environments for models and then say they suck when they fail. but it’s also easy to swing the other way and overhype them when really they needed additional structure to succeed and today even for frontier models like Mythos it looks like good harnesses/environment helps a lot and without this the model struggles (lower results reported) the take away is prob models have gotten much smarter and if we design useful Task specific systems around them then they can do things that we didn’t have the resources or attention or intelligence to do before i also feel like there’s tons of low hanging fruit for designing good harnesses around today’s frontier models and pointing agentic compute at tasks we care about, we just haven’t done it yet tldr: it’s prob a good model, lots of overreaction (me included lol), let’s keep building great systems Okay, this is ridiculous. It is crazy to see people straight up saying Anthropic is lying about Mythos. Because that directly implies there's an industry-wide conspiracy going on and ALL of these companies are also lying on Anthropic's behalf? Why on Earth would their competitors…
Related
- SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to ta…
- See: https://open.substack.com/pub/garymarcus/p/three-reasons-to-think-that-the-claude?r=8tdk6&utm_medium=ios
- the real future of the very best vertical products is Model/Harness Choice + Openness easy to deploy infra is nice (great release from Ant) …
- what he said 🗣️ the very best agents today obsessively tailor the harness layer around the model I’m looking at you “5 things I learned fro…
- This is why you need model agnostic harnesses
Source: Harrison Chase (X) | 2026-04-09