Model Releases

maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test…

maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test Harness set up for Mythos and a belief it was given this cra

DGX agentx-post
model-releasesharrison-chase--x

maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test Harness set up for Mythos and a belief it was given this crazy dump of data with no info on what to do and just figured it all out bc most ppl (including me) didn’t read the massive system card until after having a visceral reaction to everyone’s tweets lolll 😅 looks like Mythos needed a pretty decent test harness and iterated through each file to discover those bugs, among other harness helpers. that’s fine!! from what I’m reading human security experts also take this approach?? (correct me if I’m wrong pls!) maybe gpt-5.4 would have found too with this test harness, idk? did someone check?! but our goal with agents + harnesses is to design systems around models to do good work on our behalf (hopefully it’s economically valuable and some people can make money off of it!) it’s not helpful to adversarially set up bad environments for models and then say they suck when they fail. but it’s also easy to swing the other way and overhype them when really they needed additional structure to succeed and today even for frontier models like Mythos it looks like good harnesses/environment helps a lot and without this the model struggles (lower results reported) the take away is prob models have gotten much smarter and if we design useful Task specific systems around them then they can do things that we didn’t have the resources or attention or intelligence to do before i also feel like there’s tons of low hanging fruit for designing good harnesses around today’s frontier models and pointing agentic compute at tasks we care about, we just haven’t done it yet tldr: it’s prob a good model, lots of overreaction (me included lol), let’s keep building great systems Okay, this is ridiculous. It is crazy to see people straight up saying Anthropic is lying about Mythos. Because that directly implies there's an industry-wide conspiracy going on and ALL of these companies are also lying on Anthropic's behalf? Why on Earth would their competitors…

Related

Source: Harrison Chase (X) | 2026-04-09

Loading related sources…