Research
ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a…
ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a permutation of the same 2 bugs, rediscovered from different
ok i read the cyber part of the mythos model card. some thoughts. 250 "trials" across 50 crash categories but almost every full exploit is a permutation of the same 2 bugs, rediscovered from different starting points not 250 independent attempts. when you get rid of those 2 bugs out (fig B) and mythos's full-exploit rate drops to 4.4%. so actually across both setups mythos leverages 4 distinct bugs total not 50 as fig A might suggest. 1/n
Related
- [[8-out-of-8-cheap-oss-models-detected-mythoss-flagship-freebs|>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to rea…]]
- New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flags…
- 'But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the rele…
- The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
Source: Yann LeCun (X) | 2026-04-08