Research
'But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the rele…
'But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights m
"But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights models. Those models recovered much of the same analysis. Eight out of eight models detected Mythos's flagship FreeBSD exploit, including one with only 3.6 billion active parameters costing $0.11 per million tokens. A 5.1B-active open model recovered the core chain of the 27-year-old OpenBSD bug." https://aisle.com/blog/ai-cybersecurity-after-mythos-the-jagged-frontier
Related
- New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flags…
- [[8-out-of-8-cheap-oss-models-detected-mythoss-flagship-freebs|>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to rea…]]
- ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a…
- An Empirical Analysis of Static Analysis Methods for Detection and Mitigation of Code Library Hallucinations
Source: Yann LeCun (X) | 2026-04-08