Research
New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flags…
New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flagship FreeBSD zero-day, including a 3B model. Rankings reshuff
New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flagship FreeBSD zero-day, including a 3B model. Rankings reshuffle completely across tasks => the AI cybersecurity frontier is super jagged!
Related
- [[8-out-of-8-cheap-oss-models-detected-mythoss-flagship-freebs|>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to rea…]]
- 'But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the rele…
- this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us?
- ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a…
Source: Yann LeCun (X) | 2026-04-08