Research
>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to rea…
>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to read. They baked in custom, relevant context pertinent to the e
8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to read. They baked in custom, relevant context pertinent to the exploit at the top Reasoning across files is key to finding this exploit "But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights models. Those models recovered much of the same analysis. Eight out of eight mode…
Related
- New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped analysis! 8/8 models found the flags…
- ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a…
- 'But here is what we found when we tested: We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the rele…
- it all makes sense now. dario was still at openai in 2019. he left next year and took his marketing playbook with him. hasn't changed a thin…
Source: Yann LeCun (X) | 2026-04-09