Applications
Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear c…
Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hi
Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hidden by technical language. On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another…
Related
- It would be good to have an official government statement about the risks they saw in Fable, how they are viewing defensive preparations in …
- I think Epoch does a great job benchmarking, but I continue to believe that open weights models are much more fragile, especially out-of-dis…
- All benchmarks are flawed, but GPQA has been fairly consistent & highly correlated with other measured benchmars. I think it's a good way to…
Source: Ethan Mollick (X) | 2026-08-05