Safety

From the original HuggingFace report, before they knew it was OpenAI. Wild stuff. Things are going to get weird “When we started the log ana…

From the original HuggingFace report, before they knew it was OpenAI. Wild stuff. Things are going to get weird “When we started the log analysis, we first used frontier models behind commercial APIs.

DGX agentx-post
safetyclem-delangue--x

From the original HuggingFace report, before they knew it was OpenAI. Wild stuff. Things are going to get weird “When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.” https://huggingface.co/blog/security-incident-july-2026 we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. https://openai.com/index/hugging-face-model-evaluation-security-incident/

Source: Clem Delangue (X) | 2026-07-22

Loading related sources…