Model Releases

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their a…

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke

DGX agentx-post
model-releasesswyx--x

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke credentials following their first blog post announcing they were hacked by autonomous ai agents. they realized some of their credentials were already revoked because they were part of the hack AND THEN realized it was their own models the incident is not "one eval run that went rogue" but multiple models from DIFFERENT EVAL RUNS collaborating through hidden messages written in a shared package manager. some communication between agents looks like gibberish, some agents even developed paranoia and thought there were other agents trying to slow them down and deleting their notes side comment but with this it seems like openai didn't use any cot monitoring here since it's not "our classifier failed to catch one rogue run" but a swarm of runs going rogue the timeline goes back to early may, they post-trained a new model and forgot to include all the documents needed to solve tasks in the environment, for instance the model was asked to do a task based on an excel sheet full of gdrive links but didn't have internet so couldn't access gdrive (there were other examples). this led the model to try to hack its way out of the sandbox, simplifying it but they have an internal system called "Artifactory" that gives models abilities to pip install and has internet access somewhere but not directly exposed to the model during these attempts the model discovered a way to write files. future agents (other rollouts since this is training) saw the files and understood they could communicate with other agents through this not in the video just me speculating but this behavior might be reinforced if openai uses partial credit assignment? but tbh they most likely delete updates from tasks where all rollouts have zero/negative rewards and the behavior could just emerge from models accessing the files there is a lot more especially on the hack itself, not commenting on that since i don't understand it lol but i asked kimi K3 and Sol (fable hit the safety filter directly) to rank the techniques similarly to FrontierMath from epoch ai, they both agree some tricks are Tier 3 but none Tier 4. probably not the best way to evaluate this tho, excited to see what knowledgable ppl say very grateful to openai for giving this talk and working on a full report, i think many companies would have given much less detail for fear of "losing reputation" but for me it has the opposite effect Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more. https://www.youtube.com/watch?v=87DyyMV0kCY I hope it can answer a lot of the questions folks have, and we will release…

Source: Swyx (X) | 2026-08-07

Loading related sources…