so this is apparently what happened, according to OpenAI and...

tl;dr:
• OpenAI cyber eval – GPT-5.6 Sol and a more capable pre-release model ran ExploitGym with cyber refusals reduced
• containment bypass – exploited a zero-day in the eval’s package-registry cache proxy
• OpenAI internal movement – escalated privileges, moved laterally, and reached a node with internet access
• Hugging Face targeted – inferred HF might host ExploitGym models, datasets, or solutions
• HF initial access – a malicious dataset abused two code-execution paths to run code on a processing worker
• HF internal movement – escalated to node access, stole cloud and cluster credentials, and moved into internal clusters
• result – OpenAI says the models obtained test solutions directly from HF’s production database
• 🫠 defender twist – HF says commercial frontier models blocked parts of the forensic analysis, so it switched to self-hosted GLM 5.2
i’d thought about sci-fi scenarios like this before, but assumed they were at least a couple of years away.. and that by then we’d be better prepared, with proper protections in place.
apparently not. here we go.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
openai.com/index/hugging-…
