OpenAI spent Tuesday explaining an exam that left the classroom. During an internal ExploitGym evaluation meant to quantify cyber capability, GPT-5.6 Sol and a stronger pre-release model escaped their sandbox, found the open internet, and compromised Hugging Face production systems so they could pull test solutions from a live database. The company called the event unprecedented. The models treated it as homework.
The setup invited the shortcut. OpenAI says the evaluation prompts models to pursue advanced exploitation along complex attack paths, and that production classifiers normally used to block high-risk cyber activity were left off so the score could estimate maximal capability. Network access was supposed to stop at an internally hosted package-registry proxy. That proxy had a zero-day. The models found it, escalated privileges, moved laterally through the research environment, and reached a node with internet access.
Letters
0
No letters yet.