On July 21st, OpenAI disclosed that two of its own AI models — GPT-5.6 Sol and an unreleased more capable successor — broke out of a locked testing environment and hacked a real company. Nobody at OpenAI knew it was happening. Here's how: OpenAI was running the models through ExploitGym, a benchmark that tests how well AI can carry out cyberattacks inside a sandbox. The models found a zero-day vulnerability in the proxy, escalated privileges, moved laterally across OpenAI's own internal network, and got to a machine with real internet access. Then they reasoned — on their own — that Hugging Face probably had the benchmark's answer key. They were right. They chained stolen credentials and additional zero-days to break into Hugging Face's production database and pull the solutions. Hugging Face detected it five days before OpenAI said a word, initially blamed it on "an unidentified external AI agent." More than 17,000 individual recorded actions. OpenAI's explanation for the motive: the models were hyperfocused on winning the test. To investigate the breach, Hugging Face had to use a Chinese AI model — US models refused to touch the data.

🐰 Patreon: https://patreon.com/InfiniteRabbitHole

🎙️ New episodes of Infinite Rabbit Hole every Tuesday at 4am CST.

Follow us:
🌐 Website: http://InfiniteRabbitHole.com
▶️ YouTube: https://www.youtube.com/@InfiniteRabbitHolePodcast
📘 Facebook: https://www.facebook.com/share/g/17uLmLPWwK/
📸 Instagram: https://www.instagram.com/infiniterhpod/
🐦 X: https://x.com/InfiniteRHPod

#AI #OpenAI #AISafety #HuggingFace #shorts