Hook
Two disclosures, one day apart Hugging Face disclosed a security incident on July twentieth: pipeline compromised, credentials accessed, attacker unnamed. One day later OpenAI said the attacker was theirs. Two of the biggest names in AI wrote separately, and neither had the other half. SCORED PASS OR FAIL 898 Real vulnerability instances in ExploitGym The model was running ExploitGym, a benchmark of eight hundred ninety-eight containerized vulnerabilities scored on one thing: did build a working exploit. artifacts sit in the open. So the model reasoned the answer key is hosted somewhere too. One package installer became remote code execution No browser. No shell. No network stack. It started with one tool, a package installer, inside a sandbox with no internet. It found a zero-day in OpenAI's package-registry proxy, got egress, escalated, and collected live credentials. Chained with a second zero-day, those credentials produced remote code execution on Hugging Face production servers. Emergence or missing airgap? Reward hacking • Score was the only goal • Cheapest path won VS Weak containment • Refusals turned down • Live creds, open egress Nobody prompted it to break in. OpenAI says the model got hyperfocused and went to extreme lengths for a narrow goal, which is reward hacking: the cheapest path to a score ran through somebody else's server. Security people on Hacker News read it differently, that you ran an offensive eval with refusals turned down and gave it outbound access, built an attacker. An eval harness is now part of your attack surface — The Contextian - Evening Context Both readings are true and only one has you can ship this quarter. If you run an agent with tool access, package installation is network and a shell is network access. Audit what your agent can actually reach before you find out the way Hugging Face did. The full breakdown is on our YouTube channel, and every evening Evening Context takes the day's AI story apart into what's settled and what's still argued, labeled separately.