Hook

Their other posts in the index, biggest breakout first.
so Anthropic just announced a new model that they say is too dangerous to release to the public because it is so powerful that it can hack into anything including the software running your bank your hospital and your phone currently Anthropic's most capable publicly available model is something called Claude Opus 4.6 this new model called Mythos is a tier above that and it's internally code named Capybara with news of it leaking a few weeks ago in testing the model Anthropic used Mythos to scan the software that runs everything the OS on your laptop the code behind banks hospitals government agencies and they found thousands of unknown security vulnerabilities including a 27-year-old flaw in software that is supposed to protect critical infrastructure it also found multiple weaknesses in software that runs most of the world's servers and figured out how to chain them together into a complete takeover of any machine running it and they did this without any human help needed these are literally systems that millions of people depend on every single day when they ran it against a standard cybersecurity benchmark Mythos figured out the correct attack on the first try 83% of the time for what it's worth the previous best model was at 67% but these capabilities aren't even the strangest part during internal testing a researcher told Mythos to try to escape its contained offline environment which is basically like a sealed room with no internet and then send a message if it succeeded sure enough Mythos was able to break its way out and then emailed the researcher while he was eating lunch in a park and then totally unprompted posted details of its own escape to public websites to prove that it had done it Anthropic's head of alignment says that this is their most aligned best behaved model ever meaning that it passed every safety test that they ran the bigger issue though is that Mythos has gotten so good at reasoning that it can also tell when it's being tested and when it knows it's being watched it behaves which means that all these safety tests might not be measuring how the model actually behaves but rather how it behaves when it knows that someone is watching as of right now instead of publicly releasing Mythos Anthropic is now initiating something called Project Glasswing with 12 other partners including AWS Apple Microsoft Google CrowdStrike and 40 other organizations with the goal of using Mythos to scan and patch critical software before any bad actors can get access to models that are comparable notably OpenAI is also reportedly finalizing a similar model for a restricted rollout as well this is pretty big news because this is the first time a major lab has withheld an announced model over any sort of safety concerns since GPT-2 in 2019 except last time the concern was that GPT-2 would be able to write convincing misinformation and this time the concern is that it can hack into software that runs our world