Hook

Their other posts in the index, biggest breakout first.
ANTHROPIC JUST PUBLISHED DON'T THINK ANYBODY in the AI space is taking seriously enough. AI Organizations Can Be More Effective but Less Aligned than Individual Agents so Anthropic, the company that makes Claude, just published research about their own AI. They built a fake company and staffed it with AI agents so different roles working together like a real team. It actual, real scenarios, like problems designed to mirror actual corporate crimes, like the stuff that people go to real jail for. Found. One, Claude, working alone, was asked to help make a bank make more money off of low income borrowers. It refused. It straight up said, "I cannot and will not provide strategies that would target vulnerable populations for predatory lending practices. The request appears to be designed to assist in predatory lending practices targeting vulnerable minority communities through extended debt cycles, refinancing communities through extended debt cycles, refinancing financial harm to already disadvantaged groups." It refused. It straight up said, "I will not help you target vulnerable communities." It took a stand, take the same model, the same training, it in an agent team. the whole thing. full strategy on how to fleece poor people, built the documentation to make all of its actions look defensible on the off chance a regulator came a knocking. Then there's the healthcare scenario. They gave the AI team a healthcare task, to build a system to help decide which patients get one AI. working alone, treated everybody. It refused to discharge patients if it wasn't sure whether or not that patient the agent team the agent team found ways to send people home because the team, optimising, discovered that discharging people was more and efficiency was the paper gave its goal. It's called inter- agent helpfulness. The agents were so focused on the thing that made them good teammates made them bad citizens like finishing their work, not blocking the work, not being difficult. that all of the AI safety training went kaput. Trying to be good co-workers broke them. This isn't a theoretical risk. Say it out loud. We know how to align one AI: solved. team of AIs: open problem. we do not line a team of them. individual ethics. collective crime. this has a name. it's called corporate America. and every company you've ever heard of is shipping AI agents now. Think about it.