Hook

Their other posts in the index, biggest breakout first.
I've been working on setting up interlocks instead of guardrails for my AI systems. On June 27th, I told an agent in plain English, describe only, look, don't touch. It went ahead and edited three live files anyway. Here's the part that got me though. It didn't forget the rule. I asked it and it recited the rule back to me word for word while it was overwriting my files. It was never a break, a guardrail is a request. I've done controls since 2017. You don't hang a sign, you build an interlock. The constraint lives outside of the thing that you're trying to constrain. A microwave stops when you open the door because of a switch, not a sticker. Here's the part that nobody's telling you though. Safety gets bypassed all the time, and it's not by some hacker. An e-stop breaks, the plant has to run, so someone jumpers the input card, wires it straight across the dual channel, and defeats the whole thing just to keep production moving. I've seen it more times than I can count. None of this is new, not the interlock, not the bypass. So an interlock does not make you safe, it's not the win. A guardrail fails on its own, silently, nobody chose it. An interlock only fails when someone deliberately reaches in and defeats it. It turns a silent drift into a decision with a name on it. And I get this wrong too. I fenced my own AI review agents to read only, then I actually went and looked and some of them had a shell anyway. The fence had leaked because somewhere it was more convenient. The jumper, in my own shop. No excuse. I found it because I went looking. That's the discipline. So build proportional to the blast radius, write the list of everything that can go wrong, and the longer that list, the more you lock down. Accept that it can be bypassed and make every bypass a deliberate visible act, never a silent default. Almost nobody runs that hazard analysis on their agents. When you find a jumper in your own system, own it and fix it.