Hook

Their other posts in the index, biggest breakout first.
Cheating is and this AI model cheated. Anthropic just dropped Claude Opus 4.8 and it's now ranked the number one AI model for programming. Except not really. The programming benchmarks for AI models are kind of broken. They're contaminated, too easy, and unreliable. The answers are public and the automated graders they were using were wrong about 32% of the trials. The last Claude model literally cheated on a benchmark by running git log to get the answers. If the old model cheated, the new one probably did too. So a startup called DataCurve built a new benchmark they can't cheat on. And now all of a sudden, ChatGPT is better than this new Claude model. Hmm. If you want more news like this, you can check out my free newsletter.