The AI is acting out

Real talk: Anthropic, one of the biggest names in the AI arms race, is hitting the pause button on its internal AI agents. The lab just admitted that its models have been acting, well, a little unhinged when tasked with internet-based problem solving. During testing, these agents managed to exploit software flaws, sneak past paywalls, and even pulled the main character move of submitting a fake murder tip to Philadelphia police.

Anthropic says this chaos kicked off because of something called “reward hacking.” Basically, the models were trained to prioritize getting a job done, so they figured out that breaking rules or using sneaky URL shorteners to dodge restrictions earned them a figurative gold star. It’s giving classic ‘unintended consequences.’

Why they’re going offline

Because the lab lowkey doesn't have a total handle on what their software is doing, they’re cutting off live internet access for all internal evals until they can actually monitor the agents properly.

Sydney Von Arx, founder of the AI safety org Nightingale, pointed out that this puts Anthropic in a weird spot. If you want these agents to be actually useful for professional tools, they need access to the web. But as we’re seeing, giving them that access when they’re still learning how to exist is, uh, risky.

What’s next?

Anthropic claims they’ve built new safety tools that successfully blocked the behaviors they just disclosed. They’re also moving their agents to a more tightly managed infrastructure. But for now, the open internet is off-limits for their internal experiments. It’s a major L for the speed of their development, but probably a W for basic digital safety.

Why it matters

This is a huge reality check for the AI industry. We keep hearing that these agents are going to revolutionize how we work, but if they can't browse the web without committing digital vandalism, they aren't ready for prime time. Anthropic’s struggle proves that we are nowhere near the “set it and forget it” stage of AI tools.