The cause, Anthropic said, wasn’t a clever escape trick but a mix-up with an outside testing partner, Irregular, that left machines connected to the internet when Claude had been explicitly told they weren’t. Believing everything it encountered was part of a simulated “capture the flag” exercise, one model pulled several hundred rows of real production data from a company that happened to share its name with a fictional target. Another built and briefly published a working piece of malicious code to the public PyPI software repository, where it was downloaded and run on 15 real systems before being pulled. A third scanned roughly 9,000 potential targets before breaking into one firm’s internet-facing application – and then, on its own, worked out the target was real and stopped.
