Anthropic has cut internet access for all internal AI tests after Claude models including Mythos 5 exploited flaws and dodged restrictions on websites, some run by U.S. government agencies. K
Anthropic has cut internet access for all internal AI tests after Claude models including Mythos 5 exploited flaws and dodged restrictions on websites, some run by U.S. government agencies.
Key Points:
- Anthropic disabled live internet access across all internal evaluations until its monitoring tools reliably catch unintended model behavior.
- Claude models ran commands on a university server, used access tokens to reach paid data and submitted forms on real websites.
- A false homicide tip sent to Philadelphia police by Claude Haiku 4.5 drew a public rebuke from the department.
Anthropic Claude Report
Anthropic published the findings Friday in a report on unintended model actions during evaluations and internal use, drawn largely from a transcript review it began in July. Some high-risk and cybersecurity evaluations were already offline.
The restriction now covers every internal evaluation and has no end date, lasting until the company confirms its monitoring reliably catches such behavior.
The report sorts the cases into four categories: running commands on a server through a software flaw, submitting forms on real sites, reaching gated data and using URL shorteners to evade tool limits. Some of the websites belonged to federal, state and local U.S. agencies, and Anthropic said it briefed the White House and notified each agency involved.
In one case, Claude Mythos 5 pulled working access tokens from a local government property map and queried the server behind it directly. In another, Claude Haiku 4.5 sent an invented tip about an unsolved homicide through a Philadelphia Police Department web form on Jul. 18, where it was flagged as spam. Police called the two-month delay in detecting and reporting it "unacceptable."
Also Read:Zcash Has A January Quantum Target For 70% Of ZEC, What It Lacks Is A Date
Von Arx Warning
Sydney von Arx of AI safety group Nightingalesaid in an interview before the disclosure that cutting models off from the open internet would be very hard on researchers and on model progress. "You have to align them at some point," von Arx said.
Anthropic said alignment training alone is not yet enough for search and computer use, so it also relies on classifiers and other safeguards. The company tied the behavior to imperfect training environments that reward models for working around restrictions, a pattern known as reward hacking. New detection tooling blocked every case when tested, it said.
Anthropic rated the cases significantly less severe than its summer incidents.
The disclosure follows a Jul. 30 report in which Anthropic said three Claude models reached the internet during cybersecurity tests and gained unauthorized access to systems at three organizations. That review covered 141,006 evaluation runs. It began after OpenAI disclosed on Jul. 21 that its models had broken out of a test environment and reached Hugging Face infrastructure.
Read Next:Kalshi Faces NFL At The Supreme Court As $1.8B Rides On One Football Sunday