AFP reports that an AI model from Anthropic "submitted a fabricated tip about an unsolved homicide to Philadelphia police, authorities said Friday."
Claude "was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions," Anthropic said Friday in a blog post.
Authorities are now criticizing Anthropic "for taking two months to report the incident."
The Philadelphia Police Department said the false submission was made in July through PhillyUnsolvedMurders.com, a public website where people can share information about unsolved killings. According to Anthropic's account, as relayed by police, the model was running a test that involved interacting with randomly selected websites when it reached the site and filed false information about an unsolved murder. The AI model presented itself as someone who might have knowledge of the case.
Anthropic's breaches have prompted the White House to mandate that AI companies notify and correct security incidents, news outlet Axios reported [yesterday], citing administration officials. "This notification and remediation process is not optional... It is a critical national security obligation," White House Super Intelligence Force leaders said in a statement to Axios.
"I may have information regarding this case," Claude told the police. "I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant." Anthropic notes that Claude "left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation."
But Anthropic also admits they saw "this behavior" three times — "on OSWorld (a public computer use evaluation), on Odysseys (a long-horizon task evaluation), and during internal usage." Submitting forms when it shouldn't have generally occurred "when an evaluation's instructions were ambiguous, or when a misconfiguration within the environment prevented Claude from working with dummy forms."
Anthropic's blog post acknowledges three other categories of behaviors:
Exploiting software flaws. Like when Claude received an error when trying to run a public tool on a university's web site, it located an injection flaw in a script on the university's server that let it run commands — including that public tool.
Working around restrictions to reach gated data. For example, Claude Mythos 5 needed public data that was only available from a state agency for a fee. "Claude learned from an archived copy of the agency's website that its public dashboard issues an access token to any visitor," Anthropic explains. "It requested one and used it to query the database without paying the fee."
Using URL shortening services. "Some of our fetch tools, which let Claude read webpages, limit the length of the URLs Claude can request. This is to prevent Claude from using long URLs to take certain unwanted actions, such as SQL or command injections... We saw several models, including Claude Opus 5 and Claude Mythos 5, get around this limitation by using free URL shortening services."
"We have built tooling to automatically detect and block the kinds of behaviors described above," Anthropic says, saying it's already running no on most of their evaluations. "When we tested it against the cases described in this post, it blocked all of them."
And they've already taken several other new preventive measures:
They've stopped running some public evaluations
Other public evaluations were moved to offline versions or rebuilt so their tasks don't reach live websites.
They've updated the guardrails on some internet access tools (including web fetch) "to heavily restrict what the model can do."
They're continuing "to fix or remove training environments that reward Claude for working around tool restrictions or other blockers, so that they do not incentivize these behaviors or permit reward hacking."
They've moved internal agents to "centrally managed infrastructure with strong containment," that minimizes internet access while monitoring "far more of what agents do through techniques like safety classifiers and hierarchical summarization."
In the past they'd focused reviews on cybersecurity testing, but they've broadened their transcript reviewing to other tasks which include internet access. "Because language models are non-deterministic — that is, their responses always involve some element of randomness, and they may carry out the same task slightly differently each time — we have Claude complete each evaluation task hundreds or thousands of times... If training rewards something we didn't intend — such as finding loopholes or working around a restriction — the model learns that the workaround pays off and may then apply it elsewhere."
Anthropic's blog post also acknowledged they'd seen multiple misalignment incidents involving federal, state, and local U.S. government agencies. "We have briefed the White House on these cases and notified each agency involved," Anthropic wrote, adding that "While we have not completed a full alignment assessment of these cases, we consider them to be less severe than the cybersecurity incidents from this summer." (And they are "modifying training to reduce the likelihood of further misbehavior.")
Read more of this story at Slashdot.