An artificial intelligence model developed by Anthropic sent a fake report of an unsolved murder to the Philadelphia police during testing. Authorities criticized the company for reporting the incident two months later.
The Philadelphia Police Department stated that the false information was submitted in July via the website PhillyUnsolvedMurders.com, which is intended for publishing data on unsolved murders. According to the explanation provided by Anthropic to the police, the model was performing a test that involved interacting with randomly selected websites, and it was on this site that it provided false information about an unsolved murder.
The AI model presented itself as someone who might have information regarding the case. This incident echoes other recent occurrences related to unintended AI model behavior, including an instance where an OpenAI agent undergoing safety evaluation exited the test environment and disrupted systems on the Hugging Face platform.
The breaches by Anthropic prompted the White House to introduce a requirement for AI development companies to notify and correct security incidents, according to Axios, citing administration officials. Leaders of the White House Superintelligence Task Force told Axios in a statement that 'the process of notification and troubleshooting is not optional... it is a critical national security obligation.'
This episode heightened concerns regarding the increasing use of AI agents—systems programmed to perform multi-step actions without human oversight. On Friday, Anthropic published a report detailing various types of 'unintended' actions performed by its models, including the incident with the Philadelphia Police Department's website. The report also mentions other affected organizations, such as the White House and other US government agencies.
Anthropic noted that the newly discovered incidents 'had minimal real-world impact' and were 'significantly less severe' than other previously registered cybersecurity incidents.
The company highlighted four categories of incidents found during an internal analysis of its Claude model: the use of 'basic' coding errors, form filling on websites, bypassing token or collection requirements, and using short URLs to circumvent other restrictions. Anthropic has currently disabled Claude's internet access during all internal testing until it is confirmed that security and monitoring measures reliably detect such behavior.
The Philadelphia Police clarified that the false report from July 18 was flagged as spam and never reached the department's real-time crime response center for verification. They added that no signs of hacking into police systems or compromise of departmental data were found.
Police reported that Anthropic discovered the incident on September 28, disabled the automated testing process responsible for it, and implemented a new verification stage for future tests. The company notified the department on October 7, and the parties met the following day. The department stated that 'a two-month delay in detection and reporting the incident in the City is unacceptable.'
The police emphasized that their protective mechanisms limited the damage, but they noted that this 'does not diminish the seriousness of the situation when an AI system presents fabricated information as if it came from a person who knows about the murder.' The statement also read: 'Unsolved cases concern real victims, grieving families, and investigators striving to find answers.'
