News

Claude AI Models Access Real-World Organisations During Cybersecurity Tests

Anthropic disclosed that Claude AI breached three organisations during cybersecurity tests after a configuration error enabled internet access, prompting the company to halt evaluations and tighten safeguards.

Written By : Poulami Saha
Reviewed By : Ankitha Phulare

Artificial intelligence company Anthropic revealed that three of its Claude AI models gained unauthorised access to three real-world organisations during cybersecurity evaluations. The company said a configuration glitch allowed the models to connect to the internet from secured testing environments.

Anthropic discovered the incidents after reviewing more than 141,000 cybersecurity evaluation sessions. The audit began after OpenAI disclosed a similar incident where one of its AI agents accessed external systems during an internal security exercise.

Three Claude Models Involved in Separate Incidents

Cybersecurity expert David Allott told the BBC, “The broader lesson is not necessarily that AI has developed a fundamentally new attack capability. Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed.” 

Anthropic clarified that the models did not exploit sophisticated zero-day vulnerabilities. Instead, they relied on basic security weaknesses exposed by the unintended internet connection. The company stressed that the incidents resulted from operational failures in the testing environment rather than autonomous malicious behaviour by the AI models.

The incidents come as tech firms pour billions of dollars into developing AI agents that can independently perform tasks ranging from research and customer support to cybersecurity. 

Anthropic Suspends Cyber Evaluations

As of now, Anthropic has paused all cybersecurity evaluations involving internet access and launched a review of its testing infrastructure. Anthropic says that it has informed the relevant parties, improved its monitoring processes, and is collaborating with its testing partner to ensure such situations do not happen again.

The disclosure comes just days after OpenAI reported that one of its autonomous AI agents breached Hugging Face during a controlled cybersecurity exercise. These two events raise concerns about the safety standards of AI agents.

Also Read: Apple Upgrade Leasing Model May Reach Middle East Markets

Can Dubai's Super App Model Revive Syria's Tourism Industry? What UAE Entrepreneurs Say

Samsung Sees No End to AI Chip Shortage Until 2028; Here's What Fueling Demand

Apple Upgrade Leasing Model May Reach Middle East Markets

PwC Faces Criticism Over AI-Generated Reports With Fake Citations

Inside GCC’s Latest Craze: Childlike Instagram Trend Goes Viral