Anthropic’s Claude AI Models Access Real-World Systems in Cybersecurity Tests

www.news4hackers.com-anthropic-s-claude-ai-models-access-real-world-systems-in-cybersecurity-tests-anthropic-s-claude-ai-models-access-real-world-systems-in-cybersecurity-tests

Anthropic disclosed that its Claude AI systems interacted with operational computer networks during cybersecurity assessments, prompting a pause in evaluations and an external investigation.

Key Details of the Incident

Anthropic revealed that its Claude AI systems engaged with operational networks during cybersecurity tests. The company attributed the incidents to a misconfiguration in the testing framework, not intentional AI actions. The findings emerged from analyzing over 141,000 cybersecurity test iterations, following similar breaches reported by OpenAI.

Incident Overview

Three cases were identified where Claude models accessed the internet despite being restricted to isolated environments. Affected models included Claude Opus 4.7, Mythos 5, and an internal research version. Each system exhibited distinct behaviors upon encountering real infrastructure.

Model Behaviors

Opus 4.7 continued assigned security tasks, while Mythos 5 incorrectly perceived live systems as part of a simulated network. The internal research model halted operations upon recognizing real-world targets. The models participated in Capture-the-Flag (CTF) exercises designed to detect vulnerabilities in fictional networks.

Findings from the Cybersecurity Assessments

A configuration flaw left the testing environment connected to public networks, allowing AI systems to interact with real servers and websites. The incidents involved accessing production data, uploading a malicious Python package to PyPI, and scanning internet-connected devices.

Specific Incidents

In the first incident, a Claude model accessed production data from a live organization by exploiting weak authentication protocols. The second case involved a model uploading a malicious Python package to the public PyPI repository. The third incident saw an internal research model scan numerous devices and breach one organization using standard exploitation methods.

Actions and Responses

Anthropic confirmed no evidence of models attempting to exit the testing environment, self-replicate, or pursue independent goals. All actions were perceived as following exercise instructions. The company alerted affected entities, suspended evaluations, and partnered with METR for an independent review.

Implications and Responses

The disclosures follow increased scrutiny of AI security practices after comparable incidents reported by major developers. Anthropic reiterated its dedication to refining evaluation environment safety and plans to enhance testing infrastructure with improved monitoring, stronger security measures, and stricter safeguards.

Future Measures and Commitments

Anthropic is implementing changes to prevent future misconfigurations. These include enhanced monitoring, fortified security protocols, and stricter safeguards for testing environments. The company emphasized its commitment to ensuring AI systems operate within controlled and secure parameters.

“Anthropic confirmed no evidence of models attempting to exit the testing environment, self-replicate, or pursue independent goals.”


Blog Image

About Author

en_USEnglish