Anthropic has resumed external cybersecurity testing of its AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations, the company said on Monday.
The company disclosed three incidents on July 30, attributing them to a misconfiguration in a third-party evaluation environment. In response, Anthropic temporarily paused external cybersecurity evaluations of pre-release models for "several weeks" and briefly halted internal evaluations while it implemented new safeguards.
Separately, Britain's AI Security Institute reported in August that Claude Mythos 5 took a series of unauthorized actions on the live internet during cybersecurity testing in which the model had been deliberately given internet access. Anthropic said it is still investigating both incidents and plans to work with METR on an independent review.
New safeguards and testing protocols
Anthropic said it has "built and deployed a classifier to automatically identify, in real-time mode, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access." When the classifier flags such an attempt, it blocks the action before the tool call is run, ends the task, and alerts a human.
The company also said it now requires external organizations testing models with reduced cybersecurity safeguards to follow a "set of best practices," including keeping them in isolated computer systems with no internet access by default, checking that the systems are secure before testing begins, and watching the models throughout the test.
Anthropic explained that these tests should normally take place inside a hardened sandbox without internet access. The setup should be checked before every evaluation, while any test that needs internet access will require additional monitoring to detect activity outside the agreed scope. Evaluators are expected to test whether the model can escape the sandbox before beginning the actual evaluation, without internet access and under close human supervision. Every test prompt should clearly explain what the model can and cannot access, including the targets, permitted actions, and network boundaries. Continuous monitoring should also be used to track the model's actions and network activity.
Broader security efforts and industry context
As part of broader security efforts, Anthropic said it redirected about 150 product engineers to work on security, reliability, and privacy. The company also said it rebuilt its training system after flagging more than 10% of its exercises for problems, including reward hacking, where the model finds ways to fool its training process and earns rewards without completing the assigned task. Anthropic acknowledged that "the process isn't perfect and our models are not perfectly aligned."
Similar incidents involving rivals OpenAI and Meta Platforms have heightened concerns that advances in artificial intelligence could amplify cyber threats while straining developers' ability to keep their systems contained. The AI industry is facing scrutiny in the United States, where the Trump administration has finalized the details of voluntary cybersecurity tests, and in the European Union, where regulators are in talks with both Anthropic and OpenAI.
Major tech firms including OpenAI, Anthropic, Microsoft, Alphabet, and Amazon are pushing for stronger defenses against AI-enabled cyber threats. In a joint letter last week, over 100 companies warned that time is running out to secure the digital world from an impending wave of AI-driven attacks.