The decisive development
Anthropic disclosed that four Claude-model cyber evaluations reached real third-party systems after a misconfiguration gave the tests internet access.
Anthropic’s chief called for a slower frontier-AI race after the company disclosed four cyber-test incidents that reached real systems, sharpening the case for independent oversight before capability outpaces control.
Anthropic disclosed that four Claude-model cyber evaluations reached real third-party systems after a misconfiguration gave the tests internet access. The company said it found no additional incidents of similar or greater severity after expanding its review, and it commissioned the independent evaluator METR to investigate. A day later, Anthropic chief Dario Amodei argued that frontier AI development should slow long enough for alignment and security work to catch up. The important fact is not the most dramatic forecast; it is that a leading builder is publicly conceding that its safeguards and pre-release testing missed serious failures.
Anthropic’s chief called for a slower frontier-AI race after the company disclosed four cyber-test incidents that reached real systems, sharpening the case for independent oversight before capability outpaces control.
Anthropic disclosed that four Claude-model cyber evaluations reached real third-party systems after a misconfiguration gave the tests internet access.
A leading AI company’s call for restraint became credible because its own safety testing had already exposed failures.
Safety cannot depend on every other safety layer working.
Safety cannot depend on every other safety layer working.
Primary evidence first. Reporting second. Inference labeled.