Anthropic’s Safety Pledge Matters Only If Outsiders Can Make It Uncomfortable
Permanent outside reviewers could improve AI accountability. Their real value will be the ability to publish findings a company would rather keep quiet.

Anthropic, which makes the Claude AI assistant, has promised to give outside safety reviewers continuing access inside the company. Chief executive Dario Amodei announced the proposal on September 12, according to Associated Press reporting. The reviewers have not been named and the essay sets no start date.
The useful part of this promise is also the part most likely to create friction. An outside team could tell the public something the company did not choose to announce. That is a meaningful change in who controls the account of AI development. Anthropic deserves credit for proposing it. The public should reserve its trust for the arrangement that actually emerges.
The commitment matters because testing a finished chatbot gives outsiders only a partial view. Amodei wants reviewers to examine how advanced AI is developed, with access similar to employees who assess risks. He is also urging broader cooperation to slow capability growth while safety work catches up.
The public needs a second account
A company report and an outside assessment can contain the same facts yet serve different purposes. The company decides which questions to emphasize and which uncertainties deserve space. A reviewer with meaningful independence can start somewhere else: with an unexplained failure, a missed record or a disagreement that the official account treats as settled. That freedom to choose the question is part of what makes scrutiny valuable.
Consider a hypothetical deployment that passes its planned tests but produces a troubling result just before release. A reviewer seeing only the finished product might never learn how that result was handled. Someone with access during development could ask whether it prompted another test, a restriction or a decision to proceed. The issue is the decision process, not just whether the final model can produce an impressive answer.
His proposed contract would allow reviewers to publish key findings without Anthropic controlling their conclusions. The company could redact security-sensitive, legally privileged, commercially sensitive or third-party confidential material. Unfavorable findings alone would not qualify, and reviewers could disclose when redactions removed information important to their conclusions.
The publication provisions are therefore more consequential than desks or access badges. Access allows reviewers to learn; publication allows other people to benefit from what they learn. An arrangement that delivers the first while frustrating the second could become an excellent private consultancy and a poor accountability mechanism. Those are different services, and the public should not mistake one for the other.
Secrecy cannot settle every dispute
There are legitimate reasons to withhold details that could expose customers or help someone reproduce a dangerous exploit. Demanding that every underlying document become public would ignore those interests. But accepting any invocation of confidentiality would create the opposite problem. A broad label such as commercial sensitivity can cover information that is both important to outsiders and inconvenient to the business.
Those terms make publication rights a useful test of the pledge. An office badge does little for public accountability if important findings never become public. The stated right to flag consequential redactions gives readers something concrete to watch when the first reports appear.
The difficult case will be a disagreement about what readers need to know. A report could say that a weakness exists without explaining how to exploit it. It could describe a decision without identifying a customer. Whether those distinctions work will depend on the actual dispute and the reviewers’ freedom to describe its effect. A promise that everyone will cooperate is least informative when cooperation breaks down.
Anthropic has already signed an agreement with METR, an outside research organization, to investigate specific cybersecurity testing incidents. Its initial term is eight weeks, with extensions possible by agreement. That investigation is separate from the promised continuing review team; METR has not been identified as that team.
Useful scrutiny does not require a scandal
It would also be a mistake to judge the reviewers only by whether they produce alarming headlines. An assessment finding that a company identified a problem early, changed its plans and documented the result could provide valuable evidence of competent management. Independence should make favorable findings more credible too. What matters is whether the reviewers could have reached and published a different conclusion.
AP reports that OpenAI chief executive Sam Altman also committed to employee-like access for outside evaluators. That does not establish common contractual safeguards across the companies. Naming reviewers, beginning their work and publishing findings would turn the commitments into something readers can assess.
Readers do not need to decide whether they share every part of Amodei’s argument about the pace of AI to assess this narrower proposal. A public technology deserves an account that its developer does not entirely control. The decisive milestone will be the first substantive report, especially where it disagrees with the company. Until then, Anthropic has offered a promising institution. It has not yet supplied the accountability that institution is meant to provide.
