What OpenAI and Anthropic Could Do to Make Their AI Safety Claims More Credible
What OpenAI and Anthropic Could Do to Make Their AI Safety Claims More Credible
- OpenAI and Anthropic already publish safety frameworks, evaluations and risk information, but most important decisions still remain largely inside the companies.
- Public trust would improve if truly independent evaluators had stronger access before major model releases.
- Companies should publish failed evaluations and serious incidents, not only successful mitigations and polished system cards.
- Safety thresholds should create clear, binding consequences such as delayed deployment or restricted capabilities.
- The strongest evidence of seriousness would be accepting real commercial costs when a model cannot yet be demonstrated to be safe enough.
OpenAI and Anthropic have spent years publishing increasingly detailed frameworks describing how they intend to manage the risks of advanced artificial intelligence. OpenAI uses its Preparedness Framework and Frontier Governance Framework, while Anthropic operates its Responsible Scaling Policy, Frontier Safety Roadmap and public risk-reporting system.
Those efforts matter. Both companies conduct capability evaluations, use external testers, publish model information and acknowledge risks involving cybersecurity, biological misuse, autonomous behavior and loss of control.
But public skepticism persists for an obvious reason: the organizations developing the technology are still largely responsible for deciding whether their own safeguards are adequate. Trust would become stronger if safety claims were increasingly backed by institutions and rules that the companies themselves could not easily override.
1. Give Independent Evaluators More Authority Before Release
External testing is useful, but public confidence would rise considerably if qualified outsiders could evaluate frontier models independently before deployment and publish meaningful conclusions without company control.
OpenAI explicitly supports third-party evaluations and has argued that independent assessments are important for testing claims about frontier capabilities and safeguards. Anthropic also incorporates external review into parts of its Responsible Scaling Policy and risk-reporting process.
The next step would be to make this process more institutionally independent. Evaluators should have sufficient time, technical access and freedom to design their own tests instead of merely reproducing company-designed benchmarks.
For the most capable models, governments, academic laboratories and specialized safety institutes could receive controlled pre-release access under security restrictions. Their findings should then become part of the release decision rather than optional commentary after the company has already decided to deploy.
That would change the public perception from “the company says its model is safe enough” to “multiple independent groups examined the evidence before deployment.”
2. Publish the Bad Results, Not Only the Reassuring Ones
Transparency becomes convincing when companies disclose information that makes them look worse, not merely when they publish documents supporting a successful launch.
Both companies already publish substantial safety material. OpenAI releases system cards and Preparedness findings for frontier models. Anthropic publishes model information through its Transparency Hub and has introduced recurring Risk Reports under its Responsible Scaling Policy.
A stronger standard would require prominent disclosure of failed evaluations, unresolved weaknesses, evaluation uncertainty and cases where safeguards performed worse than expected.
This matters because frontier-model evaluations are imperfect. Anthropic has publicly discussed cases in which models appeared aware that they were being evaluated, which can reduce confidence that benchmark behavior reflects real-world behavior. OpenAI has similarly acknowledged limitations in frontier evaluations and argues that modern agentic systems require more realistic testing environments.
Publishing uncertainty does not necessarily weaken trust. Pretending uncertainty does not exist does. A safety report that openly says “we do not yet know” can be more credible than twenty pages of immaculate corporate optimism.
3. Make Safety Thresholds Binding Rather Than Advisory
The public is more likely to believe a safety framework when crossing a threshold creates an automatic consequence rather than another internal meeting.
OpenAI's Preparedness Framework classifies serious frontier capabilities using thresholds including High and Critical. Models reaching these levels require stronger safeguards, and Critical capability can trigger safeguards even during development.
Anthropic's Responsible Scaling Policy similarly links stronger model capabilities to stronger security and deployment requirements through its AI Safety Level system and other capability thresholds.
The credibility problem appears when company leadership retains broad discretion over whether mitigations are sufficient. A stronger system would establish predetermined consequences: mandatory external review, reduced access, deployment restrictions or a release delay when specific thresholds are crossed.
The rules could still allow emergency exceptions, but those exceptions should require written justification and later public disclosure. Otherwise, a safety framework risks becoming a sophisticated document explaining why management ultimately remains free to do whatever management decides. Humanity did somehow invent paperwork before accountability.
4. Report Serious AI Incidents Quickly and Consistently
AI companies should adopt something closer to aviation or cybersecurity incident reporting: significant failures should trigger standardized disclosure, investigation and corrective action.
Frontier AI systems are now capable enough that unusual behavior can have consequences beyond embarrassing chatbot responses. Recent cybersecurity evaluations have demonstrated why containment failures, tool access and autonomous action deserve formal incident-management procedures.
OpenAI's Frontier Governance Framework includes incident-response practices, while Anthropic's Responsible Scaling Policy and Transparency Hub describe risk management and reporting systems.
Public confidence would improve if both companies committed to a standardized incident disclosure timeline. Serious events could be reported initially with limited detail when security requires confidentiality, followed later by a fuller investigation describing what happened, why existing controls failed and what changed afterward.
The goal should not be perfection. Complex systems fail. The more useful test of seriousness is whether a company makes failures visible and demonstrably learns from them.
5. Prove That Safety Can Override Commercial Pressure
The most convincing safety commitment is expensive. Companies build credibility when they can show that an unresolved risk actually delayed a product, restricted access or changed development plans.
OpenAI provided one recent example of this principle when it said it temporarily slowed parts of its model-development work after cybersecurity incidents and evidence that an upcoming model might reach its Critical cybersecurity threshold.
Anthropic's Responsible Scaling Policy is also explicitly structured around increasing safeguards as capabilities rise, and the company now publishes Frontier Safety Roadmap goals, risk reports and a noncompliance reporting and anti-retaliation policy.
Those mechanisms become much more persuasive when outsiders can observe situations in which they impose real costs. That could mean delaying deployment, limiting an especially capable feature, spending substantially more on security or allowing independent reviewers additional time before launch.
Employee protections matter as well. Researchers and engineers need credible channels for reporting safety concerns without fearing retaliation. Anthropic already publishes a specific noncompliance and anti-retaliation policy connected to its Responsible Scaling Policy. Making similarly strong protections standard across frontier laboratories would provide another independent source of accountability.
Key Takeaways at a Glance
- OpenAI and Anthropic already have substantial frontier-safety frameworks, so the next credibility step is stronger external accountability.
- Independent evaluators need meaningful pre-release access and freedom to design their own tests.
- Failed evaluations, uncertainties and serious incidents should receive more systematic public disclosure.
- Risk thresholds become more credible when crossing them automatically triggers restrictions or additional review.
- The clearest evidence of seriousness is a willingness to accept delays, costs and lost revenue when safety evidence is insufficient.
| Action | Why It Builds Trust | Strongest Version |
|---|---|---|
| Independent testing | Reduces reliance on self-assessment. | Pre-release access with independent publication rights. |
| Failure disclosure | Shows transparency is not selective. | Publish failed tests and unresolved risks alongside successful ones. |
| Binding thresholds | Turns principles into consequences. | Automatic delay or restricted deployment when thresholds are crossed. |
| Incident reporting | Lets outsiders see how failures are handled. | Standard reporting deadlines and independent investigation. |
| Commercial sacrifice | Demonstrates safety can override launch pressure. | Publicly documented delays or restrictions when evidence is insufficient. |
AI Safety Becomes Credible When Companies Give Up Some Control
OpenAI and Anthropic do not suffer from a shortage of safety documents. Both organizations now maintain sophisticated systems for evaluating frontier capabilities, strengthening safeguards and communicating some of their findings publicly.
The remaining trust problem is institutional. A company cannot fully eliminate skepticism when it develops the model, evaluates much of the model, defines the acceptable level of risk and makes the final commercial decision about deployment.
That does not mean frontier AI companies should surrender every technical decision to regulators or outside committees. It means the highest-risk decisions should gradually become more externally verifiable and less dependent on corporate discretion.
The decisive evidence will not be another promise that safety comes first. It will be the moment when safety requirements visibly cost a company time, money or market advantage and the company follows them anyway.
Sources
OpenAI • Our Updated Preparedness Framework
OpenAI • Frontier Governance Framework
OpenAI • A Shared Playbook for Trustworthy Third-Party Evaluations