Preparedness Framework vs Responsible Scaling Policy is a comparison of the two best-known rulebooks AI labs use to decide when a model is too dangerous to release without extra safeguards. OpenAI's Preparedness Framework sorts capabilities into risk categories with "High" and "Critical" thresholds that trigger safeguards before deployment, or even during development. Anthropic's Responsible Scaling Policy started the same "if this capability, then these safeguards" approach in 2023, but its third version, from February 2026, shifted toward public roadmaps, regular Risk Reports and a clear split between what Anthropic will do alone and what it thinks the whole industry should do.
Both are voluntary company policies, not laws. This guide compares them as of October 5, 2026, adds Google DeepMind's Frontier Safety Framework, and explains how they shaped this autumn's releases of GPT-6 Astra, Claude Fable 5.1 and Gemini 4 Argon.
Preparedness Framework vs Responsible Scaling Policy at a glance
| OpenAI Preparedness Framework | Anthropic Responsible Scaling Policy | Google DeepMind Frontier Safety Framework | |
|---|---|---|---|
| First published | December 2023 | September 2023 | May 2024 |
| Current version discussed here | Version 2, April 2025 | Version 3.0, February 2026 | Version 3.1, April 2026 |
| Main risk areas | Biological and chemical, cybersecurity, AI self-improvement | Chemical and biological weapons, plus wider catastrophic risks | Misuse, harmful manipulation, machine learning R&D, misalignment |
| Thresholds | High and Critical capability levels | AI Safety Levels (ASLs), with ASL-3 active since May 2025 | Critical Capability Levels, plus Tracked Capability Levels |
| Who reviews | Internal Safety Advisory Group | Internal processes, with external review of Risk Reports in some cases | Safety case reviews before external launches |
| Public reporting | Capabilities and Safeguards Reports, system cards | Risk Reports every 3 to 6 months, Frontier Safety Roadmap | Model cards and framework reports |
Put simply: all three tie stronger safeguards to dangerous capabilities. OpenAI's version is the most operational, with named categories and two thresholds. Anthropic's latest version is the most candid about uncertainty and leans on transparency. Google's adds risks the others treat lightly, such as harmful manipulation.
How OpenAI's Preparedness Framework works
OpenAI's updated Preparedness Framework, published April 15, 2025, tracks capabilities that meet five criteria: the risk must be plausible, measurable, severe, net new, and instantaneous or irremediable.
- Tracked Categories are areas with mature evaluations and safeguards: biological and chemical, cybersecurity, and AI self-improvement.
- Research Categories are emerging areas OpenAI studies but does not yet gate on, including long-range autonomy, sandbagging, autonomous replication and adaptation, undermining safeguards, and nuclear and radiological risks.
- High capability means a model could amplify existing pathways to severe harm, and it needs safeguards before deployment.
- Critical capability means a model could introduce unprecedented new pathways to severe harm, and it needs safeguards during development as well.
A cross-functional Safety Advisory Group of internal leaders reviews the reports and recommends whether to deploy. OpenAI also says it may adjust its requirements if another developer releases a high-risk system without comparable safeguards, while still keeping protections more stringent than that rival's.
The framework in action. On August 7, 2026, OpenAI said it "cannot rule out critical cyber capabilities" for Astra, the first time it had reached that conclusion. Under the framework, Critical cyber means a model can find and build working zero-day exploits in many hardened systems without human help, or carry out novel end-to-end attacks from only a high-level goal (OpenAI). OpenAI paused Astra work that did not meet new security controls and added monitoring of the model's reasoning. When Astra launched, ordinary users got a restricted version, while vetted defenders got broader access through Daybreak. See our GPT-6 Astra vs Claude Fable 5.1 guide for what that means in ChatGPT.
How Anthropic's Responsible Scaling Policy works
Anthropic's original RSP, written in September 2023, used "if-then" commitments: if a model exceeded a capability level, a stricter set of safeguards, called an AI Safety Level, would apply. Anthropic activated ASL-3 safeguards in May 2025, mainly against help with chemical and biological weapons.
Version 3.0, published February 24, 2026, made three changes:
- Company plans vs industry recommendations. The policy now separates what Anthropic will do "regardless of what others do" from a broader capabilities-to-mitigations map for the whole industry.
- A Frontier Safety Roadmap. Public, nonbinding goals across security, alignment, safeguards and policy, which Anthropic says it will openly grade itself against.
- Risk Reports with external review. Detailed risk assessments published every 3 to 6 months, with independent expert review required in certain circumstances.
Anthropic was unusually frank about why. It said pre-set thresholds were "far more ambiguous than we anticipated," describing a "zone of ambiguity" in biology where models pass most quick tests but evidence of high risk is unclear. It also said the strongest safeguards it had envisioned for later levels "might prove outright impossible to implement without collective action."
The policy in action. Anthropic serves its most capable cyber and biology model in two forms: Claude Fable 5.1 with full safeguards for the public, and Mythos 5.1 with looser limits for vetted organizations. Our Claude Fable vs Opus vs Mythos guide explains the access programs. For Claude users, the practical effect is that some security and biology requests are refused or require verification.
Google DeepMind's Frontier Safety Framework
Google DeepMind's Frontier Safety Framework uses Critical Capability Levels. Version 3, from September 2025, added a level for harmful manipulation and expanded levels for machine learning research that could lead to misalignment, with safety case reviews before external launches and before large internal deployments. Version 3.1, from April 2026, added Tracked Capability Levels.
In practice, Google is holding back Gemini 4 Argon from general release while it strengthens safeguards, giving trusted cyber defenders early access instead; see our Gemini 4 Argon release date guide.
Key differences that matter
- Gating vs transparency. OpenAI's framework is built around clear deployment gates. Anthropic's version 3.0 still has safeguards but puts more weight on public goals and reports.
- Self-improvement. OpenAI tracks AI self-improvement as a full category. Anthropic's Dario Amodei now argues recursive self-improvement is already under way and has called to pace the frontier.
- Competitive escape clauses. OpenAI openly allows adjustment if rivals ship without safeguards. Anthropic's split between unilateral commitments and industry recommendations reflects a similar concern in a different form.
- Outside review. Anthropic requires external review of Risk Reports in some cases and is embedding evaluators such as METR. OpenAI relies mainly on its internal Safety Advisory Group plus external testing. Government testers such as the UK AI Security Institute also test frontier models. In the US, Executive Order 14409 of June 2026 asked agencies to design a voluntary framework for government access to the most cyber-capable models before release; our SI timeline lists it with other federal actions.
Strengths and weaknesses
OpenAI Preparedness Framework
- Pros: specific categories and thresholds; Critical level forces safeguards even during development; has been visibly applied to Astra.
- Cons: decisions rest with internal leaders; research categories such as sandbagging are not gated; competitive adjustment clause.
Anthropic Responsible Scaling Policy
- Pros: candid about uncertainty; regular Risk Reports; external review in some cases; ASL-3 safeguards in force since 2025.
- Cons: roadmap goals are nonbinding; higher safety levels are less precisely defined than before; relies on Anthropic's own judgment about ambiguous thresholds.
Who should read these frameworks
- Businesses choosing a model: they explain why some features are restricted and which verification programs unlock them.
- Policy and compliance teams: laws such as California's SB 53, New York's RAISE Act and the EU AI Act's Codes of Practice now require published frameworks like these, and Anthropic maintains a separate Frontier Compliance Framework for them.
- Anyone following the AGI debate: thresholds focus on specific dangers rather than labels; see has AGI been achieved.
FAQ
What is the difference between the Preparedness Framework and the Responsible Scaling Policy?
Both link stronger safeguards to dangerous AI capabilities. OpenAI's Preparedness Framework uses High and Critical thresholds in named risk categories, reviewed by an internal Safety Advisory Group. Anthropic's Responsible Scaling Policy uses AI Safety Levels and, since version 3.0, adds a public roadmap, Risk Reports and external review.
Are these frameworks legally binding?
No. They are voluntary company policies. However, laws such as California's SB 53 and New York's RAISE Act require frontier developers to publish safety frameworks, so the documents now have legal relevance.
Which models have hit the highest risk levels?
OpenAI said in August 2026 that it could not rule out Critical cyber capabilities for GPT-6 Astra, and it treats Astra as Critical. Anthropic activated ASL-3 safeguards in May 2025. Google is holding Gemini 4 Argon back from general release while it strengthens safeguards.
What is an AI Safety Level?
An AI Safety Level, or ASL, is Anthropic's name for a set of required safeguards. ASL-2 and ASL-3 were defined in detail, with ASL-3 aimed mainly at preventing help with chemical and biological weapons. Higher levels were left less defined.
Did Anthropic weaken its Responsible Scaling Policy?
Opinions differ. Anthropic says version 3.0 replaces commitments that may have been impossible to meet alone with achievable ones plus more transparency. Some critics see the change as softer. Reading the full policy is the best way to judge.