Anthropic's RSP Second Report: Self-Audited Safety Is a Variable, Not a Given
Fact: Anthropic published its second Responsible Scaling Policy risk report in June 2025. The document spans 47 pages. It details ASL-2 to ASL-4 thresholds, CBRN capability assessments, and deployment guardrails. But the one thing missing from the PDF is a single external audit signature. For a framework that claims to govern catastrophic AI risks, the absence of independent verification is a structural failure that any risk manager would flag immediately. Protocol integrity is binary; trust is a variable.
Context: The RSP framework is Anthropic's answer to the question of how to safely develop frontier AI models. It borrows heavily from biosafety level (BSL) classifications, mapping model capabilities to four tiers of security. ASL-1 is baseline. ASL-2 requires basic safety measures. ASL-3 triggers strict weight access controls, KYC for API users, and mandatory vulnerability reporting. ASL-4 is reserved for near-AGI existential risks. The first RSP was published in May 2023. The second report is meant to demonstrate that the framework is not a one-time press release but a living process. OpenAI and Google DeepMind have since released similar frameworks, but only Anthropic has published a follow-up report. On the surface, this signals commitment. But the underlying architecture of the report reveals a pattern I have seen before in crypto governance: self-assessment without independent attestation is a governance mirage.
Core: Let me dissect the RSP's second report through the lens of forensic accountability. My background in stress-testing DeFi protocols taught me that any system relying on self-reported data eventually fails when incentives diverge. The RSP's evaluation methodology is proprietary. Anthropic decides which test sets to use, what thresholds to apply, and whether a model has crossed the ASL-3 line. The report discloses that Claude 3.5 Sonnet was evaluated across CBRN, cyberattack capability, and autonomous replication. But it does not release the raw benchmark scores or the specific prompts used. In my 2020 Compound protocol stress test, I found that oracle latency could drain collateral. I submitted a 40-page report. The team dismissed it as theoretical. Six months later, a similar exploit occurred. The lesson: self-assessment by the same entity that builds the system is inherently biased toward underestimation of risk. The RSP's second report replicates this flaw at scale. The ASL-3 threshold for CBRN knowledge diffusion is defined by a set of internal red-teaming exercises. Who defines the red team's competence? What prevents the evaluators from choosing prompts that are easy to pass? The report claims that the model did not cross the ASL-3 threshold in CBRN. But without independent verification, that statement is a claim, not a fact. In my 2023 FTX forensic analysis, I traced $4.3 billion in unbacked USDC transfers. The exchange's own balance sheet showed a surplus. The on-chain data told a different story. The RSP's second report is an off-chain balance sheet. It looks solvent. But the proof is in the audit trail, which remains opaque.
Another structural flaw is the coverage gap. The RSP focuses exclusively on catastrophic risks: weapons of mass destruction, large-scale cyberattacks, and AI self-replication. It ignores the daily risks of bias, discrimination, privacy violations, and psychological manipulation. In 2025, I analyzed ten AI-crypto convergence projects. Eight used centralized cloud servers disguised as decentralized nodes. The marketing spoke of transparency and trustlessness. The reality was a rebranded web2 SaaS platform. The RSP's narrow focus on catastrophic risks is a similar rebranding. It allows Anthropic to claim leadership in AI safety while avoiding the costly and messy work of addressing algorithmic bias. The second report does not mention fairness metrics, differential privacy guarantees, or auditing for hate speech. This selective attention is a red flag. It suggests that the RSP is designed more to protect Anthropic's brand from existential reputational exposure than to protect society from the full spectrum of AI harms. Volatility is the tax on uncertainty. The RSP's second report introduces uncertainty by what it excludes.
The third issue is the lack of a binding enforcement mechanism. The RSP is a policy document, not a smart contract. There is no on-chain governance, no multi-sig requiring multiple signatories for model releases, no slashing conditions for violations. The entire framework relies on Anthropic's internal decision-making. If commercial pressure mounts—say, a competitor releases a more capable model and captures market share—what stops Anthropic from relaxing the ASL-3 thresholds? The report does not answer this. In 2024, I audited a Bitcoin ETF provider's custody solution. The whitepaper promised institutional-grade security. The actual implementation had improper key sharding. I forced them to patch it before launch. The gap between promise and practice is where risk lives. The RSP's second report is a promise. The practice remains unverified.
Contrarian: I must acknowledge what the second report gets right. The RSP is a genuine institutional innovation. It transforms abstract safety debates into operational thresholds. The fact that Anthropic publishes a second report at all puts pressure on competitors to do the same. The framework has already influenced the EU AI Act's risk classification discussions. If Anthropic opens the RSP to third-party audit, it could become a de facto standard for responsible AI development. The bulls argue that self-regulation is better than no regulation, and that the RSP's existence forces the industry to confront catastrophic risks that would otherwise be ignored. This is partially correct. The RSP has created a new category of AI safety professionals. It has driven demand for red-teaming services and model evaluation tools. The second report proves that the framework is not a one-off stunt. But the blind spot is the same one I saw in DeFi's "code is law" philosophy. Smart contract upgrade rights always sit with a few multi-sig admins. The RSP's upgrade rights sit with Anthropic's executive team. The report does not name the individuals who control the ASL thresholds. It does not provide a mechanism for the community to challenge a threshold change. Code is law, but logic is the jury. The RSP's logic is locked inside a black box.
Takeaway: The second RSP report is a signal that Anthropic is willing to be held to a standard. But the standard is self-defined and self-enforced. In the crypto world, we learned the hard way that self-custody without independent verification is a recipe for disaster. The same principle applies to AI safety. Until Anthropic invites an external auditor to verify its ASL classifications, publishes the full test sets, and commits to binding enforcement, the RSP remains a sophisticated risk management theater. The real test will come when the next generation of Claude scores above the ASL-3 threshold. Will Anthropic halt deployment and lose billions in revenue? The second report does not answer that question. Recovery is not a phase; it is a reconstruction. The AI industry needs to reconstruct its safety governance from first principles, with independent auditors, transparent metrics, and enforceable penalties. Anthropic's second report is a step toward that reconstruction, but it is not the foundation. The foundation requires a hard fork from self-regulation to external accountability.