I was halfway through a liquidity-pool oracle review when the number stopped me. Buried in the technical appendix of Anthropic's September 2025 threat intelligence disclosure โ the one every outlet reduced to "Iranian hackers used Claude to plan attacks on US Navy assets" โ was a single figure that no headline carried: their Constitutional Classifiers, the safety layer they advertise as reducing jailbreak success from 86 percent to 0.38 percent, costs roughly 23.7 percent more inference compute. That is the alignment tax. I have spent the better part of a decade watching builders quietly decline to pay it, and watching the receipts vanish into private spreadsheets. Twenty-three percent here, a few basis points of slippage there โ the arithmetic of conscience is always performed off the ledger, which is precisely why nobody can verify it. The report told me a great deal about Anthropic. It told me far more about us.
For context, the disclosure itself is unremarkable in form. Anthropic published a threat intelligence briefing describing three Iran-linked accounts banned for military-adjacent misuse, five biology-related cases flagged as potentially enabling weapons development, and several influence-operation clusters. The underlying machinery is well documented by now: RSP and its ASL-3 deployment threshold, activated in June 2024 to govern chemical, biological, radiological, and nuclear risk; the Constitutional Classifiers paper of January 2025; the standing usage policy that has become the de facto governance surface for a closed model. Three of the four frontier labs now run some version of this ritual โ OpenAI's Disruption Reports began in June 2025, Google's Threat Intelligence Group publishes its Adversarial Misuse series, and Anthropic issues its own threat briefing. Abuse disclosure has become a structural institutional function, not an event.
What troubled me was not the cases. It was the sourcing.
Seven data points, one narrator. No third-party verification. No statement from the accused. No confirmation from any regulator. The attribution to "Iran-linked" accounts rests on registration metadata, payment rails, infrastructure fingerprints and behavioral timing โ non-semantic signals โ stitched together with the kind of semantic content review that requires reading user conversations. Anthropic is simultaneously discoverer of the abuse, provider of the abused system, and author of the disclosure standard. In smart contract terms, that is a single entity writing its own audit attestation and publishing it as fact. My 2017 work on the ZEIP-20 standardization group taught me exactly how dangerous that configuration is. I reviewed 150-plus token proposal drafts over six months and identified 42 critical edge cases in transfer logic โ every single one of which, in the absence of external review, would have shipped as a silent lopsided advantage to centralized validators. The vulnerability was never in the arithmetic. It was in the fact that no independent party could inspect the claim.
Here is the part the crypto press missed while it was busy resharing the Iran headline. The biology cases, by Anthropic's own admission, were "difficult to determine" as weapons-related. A request to enhance mosquito-borne disease transmission reads as a dual-use prompt โ legitimate vector control research on one side, potential weaponization on the other. Calling that a dangerous use case is not a finding. It is an intent inference dressed as a fact. And an intent inference is exactly the class of signal that on-chain attestation was invented to challenge, because inference without a verifiable execution trace is opinion wearing a lab coat.
Which brings me to what I actually care about, and to the silence between the blocks.
Blockchain's central promise โ the reason I left a comfortable municipal auditing job in 2016 to build educational infrastructure in Nairobi โ is that a claim leaves a trail. Every state transition is reproducible. Every transfer has a signature. The reason "code is law" carries moral weight is not that code is objective, but that code is inspectable by anyone with the patience to read it. That is also the reason I distrust it. In 2017, four of those 42 edge cases involved upgrade rights hidden behind three-of-five multi-signature schemes. The contract said governance was decentralized. The deployment script said four human beings and a hardware wallet. Code is law, but only when the law is legible โ and upgrade authority is the one clause that is never in the text.
So when Anthropic tells me a threat actor's activity was "detected" and "blocked" and "reported to authorities," using the completed tense, my auditor reflex asks the only question that matters: when? From use to detection to ban to referral is a chain of events with duration. The report collapses that duration to zero. That collapsed interval โ the window in which the abuse was live and the defense was blind โ is the equity clause. It is the only number that would let an enterprise buyer, a regulator, or a competing lab actually price Anthropic's security posture. And it is absent.
I understand why. If the classifiers are advertised at 0.38 percent jailbreak success and the report simultaneously documents behavior that clearly got past the classifiers, the two claims sit in unresolved tension. Either the general jailbreak resistance does not transfer to motivated, adversarial, multi-turn attacks โ plausible, and honestly the most interesting technical question here โ or the detection is retrospective, caught by human analysts hunting rather than by automated systems alerting. Anthropic does not say which, and the omission is the story. The privacy boundary is where the real argument lives. To identify accounts "shaping public opinion," the classifier must analyze conversation content semantically. That is a surveillance capability sitting adjacent to a privacy promise, and the report does not mention it once. Forty-two edge cases taught me that the unstated clause is always the operative one.
The economics run deeper than the ethics. This report arrived nine days after Anthropic's $13 billion Series F closed at a $183 billion post-money valuation. Across the industry, abuse disclosure has three simultaneous functions that nobody separates: trust capital marketed to compliance-sensitive enterprise buyers, evidence of due diligence assembled prospectively for future litigation and regulatory inquiry, and a legislative posture argument that closed labs self-govern well enough to justify lighter external mandate. All three are legitimate. All three are also forms of narrative production, and the distinction between disclosure and marketing collapses the moment the framing is selective. The report's three themes โ biology, military targeting, influence operations โ map onto national-security anxiety almost perfectly. Fraud, spam, emotional manipulation and non-consensual content, which dominate real-world abuse volume, receive roughly one word. The topic selection is regulatory-agenda shaped, not risk-distribution shaped. Anyone who has watched an NFT collection pump on a story rather than an artwork recognizes the pattern.

I recognize it more painfully than I would like. In 2021, I helped ten Kenyan digital artists structure a DAO-governed royalty system for the Savanna Voices collection: seventy percent of secondary sales returning to creators, hard-coded. Twelve hundred items sold in forty-eight hours, $150,000 raised. The code worked exactly as specified. Within weeks, the speculative narrative swallowed the artistic one, engagement withered, and the royalty contracts sat on-chain executing perfectly for a community that had already left. The technical mechanism was flawless and the human outcome was hollow. That failure reshaped every opinion I hold about disclosure narratives. A truthful report about a broken system is worth more than a beautiful report about a healthy one โ but only if someone outside the system can check it.
And that is the gap the industry refuses to name. When a chain claims decentralization, an auditor can read the upgrade proxy. When an oracle claims low latency, a user can re-derive the price feed. When Anthropic claims responsible disclosure, what can anyone derive? Nothing independent. There is no on-chain analog for "how good is your classifier" because the evidence base is private conversations and proprietary model weights. The closest emerging tools โ zero-knowledge machine learning attestations, deterministic replay of inference, on-chain model cards โ remain experimental and, more damningly, they solve the wrong half of the problem. They could verify that a model ran. They cannot verify how well it refused, because refusal is contextual and the context is exactly what privacy forbids publishing.
So here is my contrarian position, and it will not please either camp. The instinct to demand that AI labs publish everything โ case counts, jailbreak methods, false-positive baselines, threat actor indicators โ is right in spirit and unworkable in practice, because the verification would require the surveillance the disclosure is meant to prevent. Symmetrically, the instinct to accept self-reported safety metrics at face value because they arrive in a PDF with a formal typeface is a category error. We have substituted the appearance of transparency for the apparatus of verification, and we did it in a field that spent fifteen years building the machinery for exactly this and then forgot to deploy it. Cross-lab information sharing exists largely in press releases; Frontier Model Forum working groups produce frameworks, not shared IOC feeds. The five-biology-case question I most want answered โ what is the false positive baseline? โ is unanswerable from outside. Ask the same question of any security disclosure, and you get the same silence.
What stays with me is the older, quieter memory. During the 2022 winter, donations to my educational platform fell sixty percent. I cut the team to four and rewrote forty percent of the curriculum toward risk management and governance, because teaching people to value code they could not read felt like building a library that no one could enter. Building libraries where others build empires was never a slogan; it was the only response I could think of to a market that measures everything and verifies almost nothing. Anthropic has built an impressive institutional function. I would like to believe it. But believing is not my job โ and it should not be yours.

The question I will hold through the next cycle, bull market euphoria notwithstanding, is this: when the first regulated AI audit regime arrives โ California's SB 53 was signed two weeks after this report, the EU AI Act's GPAI transparency obligations took effect in August 2025 โ who will hold the stopwatch? A vendor grading its own paper is not an audit. It is a sermon. And tracing the moral code behind every token is only a virtue when the someone tracing it does not stand to profit from the conclusion.