InSerHappy

The Trust Protocol: Why a16z's $40M Bet on Vals AI Signals the Next Infrastructure Layer for AI—and Crypto's Parallel

CryptoSam Metaverse

The narrative shift is subtle but seismic. Last week, a16z quietly led a $40 million Series A into Vals AI, a company that promises to evaluate AI models using real code pulled from GitHub’s live pull requests. The pitch is deceptively simple: instead of trusting static benchmarks like GSM8K or HumanEval—which the industry knows have been gamed, leaked, or simply overfit—Vals runs models against tasks extracted from actual developer workflows. It’s a move from academic leaderboards to production truth. And for anyone who has watched the crypto market cycle through similar trust crises, the pattern is hauntingly familiar.

Context: The Benchmark Apocalypse

Over the past three years, I’ve tracked the slow erosion of trust in AI evaluation. Every major model release—GPT-4, Claude 3, Gemini Ultra—came with a parade of benchmark scores. But by 2024, the cracks were visible. GSM8K, a math reasoning benchmark, was found to have overlaps with training data. HumanEval, the coding standard, saw models achieving near-perfect scores through prompt engineering tricks rather than genuine understanding. The industry’s response was to create harder, more dynamic benchmarks like SWE-bench, which evaluates models on real GitHub issues. But SWE-bench itself is a static dataset—once released, it becomes a target for memorization.

Enter Vals AI. The company’s core innovation is not a new model architecture but an evaluation infrastructure that treats every GitHub repository as a potential test set. By pulling historical pull requests (PRs) from any public or private repo, Vals generates a unique, task-specific evaluation: the model must solve the coding problem described in the PR, and the solution is automatically checked against the hidden test suite that was merged into the repo. This is engineering-level innovation, not algorithmic breakthrough. But in a market starving for trust, that might be enough.

The Trust Protocol: Why a16z's $40M Bet on Vals AI Signals the Next Infrastructure Layer for AI—and Crypto's Parallel

Core: The Mechanism of Trust

What Vals has built is a narrative machine disguised as a test harness. The company’s CEO, Vals Smith, has publicly stated that top-tier model vendors—OpenAI, Anthropic, Google, Meta, xAI—now cite Vals’ results in their model cards. If true, this is a watershed moment. It means the industry is beginning to accept external evaluation as a standard part of the credibility stack, analogous to how crypto protocols once began accepting smart contract audits as a prerequisite for listing.

But here’s where the story gets tangled. Vals’ technology relies on access to private repositories. The company claims it can evaluate models on any codebase, including proprietary ones, by using the pull request history. This is a powerful selling point for enterprises: “Get a score on your own code, not some generic benchmark.” Yet it also introduces a subtle risk—the very data used to evaluate might have been seen by the model during training. If the model was trained on the same public GitHub repositories from which Vals extracts PRs, the evaluation is contaminated. Smith’s team argues they can filter by timestamps and use only PRs created after the model’s training cutoff. But how many enterprises trust that? And how many have the resources to verify?

During my years auditing DeFi protocols, I learned that the gap between claimed security and actual security is often measured in unverified assumptions. Vals’ evaluation faces the same problem. The company has not disclosed the full methodology for how it selects PRs, how it prevents leakage, or whether the hidden tests are truly hidden from model vendors. The lack of transparency is a red flag, but it’s also a predictable one—Vals is a startup, not a regulated auditing firm. The question is whether the market will demand independent verification of the verifier.

Yield wasn’t the only thing that collapsed in 2022. Trust in benchmarks did too. Vals is trying to rebuild it, but the foundation is still code on a server.

Contrarian: The Independence Paradox

Here’s the counter-intuitive angle that most coverage misses: Vals AI is funded by a16z, a firm that also invests in many of the AI companies Vals evaluates. The same a16z portfolio includes OpenAI, Anthropic, and several others. This creates a structural conflict of interest. If Vals gives a favorable evaluation to a16z-backed model, it benefits the entire ecosystem. If it gives a harsh one, it risks alienating a major investor’s other holdings. The company’s response will be that it operates independently, with Chinese walls. But the crypto industry has taught us that Chinese walls are only as strong as the incentives to break them.

Moreover, the business model itself is ambiguous. Vals might charge enterprises for evaluations, but it could also charge model vendors for being listed in its “model card” ecosystem. The article from the Web3 monitoring channel notes that the company’s revenue growth claim—“8x this year’s revenue compared to 2025’s full-year”—is vague and unverified. The language suggests a startup desperate to signal traction, but the actual numbers are hidden. In a bear market, survival matters more than gains. Vals may be burning cash to acquire enterprise customers, but the unit economics are unknown.

Another blind spot: the human labor cost. Vals claims to evaluate models across domains like finance, law, and healthcare. But creating robust, domain-specific test sets requires expert annotation. A medical evaluation must be validated by doctors; a legal one by lawyers. The article does not mention whether Vals employs such experts or relies on automated generation. If it’s the latter, the quality of evaluations in high-stakes domains is suspect. The crypto parallel is clear: many DeFi protocols claimed to be “audited” but the audits were shallow, missing critical vulnerabilities. The same risk exists here.

The Trust Protocol: Why a16z's $40M Bet on Vals AI Signals the Next Infrastructure Layer for AI—and Crypto's Parallel

Takeaway: The Next Frontier

Where does this leave us? Vals AI is a bet on the thesis that AI evaluation will become a multi-billion dollar infrastructure layer, just as smart contract auditing became a necessary expense for DeFi. But the path to credibility is long. The company needs to prove that its evaluations are resistant to gaming, that its independence is real, and that its revenue is sustainable. The market’s reaction will be a leading indicator of whether AI vendors are willing to submit to external scrutiny—or whether they will continue to control their own narrative.

Yield wasn’t the only thing that collapsed in 2022. Trust in benchmarks did too. Vals is trying to rebuild it, but the foundation is still code on a server.

For crypto, the lesson is immediate. The same infrastructure that verifies AI models—decentralized evaluation, on-chain attestation, cryptographic proofs of performance—is exactly what the blockchain industry needs to verify AI-generated content in a zero-trust world. I’ve been writing about this convergence since moving to Tel Aviv in 2025, and Vals’ funding is yet another signal that the market is ready for truth-as-a-service. The question is not whether we need it, but whether we can build it without recreating the centralized gatekeepers we’re trying to escape.

Yield wasn’t the only thing that collapsed in 2022. Trust in benchmarks did too. Vals is trying to rebuild it, but the foundation is still code on a server.

In the end, Vals AI is a narrative. The story of a startup that promises to make AI honest. But narratives in crypto have taught us that the most compelling stories are often the ones that hide the most risk. The next year will tell us whether Vals is the beginning of a new infrastructure layer or just another footnote in the cycle of hype and disillusionment. I’ll be watching the pull requests.

Market Prices

Coin Price 24h
BTC Bitcoin
$76,679.3 -1.67%
ETH Ethereum
$2,461.3 -1.58%
SOL Solana
$100.48 -0.71%
BNB BNB Chain
$718.5 -0.22%
XRP XRP Ledger
$1.42 +2.03%
DOGE Dogecoin
$0.0827 -1.14%
ADA Cardano
$0.2052 -1.49%
AVAX Avalanche
$7.56 +1.25%
DOT Polkadot
$0.9895 -1.99%
LINK Chainlink
$11.42 +0.71%

Fear & Greed

69

Greed

Market Sentiment

Event Calendar

{{年份}}
12
05
halving BCH Halving

Block reward halving event

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

18
03
unlock Sui Token Unlock

Team and early investor shares released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

28
03
unlock Arbitrum Token Unlock

92 million ARB released

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

🧮 Tools

All →

Altseason Index

42

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$76,679.3
1
Ethereum ETH
$2,461.3
1
Solana SOL
$100.48
1
BNB Chain BNB
$718.5
1
XRP Ledger XRP
$1.42
1
Dogecoin DOGE
$0.0827
1
Cardano ADA
$0.2052
1
Avalanche AVAX
$7.56
1
Polkadot DOT
$0.9895
1
Chainlink LINK
$11.42

🐋 Whale Tracker

🔵
0xb533...7634
12h ago
Stake
621,363 USDC
🔵
0x82d2...b202
1h ago
Stake
7,408,607 DOGE
🔵
0x4f3e...8d1a
5m ago
Stake
544,612 USDT

💡 Smart Money

0x6e0a...0359
Market Maker
+$0.8M
73%
0x80e4...2440
Market Maker
+$4.1M
70%
0x486c...7b43
Market Maker
+$3.5M
83%