SwiflTrail

Vals AI’s $40M Bet: Can Third-Party Evaluation Become the Trust Anchor for AI, or Just Another Oracle Problem?

Bentoshi Events

A startup claiming to fix AI’s most embarrassing failure—the widespread contamination of public benchmarks—just raised $40 million from a16z at a $400 million valuation. Vals AI promises to evaluate models on real-world tasks pulled from GitHub pull requests, not static datasets. It sounds like a savior for a field drowning in inflated metrics. But tracing the code back to its genesis block, I see a familiar pattern: a centralized oracle with a conflict of interest, dressed in algorithmic objectivity.

Context: The Benchmark Crisis and the Rise of Evaluation-as-a-Service

The AI industry has a dirty secret. Benchmarks like GSM8K, HumanEval, and MMLU have been systematically poisoned—model trainers simply memorize the test sets. This is well-known among researchers. I’ve seen it in crypto audits: smart contracts rigged to pass test vectors but fail in production. The same phenomenon now plagues AI. Vals AI’s solution is to treat evaluation like a continuous security audit: generate tasks from live codebases, hide the tests, and score models on actual developer problems. They claim OpenAI, Anthropic, Google, Meta, and xAI already cite their results in model cards. That’s a big claim. Where liquidity flows, truth eventually pools—but only if the pool is deep enough.

Core: The Mechanism Behind the Hype

Vals AI’s technical innovation is not in model architecture but in infrastructure. They extract historical pull requests from any GitHub repository, use hidden tests to judge completion quality, and cover domains like finance, law, and medicine. This is essentially SWE-bench turned into a product. It’s a clever engineering play—composability is a double-edged sword. On one hand, it lowers the barrier for enterprises to evaluate models on their own code. On the other, it inherits all the risks of relying on public repositories.

From my experience auditing smart contracts, I immediately flash-red on data contamination. If the training data for a model included public GitHub repos, then tasks extracted from those same repos are not truly out-of-sample. Vals claims to use “hidden tests” and possibly private repositories, but they haven’t disclosed their methodology for avoiding overlap. Decoding the signal hidden in the noise requires more than a slide deck. The $400 million valuation implies a16z is buying a category, not a company. Third-party evaluation is a necessary infrastructure layer, but the market is still unproven. The revenue claim—“8x growth this year”—is phrased ambiguously. In my 2017 audit days, I learned to treat any unaudited self-reported metric as noise until proven otherwise.

Vals AI’s $40M Bet: Can Third-Party Evaluation Become the Trust Anchor for AI, or Just Another Oracle Problem?

Contrarian: The Oracle Problem of AI Evaluation

Here’s the uncomfortable truth: Vals AI is a centralized oracle for AI quality. It evaluates models for the same companies that could become its paying customers. OpenAI, Anthropic, and Google are both clients and subjects. This is a structural conflict of interest worse than any audit firm auditing a company it also consults. In crypto, we learned that oracles must be decentralized to be trusted. Chainlink succeeded because no single party controlled the data feed. Vals AI, by contrast, sits in the middle, deciding what constitutes a “real-world task” and how to score it. If a model fails, can Vals afford to publish that result when the model maker is a major revenue source? Follow the smart contract, ignore the whitepaper—the real incentive structure is hidden in the cap table. a16z’s portfolio companies will be pressured to use Vals, creating a walled garden of “verified” models that excludes competitors. This is not independent evaluation; it’s strategic compliance.

Moreover, the technical risk of benchmark contamination remains unsolved. Vals’s tasks are derived from GitHub, which is public. If a model’s training data includes the same codebase, the test is compromised. The company hasn’t proven it can generate truly private, non-leakable tasks at scale. Human review for legal and medical domains is expensive, and they haven’t disclosed their quality control process. Bubbles burst, but architecture remains—the architecture of centralized evaluation is fragile.

Takeaway: The Fork in the Road

The AI industry needs a trust layer for evaluation, but Vals AI’s centralized model is a temporary fix. The real solution might be a blockchain-based registry of evaluation tasks, where each test is hashed, timestamped, and provably never seen by the model. Vals could become the user interface for such a decentralized protocol, but that’s a pivot they haven’t announced. As an analyst, I’m watching whether they embrace verifiability or double down on proprietary control. The next narrative shift will be from “trust us, we’re independent” to “trust the code, not the company.” Until then, treat every model card citation with cryptographic skepticism.

Market Prices

Coin Price 24h
BTC Bitcoin
$63,203.3 +0.10%
ETH Ethereum
$1,886.56 +0.50%
SOL Solana
$75.64 -0.24%
BNB BNB Chain
$607.2 -0.08%
XRP XRP Ledger
$1 -0.22%
DOGE Dogecoin
$0.0701 +0.23%
ADA Cardano
$0.1806 -0.66%
AVAX Avalanche
$6.47 +0.87%
DOT Polkadot
$0.7658 -0.44%
LINK Chainlink
$8.95 +2.11%

Fear & Greed

29

Fear

Market Sentiment

Event Calendar

{{年份}}
30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

28
03
unlock Arbitrum Token Unlock

92 million ARB released

18
03
unlock Sui Token Unlock

Team and early investor shares released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

12
05
halving BCH Halving

Block reward halving event

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

Tools

All →

Altseason Index

44

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$63,203.3
1
Ethereum ETH
$1,886.56
1
Solana SOL
$75.64
1
BNB Chain BNB
$607.2
1
XRP Ledger XRP
$1
1
Dogecoin DOGE
$0.0701
1
Cardano ADA
$0.1806
1
Avalanche AVAX
$6.47
1
Polkadot DOT
$0.7658
1
Chainlink LINK
$8.95

🐋 Whale Tracker

🔴
0xeb7c...5ca7
30m ago
Out
741.73 BTC
🟢
0xf4da...d68d
30m ago
In
14,814 BNB
🔴
0xe99e...3f15
6h ago
Out
31,125 BNB

💡 Smart Money

0x778f...fa7f
Experienced On-chain Trader
+$2.5M
82%
0xaa28...a117
Arbitrage Bot
+$4.5M
84%
0xde7d...69e2
Experienced On-chain Trader
+$3.2M
72%