In the quiet of a September press release, OpenAI announced that GPT-6 Astra had scored 98.6% on ARC-AGI-3, the industry's most grueling benchmark for abstract reasoning. The number was arresting. Anthropic's Claude Opus 5, the previous leader, managed 30.2%. A jump of nearly 70 percentage points in six months would be, if true, the kind of discontinuity that rewrites entire fields. [[14]]

But here is where the code diverges from the headline. The 98.6% was not achieved under the standard ARC-AGI-3 harness. It was achieved using OpenAI's proprietary Responses API harness, with two settings modified: continuous conversation retention and custom compaction. Under the standard harness, the same model scored 62.7% at a cost of $26,000 per game. Under the provider adapter harness, it hit 99.9% at $19,000. [[18]] The ARC Prize Foundation itself acknowledged that "those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model." [[16]]
The benchmark measures Astra and OpenAI's agent systems together, not the model in isolation. This is not a footnote. It is the entire story.
Tracing the code back to the silence of 2017
Eight years ago, during the ICO mania that consumed crypto, I spent three months reverse-engineering Bancor's V1 Solidity contracts. I found seven integer overflow vulnerabilities that their marketing materials had no incentive to disclose. The pattern was already forming: metrics are optimized for the audience, not for truth. TVL was inflated by counting the same liquidity across multiple pools. Transaction volumes were washed through self-trading bots. The numbers were technically real in isolation, but meaningless in context.
Today, the same pattern repeats in a different domain. The ARC-AGI-3 benchmark was designed precisely to resist memorization and pattern-matching — to put models in "uncharted interactive territory where they cannot simply rely on training data to find an answer." [[11]] The whole point of the test is to measure pure reasoning capability, stripped of environmental advantage. When OpenAI modifies the environment (retaining context between turns, compacting long conversations), it is not measuring the same thing ARC-AGI-3 was designed to measure. It is measuring a system that includes the model, plus a custom scaffold, plus memory persistence, plus prompt engineering — an opaque stack where the contribution of each layer is indistinguishable.
In the quiet, the protocol reveals its true intent
Now transpose this pattern onto blockchain infrastructure. In May 2025, the Bank for International Settlements published a working paper titled "Towards Verifiability of Total Value Locked in Decentralized Finance." The findings were damning. Out of 939 DeFi protocols studied, only 46.5% had TVL figures that could be independently verified using purely on-chain data. 10.5% of protocols relied partially or entirely on data extracted from external servers for their TVL computations. 68 non-standard calculation methods existed, each producing different numbers from the same underlying data. [[1]][[5]]
The BIS researchers introduced a new metric they called "verifiable Total Value Locked" (vTVL) — TVL computed using only on-chain data and standardized balance queries. The gap between published TVL and vTVL is a direct measure of opacity in the system. It represents the exact same phenomenon as the gap between GPT-6 Astra's 62.7% under the standard harness and its 98.6% under the custom one. Both numbers are real. Both are produced by real systems. But one is a measure of the underlying capability, and the other is a measure of the capability plus an engineered environment designed to maximize the output.
Authenticity is not minted, it is verified
The ARC Prize Foundation's own blog post on Astra is remarkably candid. It states that "saturating the benchmark would not represent proof of achieving AGI" and that Astra represents "meaningful progress towards generalization" — not a breakthrough, but a step. [[17]] OpenAI's President Greg Brockman acknowledged during the press briefing that AGI remains a "gray, fuzzy thing," even as the launch materials positioned Astra as the model that marked its arrival. [[16]] The tension between the headline ("98.6% on ARC-AGI-3") and the caveats ("under our custom harness, with our settings, as part of our system") mirrors the tension between DeFiLlama's published TVL and the actual on-chain balance that a standardized query would return.
In both cases, the gap is created by the same mechanism: the choice of measurement methodology is itself a form of optimization. When protocols self-report their TVL to aggregators using custom calculation methods, they are selecting the methodology that produces the most favorable number. When OpenAI selects a custom harness with context retention and compaction, they are selecting the environment that produces the highest score. Neither is fraudulent in the strict sense. Both are technically valid. But both violate the implicit contract of a benchmark — that it measures the thing being benchmarked, not the scaffolding around it.
We audit not to judge, but to understand
The parallel runs deeper. In the crypto context, the BIS study found that 240 "equal balance queries" were repeated across multiple protocols — the same underlying assets counted multiple times across different TVL figures. This double-counting inflates the aggregate TVL of the ecosystem without adding any new economic activity. Similarly, the ARC-AGI-3 benchmark, as originally designed, prohibits retaining context across actions. The ARC Prize Foundation explicitly stated that the benchmark "measures action efficiency, not just task completion" and that "a completion-only score would tell us that Astra completed an environment, but not how efficiently it learned to solve them." [[17]] By enabling context retention, OpenAI changed what the score measures — from action efficiency (how quickly the model learns) to task completion (whether it finishes the puzzle). The latter is easier, more marketable, and less informative.

This is not a technical detail. It is a structural choice about what kind of truth the industry wants to surface.
Layer two is a promise, not just a layer
The Layer2 ecosystem in crypto offers the clearest case study. In 2025, The Block's Layer2 outlook documented a stark bifurcation: most new L2s became ghost towns shortly after airdrop farming cycles, while only a small handful — Base, Arbitrum — sustained organic activity. Total Value Secured (TVS), a broader metric including bridged assets not actively deployed in DeFi, showed $40.5 billion across Ethereum L2s. But that aggregate number hides the same verification problem. Bridge flow data is publicly verifiable on-chain, yet most TVL reporting aggregates self-reported figures from protocol teams who have every incentive to maximize the number. [[3]][[6]]
The parallel with ARC-AGI-3 is precise. OpenAI reported 98.6% on ARC-AGI-3 using a custom harness. The ARC Prize Foundation reported 62.7% using the standard harness, and also noted that "we are not claiming that it is AGI." [[17]] Which number is "correct" depends entirely on what question you are asking. If you are asking "can OpenAI's system solve these puzzles?" the answer is yes, impressively. If you are asking "does this represent a fundamental advance in machine intelligence?" the answer is more qualified. The same ambiguity pervades every TVL metric in crypto. "How much value is locked in this protocol?" depends on whether you count bridged assets, staked assets, LP tokens, or any combination thereof. The number is a function of the methodology, not the reality.
Solitude clarifies the signal amidst the noise
The contrarian angle is uncomfortable. The crypto industry, which built its entire value proposition around verifiability — "don't trust, verify" — is now importing the same opacity from AI that it was supposed to eliminate. The same protocols that demand transparent benchmarks for blockchain performance are accepting opaque benchmarks for AI performance. The same investors who insist on audited smart contracts are accepting unaudited benchmark claims. The industry is becoming the thing it was built to replace.
Meanwhile, the AI industry is learning the same lessons crypto learned a decade ago: that metrics are gameable, that benchmarks create perverse incentives, and that the most impressive numbers are often the least informative. The ARC Prize Foundation's decision to publish their own analysis alongside OpenAI's launch — acknowledging the gap between standard and custom harness results — is the closest thing to a blockchain-style audit that AI has yet produced. It is a recognition that verification is not a one-time event but an ongoing process, and that the health of a field depends on its willingness to surface the gap between claims and evidence.
Every pixel carries a history we must respect
The GPT-6 Astra story is not about whether OpenAI achieved AGI. It is not about whether ARC-AGI-3 is the right benchmark. It is about a much older problem that both crypto and AI now share: the seduction of a single number that tells a clean story, and the difficulty of building systems that resist that seduction. The BIS researchers proposed vTVL as a standard for verifiable DeFi metrics. The ARC Prize Foundation published their own verification of Astra's performance under multiple harnesses. Both are attempting to solve the same fundamental challenge — how to measure capability in systems that are designed to optimize for the measurement itself.
There is no technical fix for this problem. You cannot code your way out of a misaligned incentive. The only solution is a cultural one: a commitment to publishing not just the winning number, but the methodology, the caveats, the alternative measurements, and the gap between them. It means treating verification as a core feature, not an afterthought. It means being willing to publish the 62.7% alongside the 98.6%, and explaining why both matter.