The ledger was clean, but the vision was fragile.
That sentence surfaced in my mind the moment I read Star Xu's declaration that AI now processes 95% of OKX's engineering pull requests. A single number. No methodology. No third-party audit. No definition of what "handled" actually means. In an industry that has spent the past decade learning to distrust unverified claims, the founder of one of the world's largest centralized exchanges handed the market a statistic that would be laughed out of any serious engineering review — yet it was swallowed whole within hours.

In my years auditing smart contracts, I learned that the most dangerous numbers are the ones that sound precise.
Let me be clear about what this number cannot possibly mean. A pull request is not a ticket. It is not a commit. It is a formal request to merge a set of code changes into a production repository — a repository that, in OKX's case, directly manages custody of user funds, clearing logic, and order-book settlement. "AI handles 95% of PRs" could describe three radically different realities: AI generating the code inside the PR, AI reviewing and commenting on the PR, or AI autonomously approving and merging the PR. These are not stylistic distinctions. They are the difference between a copilot and an autopilot, and between a productivity tool and an existential threat to a safety-critical system.
The context matters more than the claim. OKX is a top-three global CEX by derivatives volume. Its codebase is not a weekend DeFi experiment; it is the custodial backbone for billions in user assets. When the 2020 DeFi Summer taught me anything, it was that even audited, battle-tested contracts — Aave's, for instance, which my team arbitraged profitably for three months — carry latent fragility that only reveals itself under stress. Aave had formal verification and multiple audits. OKX is now publicly signaling that a large share of its code changes pass through an LLM pipeline with no disclosed verification standard.
Industry baselines tell a different story than the 95% figure. GitHub's own 2023 research measured AI-generated code at roughly 30–40% of new code among adopters. Google reported around 25–30% of new code. Even the most aggressive enterprise deployments rarely exceed 50% when measured honestly across all engineering work. A 95% claim is a statistical outlier so extreme that it either reflects a narrow definition — boilerplate, test scaffolding, internal tooling — or it reflects a narrative, not a measurement.
Here is where the psychological cost accounting begins. When a founder states a number like this, he is not issuing a technical report. He is issuing a recruitment signal, an investor signal, and an internal mobilization directive all at once. The number is designed to be memorable, quotable, and un-verifiable. It is a narrative hook. And in the current AI-saturated bull market — where every whitepaper has an LLM section and every exchange is racing to brand itself as an AI company — that hook is worth more than an audit.
But exchanges are not SaaS startups. They are closer to banks. If AI is truly generating the majority of code changes in a fund-custody system, then the only remaining defense is human code review — and human review at scale is exactly the bottleneck that AI efficiency promises to remove. The efficiency gain and the safety margin are, in this specific case, in direct tension. One of them has to give.
The regulation angle is equally unexamined. Traditional financial institutions operating under SOX-style internal controls must maintain traceability for every code change that touches financial reporting or custody. AI-generated code with no audit trail creates a accountability void: if something breaks, there is no author to blame, no reasoning to reconstruct, no change history to interrogate. I have seen what happens when unverified logic slips into production — in 2018, I spent six months manually auditing Power Ledger's token sale contracts and found a reentrancy vulnerability in the distribution mechanism. The team ignored it for speed. The bug was exploited in testnet. The lesson was not that code is dangerous. The lesson was that code is dangerous when its provenance is ambiguous and its review is rushed.
Code does not lie, but people certainly do. And the most polished lies are the ones told with precise numbers.
So what do we actually know? Star Xu said something. The market amplified it. No independent engineer has verified the pipeline. No third-party audit firm has reviewed the AI merge process. No GitHub metrics, no commit logs, no review-to-merge ratios have been disclosed. We are asked to accept a claim about safety-critical infrastructure on the authority of the person with the most incentive to make it.
This is not to say OKX is reckless. It may well have a robust human-in-the-loop protocol. But the burden of proof lies with the claimant, and a single sentence is not proof. The silence around the methodology is the loudest signal in this entire episode.
The broader pattern deserves attention. AI coding adoption is real and accelerating. The question is no longer whether exchanges will use LLMs, but whether they will use them with disclosed guardrails — independent verification of AI-generated code in custody-critical modules, mandatory human sign-off on any PR touching fund flows, and auditable change provenance. None of that was mentioned. What was mentioned was a percentage.
We bet on the pattern, not the hype. The pattern here is familiar: a founder produces an unfalsifiable efficiency metric, the narrative machine converts it into a trend story, and the underlying safety question never gets asked. We saw it with TPS claims. We saw it with zero-knowledge proving costs that operators quietly subsidize until the subsidy becomes unbearable. We are seeing it now with AI percentages that no one can audit.
The edge, if there is one, is not in believing the number. It is in watching what happens when the first security incident traces back to a merged pull request that no human ever meaningfully reviewed.
Until then, keep your risk parameters tight. Track the disclosures. And remember that in a system designed to move money, the only efficiency that matters is the one that survives an audit.

The summer was loud. The profits were quiet. The code, as always, is still waiting to be read.