Open Weights, Closed Margins: Xiaomi's MiMo-V2.6 and the Repricing of On-Chain Compute
Hook: The Seven-Day Signal
Over the past seven days, three of the largest decentralized compute networks recorded their biggest single-week increase in committed GPU supply since the 2025 drawdown. No demand shock preceded it. No governance vote authorized it. No foundation grant funded it. A language model was published, and operators moved.
On September 22, 2026, Xiaomi released the weights for MiMo-V2.6 under an MIT license. The headline configuration: 1.02 trillion total parameters, 42 billion active, a mixture-of-experts router that stays frozen throughout training, a hybrid attention stack carrying a one-million-token context window, and five layers of multi-token-prediction speculative decoding. The accompanying benchmark card, vendor-reported and not independently verified, places the Pro variant at 46 on the Artificial Analysis Intelligence Index v4.3, level with Grok 4.7, and ahead of Claude Opus 5 and GPT-5.6 Sol on Terminal Bench, CyberGym, DeepSWE v1.1 and AutomationBench.
Equity desks filed it under China versus America. Crypto desks filed it under DeAI. Both filings are wrong in the same direction. They treated a public good as a product launch.
Here is the mechanical fact that matters, and it is the only one I care about this week: a trillion-parameter asset with an effective marginal replication cost of zero just entered a market that has spent four years building token structures around the premise that intelligence is scarce. The weights are free. Every layer downstream of them is not. The on-chain venues that priced scarcity into their emission schedules have not yet marked that to market, and the sideways tape is giving them cover to keep not marking it.

Context: The Liquidity Map Nobody Is Drawing
Start with the plumbing. Narrative is downstream of plumbing, always.
The US export-control timeline is now dense enough to read as a yield curve. H200 imports into China were prohibited in January 2026. The third-country cloud loophole, the path that let Chinese labs rent compliant capacity in Singapore, Dubai and Riyadh, closed in May 2026. By the second half of the year, frontier training inside China runs on domestic or non-US silicon, or on the stock of NVIDIA parts already inside the border. That is not a forecast. That is a calendar with dates on it.
The response shows up in share data. Domestic GPU and AI accelerator vendors went from approximately zero percent of the Chinese market to 41 percent. NVIDIA's China datacenter revenue went to zero. Alibaba's V900 program committed to a 500,000-chip cluster with production slated for the first quarter of 2027, which is a forward commitment and not an installed base, and the gap between those two things is where most of the bullishness currently lives.
Then the cost line. MiMo-V2.6's training burn was reported at roughly $432,000 per day. That is a consumption rate, not a total. If the run lasted sixty to ninety days, the aggregate lands between $26 million and $39 million, absorbed inside a hardware company's own cloud budget. Compare that to the decade-long capital intensity of a Western frontier lab and the asymmetry stops being about talent. It is about the price of the thing that talent consumes.
Now overlay the crypto liquidity map, because this is where the two worlds actually touch and where most commentary goes soft.
The 2024 spot ETF approvals moved roughly $10 billion of institutional money into a settlement wrapper. I wrote about that at the time as a plumbing change, not a protocol change. The issuance schedule did not move. The block time did not move. The monetary policy did not move. What moved was the accessibility of the rail, and the resulting depth of the order book behind it. On-chain liquidity depth in the top two assets is meaningfully different post-2024 than it was in the 2017 ICO era, and the difference is not sentiment. It is market structure: tighter spreads, deeper books, more passive flow, lower realized volatility per unit of notional. That distinction between a rail change and a protocol change is the whole game, and it is the same distinction that applies to open weights.
Compute-adjacent tokens, the decentralized training networks, the inference marketplaces, the GPU aggregators, the storage layers that position themselves as training-data rails, trade as a single factor with two betas. They carry a beta to the NVIDIA equity complex, which is the AI-demand leg. They carry a beta to BTC's liquidity cycle, which is the risk-asset leg. In a trending tape, the first beta dominates and the sector looks like a leveraged AI proxy. In a sideways tape, the first beta goes dormant and the second beta takes over, and the sector looks like what it actually is: a high-duration, low-cash-flow basket correlated to global risk appetite.
That regime flip is why the sector has been chopping between correlation states for two quarters. It is not confusion. It is two betas fighting over the same order book.
When Xiaomi dropped a frontier-class model with a permissive license, the market reflexively bought the first beta. It bought the AI-is-becoming-abundant trade, which is the intuitive read. More capable models, cheaper, freely available, therefore more AI infrastructure, therefore buy the compute tokens.
That reflex is a category error, and it takes roughly ten minutes of stack analysis to see why. I have spent twenty-six years watching this specific mistake get made, first in equity research, then in protocol economics. When the cost of a complementary good collapses, the value does not flow to whoever was selling the expensive version. It flows to whoever owns the bottleneck that remains.
Core: Weights Are a Public Good. Tokens Are Not.
The MIT license is the load-bearing detail. Not the parameter count. Not the benchmark card. The license.
MIT means any actor, a competitor, a state lab, a hobbyist, a defense contractor, a sanctioned entity, can download, modify, quantize, fine-tune and resell the weights. There is no usage telemetry. There is no regional restriction enforceable at the model layer. There is no revocation channel. Once the file is on enough mirrors, the artifact sits outside the jurisdiction of everyone, including its author. That is not a side effect of the license. That is the license.
Xiaomi's leadership is not naive about this. Luo Fuli's team came out of the DeepSeek lineage, and the DeepSeek playbook was explicit: you do not monetize the weights, you monetize the ecosystem the weights create. Give away the artifact, tax the flow. Money sits in inference, integration, hardware attach, and the behavioral data that returns.
So price the artifact. Zero. Then ask the uncomfortable question: what does the on-chain compute sector actually sell?
I have audited this category, and I mean audited, not opined. In 2021 I ran a team of three analysts through fifty major NFT collections to test their ownership and interoperability claims. Four percent had anything resembling a real protocol for cross-application state. The rest were database rows with a marketing layer and a rented server. We titled the report The Illusion of Digital Scarcity, and it was received about as well as you would expect from a market that did not want the arithmetic. NFTs are illusions whenever the scarcity is enforced by a single operator and rented from a platform that can change the rules.
The decentralized compute sector has inherited that structure wholesale, and the surface-level resemblance is not a coincidence. Both asset classes sell a claim on a resource whose supply curve is controlled by someone else.
Take the typical decentralized training network. Strip the token. What remains is a broker. It aggregates idle GPUs, schedules jobs, pays operators in a token whose value is a function of the network's own utilization, and takes a spread. The marginal product, floating-point operations per second delivered, is priced against a global spot market. The token premium sits on top of that spot price. The premium is the entire business model.
That premium is justified by one claim, repeated across every deck I have read in four years: that decentralized supply is cheaper or more available than centralized supply. Availability is sometimes true. Cheaper is almost never true once you load the token cost and the coordination overhead.
Now Xiaomi publishes a frontier model under MIT. Alibaba commits half a million domestic accelerators. The domestic accelerator share hits 41 percent. The effective price of frontier capability lands at fourteen cents per million input tokens on the Flash tier and forty-three and a half cents on Pro.
What just happened to the scarcity premium on distributed inference? It compressed. It did not disappear, because verifiability is still a real differentiator, but the general-purpose premium evaporated. Yields are traps when the underlying service is converging toward commodity pricing and the token emission schedule was underwritten by a margin assumption that no longer clears.
This is not a price prediction. It is a statement about what the tokens are claims on. If a token is a claim on a service whose marginal cost is falling toward the cost of electricity plus hardware depreciation, then the token is a leveraged bet on that decline being slower than the emission schedule. Most emission schedules run four years. Most relevant hardware depreciation curves run eighteen to twenty-four months to a fifty percent residual. Those two durations do not reconcile. One of them has to give, and it will not be the depreciation curve.
I watched this exact mismatch with my own capital in 2020. I put $25,000 into the Uniswap V2 ETH/USDC pool and spent the following months arguing with developers in Discord about impermanent loss and oracle manipulation, because I wanted to know whether the headline APY was a measurement or a marketing artifact. The lesson was never that impermanent loss is bad. The lesson was that a headline yield is a snapshot of a regime, and the regime is a function of incentive flow that is designed to stop. When the flow stops, the yield normalizes to the underlying risk, minus the mercenary liquidity that left first.
The same arithmetic governs compute tokens. The emissions are the yield. The underlying risk is a hardware residual value curve with no continuous mark.
There is a starker version of this. Tokenized GPU receipts, the newest entrant in the real-world-asset category, are structurally the same instrument. The cash flow is a rental rate. The depreciation is real and fast. The appraiser is the issuer. If you want to know what a tokenized depreciating hardware claim looks like when the market re-marks it, you do not need a model. You need a memory.
The Four-Layer Stack and Where the Margin Hides
Break the AI stack into four layers and the MiMo release becomes legible instead of confusing.
Layer one is silicon. Layer two is weights. Layer three is inference. Layer four is agents.
MiMo-V2.6 demonstrated what happens at layer two when a well-capitalized hardware company decides to compete: the price goes to zero and the artifact becomes a customer-acquisition cost. Xiaomi is not selling weights. Xiaomi is buying developers, and it is paying in a currency that costs the company almost nothing at the margin. The $432,000 per day training burn is a marketing line item in a business whose actual revenue is phones, cars, appliances and the operating system that binds them into a single account relationship.
Layer one is where current margin sits and where the export controls actually bite. That is why the 41 percent domestic share figure matters more than any benchmark score on the card. A model that runs only on constrained silicon forces the entire domestic supply chain, design, foundry, packaging, networking, systems integration, onto a demand curve that is politically guaranteed. Every dollar of NVIDIA China revenue that vanished is a dollar of addressable market transferred to a domestic vendor. That transfer already happened. It is not a forecast, and it does not depend on whether the agentic benchmarks replicate.
Layer three is the interesting one for anyone with capital at risk on crypto rails. Inference is a metered service with real, recurring, latency-sensitive demand. It has the properties that make a genuine market: variable load, differentiated quality of service, geographically constrained supply, and a clearing price. It is also the layer where the open-weights release is unambiguously bullish on volume and bearish on price. More capable free models mean more applications built, which means more tokens generated, which means more inference consumed. It also means every provider of that inference is now competing against a subsidized price point from a balance sheet that also sells hardware.
Look at the actual card. MiMo-V2.6 Pro at $0.435 per million input tokens and $0.87 per million output. Flash at $0.14 and $0.28. Those are not cost-recovery numbers. They are strategic numbers, and they set the anchor for every competing inference provider, centralized or decentralized. A decentralized inference marketplace that pays operators a token premium on top of the same underlying hardware cost cannot clear against that anchor unless it offers something the anchor cannot: verifiable execution, censorship resistance, or jurisdictional arbitrage for workloads that cannot legally run elsewhere.
The architectural details matter here, and they are routinely glossed over. A frozen router reduces training-time communication overhead, because MoE routing triggers all-to-all exchanges across the cluster, and that fabric traffic is the single hardest constraint in distributed training of sparse models. Freezing the router is an elegant optimization if you have the compute budget to pre-train the routing distribution. It is also a plausible signal of hardware constraints, because all-to-all across a slower interconnect is exactly the bottleneck you would engineer around under sanctions. Five layers of multi-token-prediction speculative decoding cut inference latency by drafting multiple tokens per forward pass and verifying them in one shot. That is mature engineering, not a paradigm shift, and it is precisely the kind of optimization that matters when your competitive advantage has to come from efficiency rather than raw FLOPs.
Layer four is where the applications live and where the second-order effects land. The agentic benchmark cluster is the tell: Terminal Bench 2.1 at 89.9 percent, CyberGym at 94.0, DeepSWE v1.1 at 71.9, AutomationBench at 53.1. All vendor-reported. All in the automation lane. If even the ordering is directionally right, the capability frontier has moved from conversation to execution. An agent that closes a software ticket, provisions infrastructure, or locates a vulnerability is an economic actor, not a chat interface.

Economic actors need payment rails. That is where the crypto argument stops being promotional and starts being structural.
Inference as the New Settlement Layer
Here is the connection both the AI press and the crypto press missed this week.
If agents become the dominant consumers of compute, the unit of account for machine labor stops being a subscription seat and starts being a metered flow. Metered machine-to-machine flows at scale do not clear well on card networks. Interchange is absurd at machine granularity, batch settlement is too slow, and the onboarding step has no equivalent for a process. They clear well on programmable rails with sub-cent granularity, deterministic fees, immediate finality, and no counterparty onboarding requirement.
That is a CBDC argument, a stablecoin argument, and a Layer 2 argument simultaneously. It is also why I care more about the settlement layer question than any benchmark score.
I have spent the last several years on central bank digital currency design, and the recurring blind spot in nearly every retail CBDC program I have reviewed is that the programs are built for the wrong customer. A CBDC optimized for person-to-person retail payments is solving a problem that instant payment rails already solved, and it is paying a large political price for the privilege. A settlement asset optimized for high-frequency, low-value, machine-originated payments is solving a problem that nothing has solved and that the current monetary architecture is not designed to absorb.
A trillion-parameter model released under MIT does not create that demand by itself. It removes a cost floor that was blocking it. Frontier capability at fourteen cents per million input tokens means a consumer agent can afford to run thousands of calls per session. At fifteen dollars per million, it cannot, and the application never gets built. The price point is the unlock. The rail is the bottleneck.
So watch the agent-payment layer, not the model layer. Watch whether the settlement asset ends up being a bank deposit token, a regulated stablecoin, a CBDC pilot with programmability constraints, or an unpermissioned asset. Each of those is a different political settlement, and the technical choice is downstream of the political one, no matter what the whitepaper claims about throughput.
There is a hard problem buried here that the decentralized sector keeps papering over with optimistic roadmaps. Verifiable inference is expensive. If you want cryptographic proof that a specific model produced a specific output, you pay a multiple of the compute cost of producing it. That multiple has been falling, but it remains large. So the market splits into two segments: unverified inference at commodity prices, and verified inference at a premium. The premium is a trust rent, and trust rents compress as reputation systems mature and as hardware attestation improves.
Which means decentralized inference cannot win on price against a hardware company that treats models as customer acquisition. It can only compete on the two things a centralized provider structurally cannot offer: verifiability and permissionlessness. Those are real. They are also narrow markets relative to general-purpose demand. The sector keeps building general-purpose capacity and pricing it as if the trust premium applies universally. That is a mismatch between what the network can do and what the token price assumes it charges for.
There is an instructive parallel in the DEX design space. Uniswap V4's hooks turn the exchange into programmable Lego, and the design surface explodes. The complexity spike that comes with it will scare off a large share of developers, and the same dynamic is about to hit programmable inference markets. Once you can attach arbitrary logic to a settlement event, the number of things you can build goes up and the number of people who can build them safely goes down. Distribution of capability is not the same as distribution of competence.
The Fragmentation Tax
There are dozens of Layer 2s now and the same small user base. That is not scaling. That is slicing already-scarce liquidity into fragments that each carry their own bridge risk, their own sequencer dependency, and their own liquidity mining budget.
Apply the same sentence to compute networks with one word changed. There are dozens of decentralized compute networks and the same finite pool of economically deployable GPUs. Every new network splits operator attention, splits developer tooling, splits whatever token liquidity finances the hardware, and inserts an accounting and bridging layer between a job and a machine. The fragmentation is marketed as decentralization. In practice it is routing overhead with a token attached.
Scale kills decentralization, and it does so at the exact moment the sector needs scale to be competitive.
A frontier inference provider requires multi-thousand-chip clusters with high-bandwidth interconnect, because a one-million-token context window does not fit in the memory of a single node. Serving that context requires tensor parallelism across machines, efficient KV-cache management, and a fabric that does not bottleneck on the attention pass. The capital cost of that topology is not compatible with a permissionless swarm of heterogeneous consumer GPUs, and no amount of clever scheduling changes the memory wall.
So the decentralized networks do what every undercapitalized network does. They serve the workloads that fit: short-context, latency-tolerant batch jobs, fine-tuning runs, rendering, synthetic data generation, and evaluation harnesses. Real businesses with real revenue. But not the frontier. And the frontier is where pricing power sits.
Alibaba's 500,000-chip commitment is the counter-example that proves the point. If that cluster materializes at spec and on schedule, it is a centralized asset with a centralized cost structure, and the open weights run beautifully on it. The MIT license does not require decentralization. It requires hardware. Xiaomi gave away the artifact and left the capital expenditure to whoever wants to serve it. That is a clever position: capture developer mindshare, externalize the capex, and let the domestic cloud complex carry the balance sheet risk.
A permissive license on frontier weights is not a decentralization strategy. It is a distribution strategy wearing a decentralization aesthetic. It lowers the barrier to building an application. It does not lower the barrier to serving one at scale.
Governance Without Legal Status
Now the part nobody wants on the record.
Open-weight ecosystems attract governance wrappers. Foundations, councils, stewards, DAOs with token-voting parameter committees. The pitch is community control over model evolution, safety policy and roadmap priorities. The structure is almost always the same: an unincorporated association of anonymous or pseudonymous token holders voting on proposals that commit resources nobody is legally obligated to provide.
I have spent a lot of time reading the fine print on these arrangements, and the pattern is consistent. Most DAOs have the legal status of no legal status. Not a corporation. Not a partnership by design, though the law may treat it as one anyway. Frequently a foundation in a friendly jurisdiction that does not actually control the thing the token holders control. When something goes wrong, a treasury is drained, a contributor is harmed, a model output causes damage, a regulator arrives, the people who voted have no liability shield and usually no insurance. The multisig signers have personal exposure they have never examined, and the foundation directors have duties they did not read.
This is not an abstract concern in the open-weights landscape. A model with offensive cyber capability, released under MIT, with no usage telemetry, cannot be governed by a token vote. There is no enforcement mechanism. A parameter committee can publish a safety policy. It cannot revoke a weight file. Security vendors can build detection. They cannot un-release a torrent.
CyberGym at 94.0 percent is the number that belongs on the front page and is not. That is a cyber-offense benchmark. A vendor-reported score at that level, in a model anyone can download, modify and run locally, describes a capability distribution shift. I am not making a moral argument. I am making an attack-surface argument. The cost of locating a software vulnerability, generating convincing phishing infrastructure at scale, and automating intrusion reconnaissance just moved down. That cost is a component of every organization's risk model.
Where does the sector price that? It does not. The security premium is a negative externality, and negative externalities do not appear in token price until a regulator forces an internalization. When that internalization arrives, most plausibly as a weight-export licensing regime, which is the logical next step after an import ban, the open-weight ecosystem's jurisdictional arbitrage closes. Mirroring remains technically possible. Institutional deployment becomes legally fraught. The developer graph forks into a regulated half and an unregulated half.
That fork is the single most important thing to position for, and it is in nobody's model.
Contrarian: The Decoupling Nobody Priced
Consensus is broken.
The consensus holds two beliefs simultaneously. First, that Chinese frontier open weights are bearish for US AI equity valuations. Second, that they are bullish for decentralized AI tokens. Both cannot be true in the way the market trades them, and I think both are wrong.
Take the first. A frontier-class open model does not reduce demand for compute. It increases it. Cheaper inference expands the set of economically viable applications, and the applications consume more total compute than the expensive regime did. This is not a controversial claim. It is the standard reading of every cost curve collapse in the last century of industrial history, and it held for steel, for electricity, for bandwidth, and for storage. What the model does is compress the margin of the layer that sells intelligence as a service while expanding the volume of the layers that consume it. That is bad for one business model and good for silicon, networking, memory, and power.
The market traded the first half and ignored the second. It marked down the model layer and did not mark up the physical layer proportionally. That is a larger mispricing than anything in the token complex, and it sits outside crypto entirely.
Now the second belief. The DeAI trade is positioned as a beneficiary of open weights. The reasoning is that free models commoditize intelligence, and decentralized networks are the natural low-cost provider. But the anchor Xiaomi set is a subsidized price from a balance sheet that also sells phones and cars. A decentralized network paying retail hardware costs plus a token premium cannot undercut it. It can only differentiate on verifiability, which is a narrow market.
Meanwhile, the open weights reduce the value of the one asset the decentralized networks were implicitly shorting: the closed model API. If you can run frontier capability locally for the cost of electricity and a one-time hardware purchase, a large class of applications simply leaves the metered market. That demand does not flow into decentralized inference. It flows onto hardware the user already owns, or onto the cheapest compliant cloud.
The real decoupling is not China versus America. It is between the price of intelligence and the price of the substrate that produces it. Weights trend toward zero. Silicon, power, memory bandwidth and cooling trend toward the cost of capital. Value is migrating from the top of the stack to the bottom, and every token designed to capture value at the model layer is now swimming against that current.

I want to be precise about the second-order consequence, because it has a crypto shape and it is under-owned.
If weights are free and inference is metered, the scarce resource in an agent economy is not capability. It is settlement. An autonomous agent transacting with other agents needs a rail with programmable constraints, deterministic fees, and enforceable finality. That is a payments problem sitting inside a monetary system whose dominant instruments were designed for human-initiated transactions at human frequency. Banks clear in batches. Card rails carry interchange that is absurd at machine granularity. Stablecoins solve granularity and inherit issuer risk. CBDCs solve granularity and issuer risk and inherit the political constraint.
For three years I have watched central banks design retail CBDCs for a use case that instant payment systems already cover, while the machine-payment use case sits unaddressed. That is where the design gap is largest, and that is where demand is about to be created by a price point rather than a policy paper. A frontier model at fourteen cents per million tokens turns the machine-payment rail from a concept into a bottleneck.
Own the rail, not the commodity that flows over it.
There is a fair counterargument. Open weights plus domestic silicon plus a permissive license could produce a genuinely independent AI stack serving the large bloc of countries that cannot or will not buy into a US-controlled stack. That is a real market, spanning Southeast Asia, the Gulf, parts of Africa and Latin America, and it is denominated in infrastructure contracts, not tokens. Sovereign clouds, domestic fine-tunes, local-language capability. The commercial logic is sound and the demand is real.
But notice what that business requires. Contracts, service-level agreements, export compliance, legal entities, government relationships. None of which a token-voting parameter committee can provide. The market that benefits most from open weights is the market that least resembles a decentralized organization.
Takeaway: Positioning Inside the Chop
A sideways tape is not a waiting room. It is where the cost basis gets set for the next regime, and it is the only environment where you can accumulate a position without competing against momentum money for it.
Right now the market is pricing two things wrong at once. It is under-pricing the physical substrate that benefits from cheap intelligence. It is over-pricing the token structures built to sell intelligence as though it were scarce.
Three signals will settle the argument. I am watching all three rather than trading the narrative.
Independent replication comes first. Every performance number in the MiMo release is vendor-reported. Level with Grok 4.7, ahead of Claude Opus 5 on agentic suites, none of it measured by a neutral party. If blind evaluations from an independent harness hold up, the pricing pressure on closed frontier APIs is durable and the re-rating of the physical layer is justified. If the agentic scores collapse under controlled conditions, the entire event was a benchmark artifact and the sector gives it back inside a week. Track the download volume and the daily-active inference series as the honest proxy, because those are harder to curate than a scorecard.
Second, the weight-export regime. The import side is already closed. The logical next step is a licensing requirement on the export of model weights and the compute used to train them, with a compliance burden that lands on the cloud providers hosting the mirrors. That legislation forks the developer graph, and positioning means knowing which rails and hosting layers have a compliant story and which do not.
Third, the agent settlement layer. Watch which programmable asset actually clears machine-originated transactions at scale over the next eighteen months: a regulated stablecoin, a deposit token, a constrained CBDC pilot, or something unpermissioned. That choice determines which chains matter in the agent economy far more than any throughput benchmark.
One structural question to hold onto through the chop. If a trillion-parameter frontier model costs fourteen cents per million tokens to query, the weights are free, and the hardware is increasingly domestic to the jurisdiction that built it, then what exactly is the scarce asset that a compute token represents?
I have a position. I do not think most of the sector has one.