On a Tuesday morning I opened a staged analysis report sitting in my queue. Nine sections. Technical architecture. Token economics. Market structure. Ecosystem position. Regulatory exposure. Team and governance. Risk matrix. Narrative and expectation gap. Supply-chain transmission.
Each section had a table. Each table had rows. Every cell in every row carried the same value: N/A — insufficient input. The dimension notes did not say "unknown." They said, per section, exactly which inputs were missing and what each one would have changed. At the bottom, one line: REPORT STATUS: BLOCKED — INPUT DATA MISSING.
Nothing was analyzed. That was the point. And it is the rarest artifact I have seen in crypto research this cycle.
I have read a few thousand "deep analysis" reports across fourteen years of watching this industry. The overwhelming majority arrived with a thesis, a scaffold, and confident prose in every cell. Almost none of them could tell you what they would have done if the input had been empty. Most would not have noticed the input was empty.
Code does not lie, but it often omits the context. The report that refused to run was the only honest output in my inbox that week, and it produced zero words of analysis.
To understand why that matters, you have to look at how the pipeline was actually wired. Stage one was a deconstruction layer: extract title, source, information points, core claims, domain tags, identified protocols. Stage one returned nulls. Empty strings where titles belong. An empty array where the information-point list belongs. Blank fields for identified projects.
Stage two was constrained by a single hard rule — every judgment must cite a stage-one information point. Stage one produced zero citable points. Stage two evaluated the constraint, found no valid path to any of its nine dimensions, emitted the template with explicit insufficiency labels, and halted.
That is a circuit breaker. It is also, statistically, an extinct design pattern.
Here is the economics. Generating three thousand words of plausible crypto analysis costs a fraction of a cent and roughly eleven seconds. Acting on fabricated analysis costs a fund somewhere between a rounding error and a fund-ending drawdown, depending on position size and how early the fabrication is discovered. The marginal cost of a lie is effectively zero. The marginal cost of a truth is a refusal to produce — which, in any organization that measures output volume, reads as underperformance.
The asymmetry is the entire story. Every incentive in the research layer points toward generation and away from verification, and the only thing that ever stopped it was a programmer deciding to put a require() in the middle of the pipeline.
I have seen this exact failure before, in a domain where the stakes were measurable on-chain. In 2020 I spent three weeks reverse-engineering the price feed mechanisms of five lending protocols during the DeFi Summer. I was not looking for exploits. I was looking for how each feed reported its own staleness. Four of the five returned a price. That price was sometimes old. The calling contract had no mechanism to distinguish a fresh feed from a stale one, because both arrived as a uint256 with the same type signature and no accompanying timestamp in the consumption path.
In August 2020, delayed feeds produced undercollateralized positions that liquidators could not reach fast enough. The data was not wrong. The absence of data had been formatted as data. That is the bug class. It is not specific to Solidity, and it is not specific to DeFi. It is the most durable defect in this industry, and it is currently being deployed at industrial scale in the research layer that feeds allocators.
There are three architectural choices available when input is empty. Almost every pipeline picks the worst one by default.

The first option is hallucination. Fill the schema with values that satisfy the type signature. In Solidity terms, this is return 0; inside a failed try/catch. Zero is a valid price. Zero is a valid balance. Zero is a valid TVL. Downstream, nothing can distinguish "we do not know" from "the value is zero," because both are the same twenty-six bytes.
In a language model pipeline, the equivalent is filling a technical_architecture field with correct-sounding but unverifiable claims. The schema validates. The report ships. The reader has no way to know that the field was populated by a fallback branch rather than a data path.
This is not a metaphor. It is the same bug. In 2017, while I was still a final-year data science student in Ho Chi Minh City, I spent four weeks manually auditing the Solidity of three lesser-known ICO contracts. I found critical reentrancy vulnerabilities in two of them and submitted pull requests. What made those bugs exploitable was not faulty arithmetic. It was that a state variable had been read at a moment when the protocol assumed it was still valid. Stale reads formatted as current state. The 2020 oracle failures were the same pattern with more zeros attached.
The second option is a hard revert. The pipeline halts and produces nothing. Clean, type-safe, and politically unsustainable. An inbox item that generates no artifact gets deprioritized. Then someone removes the constraint because it costs throughput, and the pipeline starts hallucinating again — only now with the memory of having been "fixed." I watched this happen to a bridge team's internal review process in 2022. The gating check was removed in a sprint planning meeting. I do not think anyone in that meeting understood they were deleting a trust boundary.
The third option is typed absence: a return path that carries value | unknown | stale as distinct, mutually non-coercible states, where unknown cannot silently decay into a number. Good oracle adapters have done this since roughly 2022. latestRoundData() returns an updatedAt timestamp, and the consumer is expected to compare it against a heartbeat threshold. Read the Chainlink documentation on this and it is unambiguous: the staleness check is the consumer's responsibility. In practice, a majority of consumers do not perform it, because the function still returns a number and the number still compiles.
A typed absence at the presentation layer is the whole difference between an analysis and formatted noise. The report I opened implemented it. Every field was explicitly non-informative, and — this is the part that matters — each dimension listed the specific missing inputs required to proceed, with the corresponding trigger condition. Not "we lack data." Rather: we lack data of these five classes, and here is what each class would change about the conclusion.
That specificity is expensive. It requires the pipeline to model its own dependencies. Most systems skip it, and the cost of skipping it is invisible until someone reads the output as fact.

I want to be precise about the risk profile, because "fabricated research" sounds abstract until you map the blast radius.
| Failure mode | Mechanism | Detection difficulty | Blast radius | |---|---|---|---| | Fabricated technical analysis | Plausible scaffolding, no falsifiable claim | Very high — reads better than accurate work | Sector-wide capital allocation | | Fabricated on-chain metrics | Numbers that cannot be negatively signed | Medium — spot-checkable against a node | Position sizing, LP entry timing | | Fabricated team or backer claims | Attribution without disclosure trail | Low to medium | Retail entry near local tops | | Fabricated regulatory assessment | Jurisdiction-general language | Very high — no single reader can adjudicate | Institutional adoption theses |
The fourth row is the one that keeps me up. A regulatory claim is the hardest class of statement for any individual reader to falsify, because falsification requires legal expertise in a jurisdiction the reader probably does not reside in. It is the only category where verification cost scales with the reader rather than the writer. Which means it is the category most likely to be fabricated at scale, and the category most likely to be acted upon by large pools of capital.
Here is where my own specialty is relevant, and where I think the framing most people use is wrong.
In zero-knowledge circuit design, a proof is a statement about a witness. The prover holds the witness; the verifier does not. What the verifier learns is a boolean: the statement holds. If the prover does not have the witness, there is no proof to produce. Not a proof of emptiness. Not a proof of unknown. There is no construction in the field that proves the absence of a witness is informative. Zero-knowledge proofs let you demonstrate what you know without revealing it. They have nothing whatsoever to say about what you do not know.
This matters because it means a research pipeline can never produce a proof of insufficiency. It can only produce an assertion of it. Assertions are cheap. Copies are free. Within a week of a "we could not verify X" report, derivative reports will cite it as "X was verified as absent" — a semantic drift that becomes irreversible the moment the artifact exists as text. The honest report is honest for exactly as long as nobody re-summarizes it.
I learned a version of this in 2022, when I spent two months triaging the source code of legacy Ethereum Layer 2 bridges during the depths of the winter. I found three critical security flaws in a bridge that was, at the time, popular. I was dismissed — not on the merits, which were never disputed in writing, but on the question of whether I had standing to report. The team did not evaluate the finding. They discarded the channel. That is revert() applied to a person, and it is the same design flaw as a silently dropped error: the output never reaches a consumer who could act on it.
I published pseudonymously instead. It gained traction among security researchers. Competence, not identity, is what cryptography rewards — but only if the finding survives the routing layer.
Now the part I think the report itself got right by accident, and wrong by design.

Stage two behaved correctly. But stage two's constraint only existed at stage two. Upstream, stage one returned a well-formed empty schema. It did not raise. It did not error. It returned nulls where strings belong and empty arrays where lists belong, and the pipeline typed them as valid without comment. If stage one had thrown — "zero information points extracted from provided content" — a human would have looked at the input within the hour.
Instead, the failure was formatted as a successful parse. The only thing standing between that and a fabricated nine-dimension report was a second-stage constraint that a future maintainer will remove the first time it costs them throughput. That is the actual vulnerability: the contract between stage one and stage two had no require(). The honest report is not a safety property. It is a coincidence of two layers that happened to disagree in a survivable direction.
And there is a reflexive trap worth naming. Refusal is now content. "I will not analyze this" is a shareable artifact, and the cheapest way to signal rigor is to produce nothing about something. Expect a wave of performative blocking — pipelines that decline to analyze real inputs because declining reads as integrity. The signature that separates practice from theater is not the refusal. It is the specificity of the missing-input list. Theater says "insufficient data." Practice says "insufficient data of these nine classes, and here is what each would change."
Over the last year I designed a privacy-preserving compliance layer for an institutional DeFi platform — solvency verification without transaction-history exposure. The hard part was never the circuit. The hard part was specifying every edge case so that an absence of evidence could not be read as evidence of compliance. That is the same problem, in a different register.
Within eighteen months I expect research pipelines that commit their input corpus to a hash before generation and publish the commitment alongside the output, so a reader can verify the artifact was derived from the corpus it claims. Not a proof — a binding. It will be adopted slowly, because it makes refusal auditable, and most producers do not want refusal to be auditable.
The exploit surface that matters in the next cycle is not the contracts. Contracts have years of adversarial pressure behind them and armies of auditors in front of them. The research layer that feeds allocators has neither. Its dominant bug class is already deployed across the industry, and it does not even look like a bug.
It looks like a report.