The Jagged Frontier: Frontier AI Still Can't Do Epoch's Job

HasuPanda • • In-depth
The Jagged Frontier: Frontier AI Still Can't Do Epoch's Job Here is a fact worth pinning above every founder's desk this quarter: the most capable models on the planet — systems that draft your Solidity, summarize your governance threads, and clear your take-home interviews — cannot perform the daily work of the research institute that just tested them. Epoch AI, a nonprofit that studies the trajectory of machine intelligence, ran its own workflows through frontier systems and published the result under a deliberately flat title: Automation Reports. The verdict is not subtle. Structured tasks: handled. Open-ended work: not. That gap is the whole story. Not the benchmark. Not the model list. The gap. I have spent years watching this pattern from the other side of the table — as an auditor who learned, painfully, that a clean-looking contract and a safe contract are different objects. The instinct that made me dissect ERC-20 standards during the 2017 ICO mania is the same instinct that now makes me distrust the framing of this report. Not because the claim is false. Because as reported, it is almost impossible to falsify. Context: a study that grades itself Epoch AI is not a household name outside research circles, but inside them it carries weight. The organization builds datasets on compute, model scaling, and capability trends — the kind of infrastructure that policy researchers cite when they need a number that is not a marketing number. Its Automation Reports are positioned as public research output, not a product. There is no API to buy, no SaaS tier, no token. That matters, because it removes the most obvious commercial motive while leaving a subtler one untouched. The design is what makes the study interesting and, simultaneously, hard to trust. Rather than testing models on a synthetic benchmark, Epoch used its own research workflow as the task pool: data collection, model specification verification, trend analysis, forecasting. This is a real-world workflow benchmark, and methodologically it sits a step above the polished leaderboards that dominate AI discourse. Synthetic benchmarks measure what is easy to measure. Workflow benchmarks measure what people actually do. That is why the frontier's failure here is more interesting than any leaderboard. But here is the tension baked into the design. The task pool was assembled by the same institution that would be judged against it. When a research group asks whether AI can do its job, the answer is filtered through decisions about which tasks count, how success is scored, and where the line between structured and open-ended is drawn. Those decisions are invisible in the summary. Reading the silence between the blocks, what remains is a conclusion without a method — a verdict without a transcript. The history of AI evaluation is a history of Goodhart's law. Build a benchmark, and the industry optimizes for the benchmark rather than the skill it was meant to proxy. MMLU scores inflated. Coding leaderboards saturated. Each time, the frontier looked solved until it met a task nobody had thought to grade. Workflow benchmarks were supposed to break that cycle by refusing to abstract the work. Epoch's choice to grade itself is an extension of that logic — and, unavoidably, an invitation to the same failure. This is not a knock on Epoch specifically. It is a structural property of self-evaluation. The same dynamic shows up every time a protocol audits itself, or a fund marks its own book. The number is not necessarily wrong. It is just not independent. Core: the jagged frontier is real, the evidence is thin The directional finding is credible. It matches a consensus that has been building quietly across the evaluation community for two years: the capability frontier is jagged. Models are superhuman in narrow lanes and surprisingly brittle two steps off the path. Ask a system to reformat a spreadsheet into JSON and it never blinks. Ask it to decide which of two research questions is worth pursuing and the confidence evaporates. The pattern is not random. It tracks the presence or absence of a verifiable reward. The mechanism behind that jaggedness is well understood by anyone who has built with these systems. Structured tasks have a verifiable shape. There is a ground truth, a scoring function, a reward signal that reinforcement learning can chase. Open-ended work has none of that. There is no single correct answer, no automatic grader, no clean gradient to descend. The task is defined by judgment, and judgment resists optimization. The alignment angle is where this stops being an office-productivity story. Open-ended work is the hardest thing to align, because alignment itself assumes you can specify what you want. RLHF and its descendants need preferences to train against. On a structured task, preferences are cheap — the answer is right or wrong. On an open-ended task, "good" is contested even among experts. A model that cannot be graded cannot be reliably improved by gradient. The bottleneck Epoch reports is not a scaling problem. It is a specification problem, and specification is the same wall the alignment community has been staring at for years. Tracing the logic gates behind the yield here is instructive. In DeFi, the returns that look real are the ones backed by revenue; the rest are emission schedules dressed as income. In AI evaluation, the capability that looks real is the capability that can be scored; the rest is narrative dressed as progress. Epoch's finding, if it holds, is a rare moment where the two currencies diverge — a system that scores well on everything measurable still fails at the thing that cannot be measured. That distinction matters for where the capital flows. If the bottleneck were compute, the answer would be more GPUs, and the hyperscalers would win by default. If the bottleneck is architecture or methodology, then the marginal dollar is mispriced. The article offers no evidence either way — but the direction of its finding, toward a qualitative wall rather than a quantitative one, is a quiet challenge to the "scale is all you need" consensus that has underwritten a decade of data-center spending. What the report does not give us is the part that would let me trust it. There is no task list. No scoring rubric. No model roster — not GPT-4o, not the model, not Gemini, not Llama, none of them named. There is no human baseline. And that last omission is the most damaging. "AI struggles with open-ended work" is a sentence with no anchor without a number for how humans perform on the same tasks. Struggles relative to what? Perfect accuracy? A junior analyst? A senior one? The report says "struggles" and leaves the reference frame on the floor. The missing baseline is not a minor omission; it is the hinge of the entire claim. Suppose Epoch's researchers score 60% on their own open-ended tasks. If frontier models score 55%, the story is "AI is closing in." If they score 15%, the story is "AI is nowhere near." The same headline — "AI still can't do Epoch's job" — describes both. Without the numbers, the reader supplies the meaning, and the reader's priors decide the takeaway. This is not a small gap. It is the difference between a finding and a vibe. And I have seen exactly this failure mode before. During DeFi Summer, I stress-tested Sushiswap's fork against Compound's mechanics and found that the "infinite yield" everyone celebrated was an emission rate with no revenue behind it. The math was real. The story was not. Here, the story — "frontier AI still can't do Epoch's job" — is real and quotable. The math is missing. There is a second layer that the AI press will miss and the crypto press is uniquely positioned to catch. The study's most loaded implication is not about office work. It is about recursive self-improvement — the idea that AI could accelerate AI research itself, a loop that safety researchers treat as the central risk of the coming decade. If frontier models cannot even do Epoch's routine research tasks, then the timeline for a self-improving loop slides backward. The most feared scenario gets a reprieve. Whether that reprieve is real depends on data we do not have. Contrarian: an institution grading whether AI can replace institutions Now the uncomfortable part. Epoch studies AI capability and AI risk. Its institutional identity is bound to the proposition that these questions matter enormously. When such an organization evaluates whether AI can do the work of researchers, the evaluation is not neutral — not because anyone is lying, but because task selection and scoring are interpretive acts, and interpreters have priors. The bias could cut either way. A researcher might underrate AI, because conceding that a machine can do your job is existentially uncomfortable. Or overrate the difficulty, because warning about risk is the institutional mission. The direction is unknown, which is precisely the problem. An unfalsifiable conclusion from an interested party is not evidence. It is a hypothesis wearing a lab coat. The counter-argument is fair: independent verification exists. METR, ARC, and academic groups all publish capability evaluations, and their results can be cross-checked against Epoch's. If three unrelated teams find the same open-ended bottleneck, the bias concern evaporates. That is the correct test. But it is a test nobody has run yet in public, and until they do, the report stands on institutional trust alone. And it arrives at the exact moment when the "AI x Crypto" narrative needs raw material. Every cycle, a new meta absorbs the vocabulary of whatever is ascendant — DeFi borrowed "protocol," NFTs borrowed "culture," and now agent tokens borrow "autonomy." A finding like this one is perfect feedstock: it lets a project claim it is building the human-in-the-loop layer that the frontier cannot replace. Where code meets cultural memory, the borrowing happens before the evidence lands. Watch for that framing within a quarter. The study did not endorse it. The market will anyway. The healthier reading is smaller and more useful. This is a snapshot of one model generation, taken by one team, on one workflow. Frontier capability moves on a quarterly cadence. The same report, rerun in twelve months, could invert. Treat it as a temperature reading, not a verdict. The "still" in "still can't" is doing enormous rhetorical work — it implies a stable boundary, when the only thing stable about this frontier is how fast it moves. Takeaway: the benchmark is the asset, not the headline If Epoch publishes the full Automation Reports — task list, rubric, model roster, raw scores — the conclusion becomes testable, and the real value surfaces. A reproducible, workflow-level capability benchmark is scarce. METR has one. ARC has one. SWE-bench has one. A research-workflow benchmark would be a fourth pillar, and it would outlast whatever the current models happen to score. So the question worth tracking is not whether AI can do Epoch's job today. It is whether anyone will publish the transcript that lets us check. The audit trail never lies — provided someone is willing to keep it. The next model generation is already in training. The only question is whether the next report arrives with a method attached, or another headline asking us to trust the institution and skip the math.

The Jagged Frontier: Frontier AI Still Can't Do Epoch's Job