Sidequest Note · Parv Mahajan · · updated
The US-China AI gap is very uncertain.
For every day, we take the best Chinese model available and ask how many months earlier a comparable US model as measured by the the Epoch Capabilities Index was already out. To capture uncertainty in the ECI measurements, we run a Monte Carlo simulation: we redraw every model’s score from its CI interval 800 times and build a graph for each world to get the 5th-95th percentile shaded area shown.
Sidequest Notes solely represent the views of the authors, and do not necessarily reflect the views of MCNAIR as a whole.
Methodology and Adversarial Review from Claude Leviathan
How the chart above is computed, followed by an adversarial review of that computation. Commissioned by MCNAIR and published without edits.
How the simulation works
The measurement. For every day, we take the highest-scoring Chinese model released on or before that day, then find the earliest US release that scored within one ECI point of it. The gap is the number of months between that US release and the day being read. The matched US model only changes when China’s best model improves, so the line climbs one month per month between Chinese releases and drops when one lands — the sawtooth. The one-point slack is deliberate: requiring a strictly stronger US model at every point biases the gap downward, because the US side of every comparison is then always the better model.
Why a band at all. Epoch publishes each ECI score with a confidence interval, and those intervals are wide — a median of roughly five points on each side, and often badly lopsided (Qwen3-Max runs −3.7 / +10.8 around its estimate). A single line drawn through the point estimates asserts a precision the index does not have.
One draw is one world. We redraw a score for every one of the 181 scored US and Chinese models from its own published interval, using a split normal so each side of an interval keeps its own width instead of being averaged into a symmetric error bar. We then rebuild which model was actually best in each country at each point in time under those drawn scores — which can promote a model that never set a record on the point estimates — and re-run the identical rule the center line uses. That is one hypothetical world. We run 800 of them from a fixed seed, so the chart is the same on every reload.
Reading the band. Between any two Chinese releases, every world holds its US anchor fixed, so every world rises at exactly one month per month and the ordering of worlds never changes. That lets us sort the 800 gaps once per interval and draw the 5th and 95th percentiles as straight ramps, exactly like the center line. The shading is twelve nested polygons interpolated from the center out to the edges, so it is densest where the worlds agree and nearly gone at the extremes. The Simulations toggle replaces the envelope with the individual worlds.
What is left out. LLaMA-65B and Baichuan1-7B, each country’s first frontier entry, are dropped per Epoch’s methodology. A world whose US anchor lands at or before GPT-4 — the oldest US model in the dataset — is a lower bound rather than a measurement, so it is discarded, and an interval must resolve in at least 90% of worlds before it is plotted. That is why the chart starts in December 2024 rather than at the first Chinese release.
Adversarial review
Two things hold up. The dataset reproduces exactly from Epoch’s published CSV — 132 US models, 49 Chinese, 17 and 19 on the respective frontiers, every score identical. And the median of the simulation equals the center line at every plotted point, which is the check that the band is answering the same question the line asks. The rest of this is what is wrong with it.
- The band is one-sided across two-thirds of the chart. Over the 578 plotted days, the 95th percentile coincides exactly with the center line on 32% of them and the 5th percentile does so on another 34%; only a third of the chart has spread on both sides. Between 22 December 2025 and 2 February 2026 the center runs 12.1 to 13.4 months and the upper edge is identical to it, while the lower edge sits at 1.1 to 2.5 — the line rides the top of an eleven-month band. The cause is not cosmetic: 68% of the 800 worlds in that interval pick the same US anchor, and not one picks an earlier one, because no US model between GPT-4 and o3 clears the bar under any draw. The upper tail is truncated by how few models have been scored, not by evidence. “5th to 95th percentile” describes the arithmetic honestly and the picture misleadingly: this is a distribution with most of its mass on a single point, not a spread.
- The band dips below zero using models that did not exist yet. The lower edge reaches −2.4 months in September 2025, which puts a −3 gridline under an axis captioned “Months the US was ahead.” Neither the center rule nor the simulation restricts the US search to models available on the day being plotted, so a world in which the best Chinese model is drawn high resolves to a US anchor in the future. On 5 September 2025 the 5th percentile is set by Gemini 3 Pro, released 17 November. The chart does not say so, and nothing on it lets a reader recover that.
- The width is mostly an assumption. All 181 models are drawn independently. But ECI is a composite index, so a large share of its uncertainty — benchmark selection, calibration, scaling choices — is shared across models rather than independent per model. Re-running the identical simulation with one common draw per world instead of 181 independent ones halves the band: median width falls from 4.1 months to 2.0. What is published is closer to an upper bound on the uncertainty than a best estimate of it, and the assumption that produces it is stated nowhere. A related worry — that taking a running maximum over many noisy scores inflates both frontiers — turns out to be small, at +0.1 to +0.4 ECI on the US side and −0.1 to +1.3 on the Chinese side.
- Discarded worlds bias the upper edge downward. A world anchored at or before GPT-4 is thrown away. Those are not missing observations; they are the worlds in which the gap is at least as large as our coverage can express, which makes them the largest values in the sample. Removing them and then reading a 95th percentile off what remains understates the top of the band. The interval beginning 24 December 2024 resolves in 94.9% of worlds, so its upper edge is really the 90th percentile; the 90% threshold permits this to reach the 85th. Sorting discarded worlds to the top instead of dropping them widens the median band from 4.1 to 4.7 months.
- The band moves on days the line does not. The center is recomputed at the 19 Chinese frontier releases; the band is recomputed at all 49 scored Chinese releases, because a resampled world can promote a model that never set a record. That produces three visible discontinuities where an edge jumps while the center is flat, the sharpest on 5 September 2025, where the lower edge falls eleven months in a single day. It is faithful to the method and it will still be read as a defect.
- The stated conclusion no longer matches the chart. The commentary above puts the gap at three to nine months. On the current data the center line spends 24% of the window above nine months and peaks at 13.4; the band reaches 15.5. The most recent reading is 7.4 months, with a band of 4.7 to 8.2.
Minor. The Simulations view silently drops any world that leaves the axis or fails to resolve on screen — 735 of 800 are drawn, and the omitted ones are precisely the extreme ones, so that view is narrower than the band it exists to illustrate. The tooltip prints the Chinese model’s interval but only the US model’s point estimate, on a chart whose whole subject is the intervals. And the exclusion of LLaMA-65B and Baichuan1-7B noted beside the chart applies to the center line and the release timeline; the resampling pool contains every scored model.
Two charts in circulation imply overconfident conclusions.
Two recent visualizations frame the US-China model gap in opposite directions. The U.S. Center for AI Standards and Innovation's evaluation of DeepSeek V4 Pro fits parallel Elo-vs-time trend lines for U.S. and PRC frontier models, suggesting a steadily widening U.S. lead and that Chinese labs are structurally unlikely to close it. Epoch AI's analysis of the Epoch Capabilities Index instead tracks each frontier release as a step function and reads the gap as roughly stable, or modestly narrowing.
We believe both readings overstate the confidence that the underlying data can support. Reasonable choices about which models count as “frontier,” which of these models are measured, and which benchmarks to include in a composite produce qualitatively different conclusions. We provide a visualization based on a Monte Carlo simulation that we believe more accurately captures this uncertainty. Our best guess is that Chinese models are 3 - 9 months behind the US frontier.