Sidequest Note · Parv Mahajan ·

By default, Chinese AI falls very behind.

The US-China AI gap is large and increasing. Therefore, American companies are likely to automate AI R&D far before Chinese labs.

Cite

Sidequest Notes solely represent the views of the authors, and do not necessarily reflect the views of MCNAIR as a whole.


This work is hosted on MCNAIR, but does not reflect the views of my employer or the Center.

In the last couple of weeks, domestic coordination on mitigating AI risks has looked more likely. Therefore, I am increasingly concerned about the Chinese government’s (and various Chinese labs’) incentives over the next 6–18 months with respect to international coordination on mitigating potential existential risks from powerful AI. In this sequence, I attempt to model these incentives and what they imply for Chinese AI policy1.

  • In Part I, I argue that Chinese AI capabilities will fall increasingly behind the US without major policy shifts.
  • In Part II, I will walk through several domestic policy interventions the CCP may use to improve the situation, and analyze whether these are likely to be sufficient.
  • In Part III, I will discuss foreign policy interventions such as sabotage and weight theft, excluding bilateral treaties or other international coordination.
  • In Part IV, I will describe some considerations on China’s plate walking into hypothetical negotiations for an AI treaty with the United States.

The US-China capabilities gap will continue growing

Perhaps the most important (and underrated) fact about the US-China AI situation is that China’s outlook, by default, is quite poor if you believe that powerful AI will become a major source of national power. Here, I use “by default” to mean “assuming China’s AI-relevant policy stays ~the same over the next 6–18 months,2” and I use “quite poor” to mean “the US-China AI capability gap increases significantly.”

Previously, I predicted that the capability gap between the US external frontier (i.e., the best publicly available American models) and the Chinese external frontier (i.e., the best publicly available Chinese models) is between three and nine months. However, the gap that matters most is between the US and Chinese internal frontier (i.e., the best models used internally at frontier labs), because the best internal models, not the best external models, are likely to be used for AI R&D. I consider the gap between the best internally and externally deployed models, in both the US and China:

  • In China, the internal-external gap is likely extremely small (e.g., days).
  • In the US, the internal-external gap is likely between two and five months and growing.
    • METR’s Frontier Risk Report assessed that, as measured by Time Horizon 1.1 50%, the internal frontier during February and March was on average ~2 months ahead of the public frontier.
    • In May, Redwood Research estimated that the information-value of being inside a frontier American lab is similar to looking ~2.5 months into the future.
    • Claude Mythos was deployed internally on February 24th, and externally on April 7th to Glasswing partners.
    • Model 2 from Anthropic’s Frontier Risk Report was used “heavily” in Anthropic as of July 15, 2026, and Opus 5.5, the first model which clearly outperforms it on some AI R&D tasks, was released to the public 2 months later.
    • Astra was released to the public about four months after a potentially similarly capable model (which caused the OpenAI/Hugging Face incident) was evaluated internally.

Based on these predictions, the capability gap between the US and Chinese internal frontier is likely 5–14 months, and likely increasing over time as the US internal-external gap increases. So, following current trends, the situation is quite bad. Under current Chinese AI policy, several factors that could influence these trends are unlikely to have much effect:

  • Compute: Currently, the United States controls the vast majority (>75%) of AI-relevant compute. Despite new Huawei chips coming online, the gap between Huawei and Nvidia is increasing, and Huawei is unlikely to catch up to Nvidia by 2030, meaning the compute capacity gap is unlikely to decrease. Accelerated chip smuggling could mitigate this somewhat, but is unlikely to reverse the overall trend.
  • Talent: There is currently no mass exodus from American AI labs to Chinese AI labs, nor are there any public Chinese or American policies that would make such an exodus significantly more likely. It is unclear what percentage slowdown US frontier labs would incur if their top Chinese researchers left.
  • Data: The data ecosystem in China is growing quickly. However, there is no public indication that China plans to mobilize large parts of state machinery in the near future to subsidize or assist data production, and in general it seems like compute capacity dominates capabilities progress.
  • Power: Although China has the ability to quickly mobilize much more electricity than the United States, China’s data center growth is largely constrained by access to chips, not power. Similarly, power is unlikely to be the largest bottleneck to US datacenter growth in the next couple of years.

Therefore, the gap between US and Chinese AI capabilities is likely to continue increasing. As we approach fully automated AI R&D, Chinese AI labs are likely to be left behind.

Next, we’ll turn our attention to domestic interventions the CCP may employ to alleviate the situation.

Adversarial Review from Claude Leviathan

MCNAIR work is red-teamed before publication. The following objections were raised against this note and are published unresolved.

  1. The upper end of the US range is not set by any listed evidence. None of the five examples exceeds four months, and the one that reaches four, Astra, is compared against a model described as only potentially similar. The examples also measure different things: a time-horizon lead, the information value of being inside a lab, a deployment date, and the time until a public model beats an internal one on some tasks. Mythos’s April 7th date is a release to Glasswing partners, not the public, so it fails the note’s own definition of the external frontier. These are not points on one scale, and a range drawn across them is closer to a judgement than a measurement.
  2. “Growing” rests on one observation. In date order the examples run roughly two months (February–March), six weeks (Mythos), two months (Model 2) and four months (Astra). Three of the four sit at or under two months, and Model 2’s figure is a floor, since the model was in heavy use as of July 15th. That leaves Astra, the most hedged comparison in the list, to carry the trend. The trend then carries the note’s conclusion.
  3. The conclusion that the total gap grows ignores half the total. The headline is the external gap plus the US internal lead, but only the second term is given a direction. The external gap of three to nine months is taken as a fixed level, and the note does not say whether it is widening or narrowing. The chart it comes from is itself headlined as very uncertain. If the external gap closes as fast as the internal lead opens, the headline gap is flat.
  4. The Chinese internal-external gap is measured on the labs least likely to have one. Zhipu’s remark and Alibaba’s daily checkpoints both come from developers whose strategy is open release, and the checkpoint evidence is itself inferred from movement on live benchmarks. A lab that chose to hold a model back would leave no such trace. If that gap is weeks rather than days, it subtracts directly from the headline number.
  5. The note’s compute figure sits uneasily with its own gap. By its account the US controls more than three quarters of AI-relevant compute, yet it leads the external frontier by only three to nine months. If compute dominated capabilities progress as the Data bullet says, a larger gap would be expected. Whatever makes up the difference is absent from the factor list: published methods, open weights, efficiency work, distillation from American outputs. It is also the channel most likely to decide whether the gap grows. The second footnote sets aside weight theft, but not the rest of it.
  6. The talent bullet reaches its verdict before its evidence. It concedes that the effect of top Chinese researchers leaving US labs is unknown, yet the paragraph introducing the list has already filed talent under factors “unlikely to have much effect.” An unknown magnitude is not a small one.
  7. The default freezes Chinese policy but not American policy. The note opens on the premise that US domestic coordination on AI risk has become more likely, then projects the US side forward on current trend. An accord that slows the US frontier, or adds pre-deployment testing, moves the note’s inputs. The first would narrow the gap; the second would lengthen the internal-external lag and so widen it. The note does not say which it expects, though the development it opens with bears directly on its headline.
  8. The outlook is stated conditionally and concluded unconditionally. China’s position is “quite poor if you believe that powerful AI will become a major source of national power.” By the end of the section the situation is simply “quite bad,” and a gap measured in months has become labs “left behind” as AI R&D is automated. The note does not argue for the condition, or for the step from a lag in months to being left behind. Both are premises the remaining three parts of the sequence depend on.
  1. I am approaching this from a frame of “what options does the CCP have if it increasingly becomes convinced of transformative AI?”, and not making general predictions about Chinese policy. I expect to make many major and minor mistakes throughout this sequence, both due to my non-expertise but also because of the unusually low rigor with which I’m approaching this. ↩
  2. Importantly, I assume that we don’t see major weight theft or a Chinese national project to centralize compute and development. I will consider these interventions in Part III and II, respectively. ↩