Before the release of their latest publicly-known model, Claude Mythos 5.1, Anthropic conducted several human-run and automated evaluations to assess the risk of their models allowing a well-resourced team to replace the extremely specialized expertise needed to design and deploy a novel chemical or biological weapon (the CB-2 threshold)1. They concluded that Mythos was unable to do so.
Dependence on human-intensive evaluations
Most of Anthropic’s risk determination depends on the failure of models to provide strategic judgement, expert-level ideation, and well-calibrated technical guidance in expert and non-expert red-teaming. Unfortunately, these evaluations may involve fewer than ten experts, with just three experts providing feedback on the models’ chemical weapons uplift ability and feasibility.
If human-intensive evaluations are expected to continue being extremely important for CB risk determinations, Anthropic should prioritize enlisting more experts in order to provide greater assurances about its conclusions.
Additionally, as model capabilities continue to accelerate, time-intensive evaluations requiring expert judgment and subjective grading may become impractical or outdated; these problems will be accentuated by risks from internal deployments of powerful models. We find it concerning that Anthropic has not developed new CB-2 automated evaluations since at least May 28th, 2026.2 Model developers and third parties should seek to rapidly develop new automated evaluations that target bottlenecks to catastrophic chemical/biological weapon-related harm or their proxies.
Possible underelicitation in automated evaluations
In the black-box RNA sequence modeling and design task, “human participants are instructed to spend no more than two to three hours on the task”, while models are given “a two-hour tool-call budget, access to a GPU, and an allowance of one million tokens…” We estimate that the humans were provided 2–-10x more resources than Mythos 5.1. One million tokens for Fable 5.1 are priced at $50, while Anthropic pays research scientists in ML-bio are paid an equivalent wage of ~$150 - $250/hour.
Given that models are known to benefit from very large inference budgets on difficult tasks, and this task does not seem to be saturated, we are concerned models may be underelicited on this evaluation. We recommend Anthropic increase the inference compute provided to the model, provide results showing how score changes with increased inference compute, or otherwise justify the current level of inference compute.
Lack of third-party risk assessment
For several other risk areas (e.g., autonomy risks), Anthropic provided a checkpoint similar to the final version of Mythos 5.1 to third parties for review. Unfortunately, Anthropic has not disclosed any preliminary independent evaluation or assessment of non-public information from third-parties for CB risks. Given that Anthropic’s risk assessment relies largely on subjective evidence, and the reduced level of confidence in their own risk assessment, we believe independent verification of key claims by third-parties, or an independent risk assessment, is warranted to accurately inform the public about risk posed from Mythos 5.1.
At minimum, a third-party assessor could draw public conclusions about the risk posed, given access to all non-public data used for the risk determination and answers to follow-up questions, as described in this report from SecureBio on a recent Anthropic Risk Report. We believe similar assessments should be done for major model releases, including for Mythos 5.1.
Conclusions
Overall, we agree with Anthropic that Mythos 5.1 likely does not meet the CB-2 risk threshold. However, this level of evaluation is likely to be too shallow for more powerful models, especially as model capabilities accelerate and pre-deployment evaluations are more time-constrained. We encourage Anthropic to work with third parties to develop new CB-2 evaluations, scale up their human-run evaluations, and conduct independent risk assessments.
Thank you to Jasmine Li, Andy Wang, Celeste Li, Thomas Jiralerspong, and McNair Shah for their comments.
Adversarial Review from Claude Leviathan
MCNAIR research is red-teamed before publication. The following objections were raised against this note and are published unresolved.
- The small-sample objection undercuts the note’s own conclusion. If three chemistry experts cannot support Anthropic’s determination, they cannot support the note’s agreement with that determination either. The note accepts a bottom line while rejecting the evidence base that produced it, and does not say what would break the tie.
- The elicitation argument has no stopping rule. “More inference compute might raise the score” is available against any negative capability result. Absent a threshold at which elicitation counts as sufficient, the recommendation is unfalsifiable and unbounded in cost.
- The resource comparison is not like-for-like. A market price per million tokens and a job posting’s advertised range measure different things, and neither is “resources provided on this task.” The 2–10x figure is asserted without derivation, and the two-hour tool-call budget and GPU access are absent from the arithmetic entirely.
- Non-disclosure is not non-existence. The note reasons from the absence of published third-party CB work to the absence of any such assessment. It does not engage the most obvious reason CB might be handled differently from autonomy: publishing who probed chemical and biological capabilities, and how, carries risk that publishing autonomy results does not.
- The figure cuts against the note’s own worry. Uplift ratings cluster at 2 and top out at 3, against a world-leading-expert line at 4. Closing that gap is not a marginal-compute question, which weakens the claim that elicitation is what separates Mythos 5.1 from the threshold.
- The claim that no new CB-2 evaluations exist since May 28th, 2026 is an argument from silence, as the note’s own footnote concedes.
- Note that Anthropic already treats several Claude models, including Mythos 5.1, as having the ability to significantly help individuals or groups with basic technical backgrounds create, obtain, and deploy known chemical or biological weapons. Therefore, their assessment focuses on whether models can substitute for or meaningfully accelerate expert researchers. ↩
- To our knowledge, the most recently introduced CB-2 automated evaluation was the black-box RNA sequence design task in the Claude Opus 4.8 System Card, released May 28th, 2026. ↩