Work
Recent work and notes.
Sidequest Notes are short-form dispatches which may not reflect the views of MCNAIR as a whole. Research posts are slower, more formal results.
Subscribe to the MCNAIR Newsletter-
Sidequest Note
You can just not do things.
We present arguments for resting to be more productive, especially in high-stress fields such as AI risk. Researchers should preserve their agency by working sustainably.
-
Sidequest Note
Yet another post on AI and mathematics.
Twenty-five Fields Medalists have condemned the way AI mathematical results are announced. The speed of progress, not the way in which this progress is communicated, should be the main focus of their ire.
-
Sidequest Note
Persona Vectors don’t work on real data, or: Why you should stop overusing synthetic data for your research.
The result where projecting a prompt’s final-token activation onto a persona vector predicts how much the response will express the trait reproduces on the paper’s own LLM-generated eval prompts, and drops to roughly zero on real chat data.
-
Research
Review of the CB risk determination in the Claude Mythos 5.1 System Card
While we agree with Anthropic’s bottom-line conclusions about chemical and biological weapon risk, we find potential errors in their automated and human-led evaluations, and find the lack of third-party independent assessments highly concerning.
-
Sidequest Note
通用漏洞评分系统 4.0 版计算器。
A Simplified Chinese edition of the NVD CVSS v4.0 calculator. All 32 metrics, 116 values and their official definitions are translated; the scoring is ported from FIRST’s reference implementation and checked against it across 304,976 vectors.
-
Sidequest Note
The US–China AI gap is very uncertain.
For every day, how many months earlier a comparable US model was already out — measured against the Epoch Capabilities Index, with a Monte Carlo band for the uncertainty in the scores.
-
Sidequest Note
Crushing Nondeterminism in LLM Inference.
In 100,000 “flip a coin” completions, Claude Sonnet 4.5 returns heads ~100% of the time and Opus 4.6 ~99.26%. MCNAIR has therefore standardized Sonnet 4.5 calls as our primary mechanism for making fair executive decisions.