TailoredByte
All insights
Sep 7, 2026

The Models You Can't Buy

Frontier labs now use their best AI models internally — for math, drugs and physics — before selling them. What the Sep 2026 numbers say about the gap. (152)

By Konrad — CEO

Przeczytaj po polsku
ON THIS PAGE
  1. 01How far ahead the internal models really are
  2. 02Compute is short, and the compute that exists earns more inside
  3. 03The narrative shift: from taking people's jobs to discoveries
  4. 04Mathematics has already sped up
  5. 05Anthropic goes into biology
  6. 06What's on Sam's and Dario's desks?

On September 6, OpenAI published a note on how its own research organization works: the median researcher burns more than $600 a day in tokens, the 90th percentile more than $7,000 a day, and for every day of human work there are 3.1 days of agent work [1].

How far ahead the internal models really are

The only measurement that exists was done by METR: four labs (OpenAI, Anthropic, Google, Meta) gave it access to their internal models in February and March, and it concluded that the internal frontier is "on average ~2 months ahead" of the public one [2]. Some AI-market analysts put it at 3–4 months for post-trained models (up to half a year for raw checkpoints); others call it "two generations".

The ten mathematical results OpenAI announced on August 1 came from "an internal version of Astra"; Astra went on sale on September 3, 33 days later [3][4]. To prepare and compute those results it must have been available internally for a while already — my bet is at least a month.

Claude formalized Fermat's Last Theorem between August 7 and 18 on "an internal research model roughly comparable to Claude Fable 5.1"; Fable 5.1 shipped on September 1 [5].

Compute is short, and the compute that exists earns more inside

Sarah Friar, OpenAI's CFO, said in April that there are "things we're not pursuing because we don't have enough compute" [6].

Add to that: if you calculate the future value of a token, a token spent on improving your own model is probably worth more than a token sold to a customer for code generation, so the compute will flow to self-improvement, not to customers.

That squares with the AI 2027 forecast (the share of compute sold externally falls from about 30% in 2024 to 13% in 2027 [7]) and with the numbers in the OpenAI note I started with. If you own one shovel and you've just found gold in your backyard, you don't lend it to the neighbor.

But this isn't only about improving models.

The narrative shift: from taking people's jobs to discoveries

Since we're on narrative, let's look at what the CEOs of the biggest players are saying. May 27, Altman: "I'm delighted to be wrong about this" — he had expected a bigger hit to entry-level jobs [8]. July 12: he's "pretty sure" AI has been net job-creating [9]. August 25, on David Senra's podcast: "we've all been too ambitious on timelines"; the economy has inertia, and that's fine [10]. August 26, in TIME, on Astra: "I expect this will be the first model where the model actually invents new things in a way that matters" [11]. Amodei, August 15: saying that AI will cure cancer "is more a cliché than it is inspiring… The thing that will work is actually curing cancer" [12]. Zuckerberg, August 10: "Invention, not automation, will be the greatest contribution of superintelligence" [13].

A year ago AI's greatest ambition was to take our jobs, and then "it transitioned beyond that in a heartbeat to: I don't even care about your job. I have deeper thoughts that I'm working on".

If a company can make money on new drugs, new physics, materials and energy, it doesn't have to make money replacing people in offices — those are enormous markets, with fewer voters to anger (71% of Americans don't want a data center near them [14]). The change of subject is visible; changing the business model may not be that easy at all.

Mathematics has already sped up

Astra, i.e. GPT-6, came out on September 3. On FrontierMath Tier 4ᵃ, the hardest closed math benchmark, it scores 97.6% [4]. On Epoch's new benchmark — 68 open Erdős problems — it solved two [15].

A day later Anthropic announced that Claude had formalized Fermat's Last Theorem in Leanᵇ: 13 million lines, over 30,000 theorems, 11 days, about 6 billion tokens [5].

OpenAI's ten August results stand on theorems by Gábor Kun and Kun–Thom from the Rényi Institute in Budapest [16]; the May Erdős result was verified by Alon, Gowers, Shankar, Tsimerman and Bloom [17]; the Fermat code was read by Buzzard [18]; the counterexample to the Jacobian conjecture and the proof of Sendov's conjecture were "digested" by Terence Tao [19]. Tao described this at the ICM congress as three stages — generation, verification, understanding — of which AI speeds up the first two, and the third stays with humans [20].

That's why the labs are buying those humans: John Jumper (Nobel for AlphaFold) moved to Anthropic, Alpöge and Furman check Claude's proofs there, black-hole physicist Lupsasca works at OpenAI for Science, Harmonic pays mathematicians stipends, Axiom publishes with Ken Ono [21]. The model generates. The human picks the question, checks, and understands. And gets paid a salary the lab covers without blinking. They can afford it.

Anthropic goes into biology

This year Anthropic is opening wet labs and hiring biologists, acquiring Coefficient Bio (reportedly about $400 million in stock), bringing in John Jumper (June), launching Claude Science and its own drug-discovery program for rare diseases (June 30), and on September 1 releasing Mythos 5.1 with the biology safeguards unlocked only for verified organizations, "in partnership with the US government." On top of that it is showing protein design with a hit rate of about 50% against the typical 10–15% [22][23]. CEO Amodei promises "incredible results in the coming months" [12].

For now this is only the beginning. Today we have one AI-designed drug in Phase III (Insilico's rentosertib, since July) [24], zero approved, and AI molecules pass Phase I 80–90% of the time and Phase II about 40% — the same as ordinary ones [25].

AI is improving chemistry. Not biology yet. For now.

What's on Sam's and Dario's desks?

September 3: "much, much, much more capable models coming soon" (Sam Altman) [26]. If what we get publicly is Fable 5.1 and, a moment later, Astra, what do OpenAI and Anthropic already have available? Will we ever find out? Probably — but this is only an opinion — we won't be buying the best models and superintelligence in a subscription, just like that. Maybe that's for the better.


Glossary

FrontierMath Tier 4 — the hardest level of Epoch AI's mathematics benchmark (research-level problems); the v2 release of June 2026 corrected errors in about 42% of the problems. OpenAI funded the benchmark and has exclusive access to a subset of the problems — Epoch discloses this itself.

Lean — a proof assistant. A proof "checked in Lean" means a computer verified every step from the axioms; it does not mean the proof is readable by humans, or new.

Sources

[1] OpenAI — "Research acceleration: The view inside OpenAI" (Sep 6, 2026): researchers' token spend, 3.1 agent-workdays per human workday, the "research intern" goal reached: https://openai.com/index/research-acceleration-view-inside-openai/

[2] METR — Frontier Risk Report (May 19, 2026): measurement of the internal-vs-public model lead ("on average ~2 months"): https://metr.org/blog/2026-05-19-frontier-risk-report/

[3] OpenAI — "Ten advances in mathematics and theoretical computer science" (Aug 1, 2026): results from "an internal version of Astra": https://openai.com/index/ten-advances-in-mathematics/

[4] OpenAI — GPT-6 Astra (Sep 3, 2026): release, 97.6% on FrontierMath Tier 4: https://openai.com/index/gpt-6-astra/

[5] Anthropic — "Formalizing Fermat's Last Theorem" (Sep 4, 2026): 13M lines of Lean, 30,300 theorems, 11 days, ~6B tokens, internal research model: https://www.anthropic.com/research/formalizing-fermats-last-theorem

[6] Business Insider (via AOL) — Sarah Friar on the compute shortage (Apr 2, 2026): https://www.aol.com/articles/openais-cfo-says-company-passing-112355988.html

[7] AI 2027 — Compute Forecast: a lab's compute allocation in 2024 vs 2027: https://ai-2027.com/research/compute-forecast

[8] The Decoder — Altman and Amodei walk back their job-loss predictions (May 27, 2026): https://the-decoder.com/sam-altman-and-dario-amodei-walk-back-their-ai-job-apocalypse-predictions/

[9] The Decoder — Altman: AI is net job-creating (Jul 12, 2026): https://the-decoder.com/openai-ceo-altman-is-now-pretty-sure-ai-is-net-job-creating-which-is-quite-the-pivot-from-predicting-mass-layoffs/

[10] The Next Web — Altman on David Senra's podcast on over-ambitious timelines (Aug 25, 2026): https://thenextweb.com/news/sam-altman-ai-timelines-wrong-inertia-senra-podcast

[11] TIME — "Inside OpenAI's Reboot" (Aug 26, 2026): Altman on Astra and AGI by year-end: https://time.com/article/2026/08/26/openai-sam-altman-interview/

[12] TechCrunch — Amodei on the "crisis of trust" and curing cancer (Aug 16, 2026): https://techcrunch.com/2026/08/16/anthropic-ceo-says-ai-backlash-is-fundamentally-a-crisis-of-trust/

[13] Meta — Mark Zuckerberg, "The Future is for Everyone" (Aug 10, 2026): https://about.fb.com/news/2026/08/the-future-is-for-everyone/

[14] Gallup — 71% of Americans oppose a data center nearby (May 14, 2026): https://news.gallup.com/poll/708620/less-support-solar-wind-energy-nuclear.aspx

[15] Epoch AI — "Announcing FrontierMath Erdős" (Sep 1, 2026): 68 open Erdős problems, Astra 2/68, all other models 0: https://epoch.ai/latest/announcing-frontiermath-erdos

[16] HUN-REN Rényi Institute — on the Kun and Kun–Thom theorems in OpenAI's proof (Aug 4, 2026): https://www.renyi.hu/en/news/openai-mathematical-breakthrough-builds-renyi-researchers-work

[17] OpenAI — disproof of Erdős's unit-distance conjecture (May 20, 2026), with verification by Alon, Gowers, Shankar, Tsimerman and Bloom: https://openai.com/index/model-disproves-discrete-geometry-conjecture/ ; Quanta Magazine — "Why the Legendary Erdős Problems Are Falling to AI" (Aug 3, 2026): https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/

[18] Kevin Buzzard — "FLT: Anthropic has beaten me to it" (Sep 4, 2026): https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/

[19] Terence Tao — "A digestion of the Jacobian conjecture counterexample" (Jul 21, 2026): https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/ ; "A digestion of the proof of Sendov's conjecture" (Aug 12, 2026): https://terrytao.wordpress.com/2026/08/12/a-digestion-of-the-proof-of-sendovs-conjecture/

[20] Terence Tao — "Mathematics in the age of AI", arXiv 2608.16753 (ICM lecture, Jul 24, 2026): https://arxiv.org/abs/2608.16753

[21] TechCrunch — John Jumper moves to Anthropic (Jun 20, 2026): https://techcrunch.com/2026/06/20/nobel-laureate-john-jumper-is-leaving-deepmind-for-rival-anthropic/ ; Anthropic — Alpöge and Furman verify Claude's result (Aug 10, 2026): https://www.anthropic.com/research/riemann-zeta ; Latent Space — Alex Lupsasca at OpenAI for Science (May 5, 2026): https://www.latent.space/p/lupsasca ; Harmonic — mathematician sponsorships (Jan 22, 2026): https://www.harmonic.fun/news/mathematician-sponsorships/ ; Axiom Math — paper with Ken Ono (Aug 18, 2026): https://axiommath.ai/research/the-weight-comes-last/

[22] SynBioBeta — Anthropic opens wet labs and hires biologists (May 7, 2026): https://www.synbiobeta.com/read/anthropic-is-hiring-biologists-building-wet-labs-and-betting-big-on-drug-discovery ; BioSpace — the Coefficient Bio acquisition (Apr 6, 2026): https://www.biospace.com/business/ai-giant-anthropic-leans-into-life-sciences-with-400m-coefficient-bio-catch ; MIT Technology Review — Claude Science and the drug-discovery program (Jun 30, 2026): https://www.technologyreview.com/2026/06/30/1139987/claude-science-is-anthropics-newest-flagship-product/

[23] Anthropic — Claude Fable 5.1 and Mythos 5.1 (Sep 1, 2026): Life Sciences Verification Program, protein design: https://www.anthropic.com/claude-fable-and-mythos-5-1

[24] Insilico Medicine — Phase III trial of rentosertib begins (Jul 7, 2026): https://www.prnewswire.com/news-releases/insilico-initiates-phase-iii-clinical-trial-for-rentosertib-its-ai-empowered-tnik-inhibitor-for-idiopathic-pulmonary-fibrosis-302819553.html

[25] Jayatunga et al. — "How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons", Drug Discovery Today 2024, DOI 10.1016/j.drudis.2024.104009

[26] Axios — interview with Sam Altman (Sep 3, 2026): https://www.axios.com/2026/09/03/axios-interview-sam-altmans-sobering-siren