Introduction
I have written about Europe’s AI sovereignty problem three times now, and every time I land in the same uncomfortable place. The two pillars Europe leans on (the closed US frontier and the open Chinese weights frontier) are both rotting, and the continent I live in still has no pillar of its own.
Each post made me restate the same facts, updated. Who shipped what this week. Which weights slipped behind an API. Which score moved. After the third pass I noticed I was doing something a blog post is bad at: keeping a running tally.
So this is the running tally. A living scoreboard. I will update it every time a frontier model lands, an AA index version bumps, or a set of weights quietly moves behind an endpoint. The date at the top of the scoreboard section tells you when I last touched it.
If you want the long argument for why this matters, read the two-pillars posts first (Green 2026a, 2026b). This page is the scoreboard. The argument lives there.
The two pillars, restated in one paragraph
Europe runs its frontier AI on two pillars we do not control. Pillar one is the closed US frontier (Anthropic, OpenAI, Google, xAI, and now Meta) accessed through APIs whose terms can change overnight based on a government we have nothing to do with. Pillar two is the open Chinese weights frontier (Kimi, GLM, and the still-promised Qwen3.8-Max weights) which ships real weights today but ships them from a jurisdiction consulting on restricting overseas access to its most advanced models (The Next Web 2026). Both pillars can be pulled. We saw pillar two wobble in slow motion when Qwen3.8-Max launched API-only with weights “promised within days”. We saw pillar one’s political risk go from theoretical to concrete when Anthropic tried corporate self-regulation for military use and got punished so publicly that no rational actor will try it again (Anthropic 2026).
Renting frontier AI from San Francisco or renting it from Hangzhou is the same dependency with different invoices. That is the whole scoreboard in one sentence. The rest is detail.
What changed
Before the table, two things in this snapshot surprised me and are worth naming, because they sharpen the scoreboard rather than blur it.
The US no longer has an open-weights frontier model. Meta’s frontier is now Muse Spark 1.2, and it is proprietary. Llama used to be the American open-weights insurance policy. That policy lapsed. The only open-weights models within shouting distance of the frontier are Chinese (Kimi K3 at 59.7, GLM-5.2 at 52.6) and a single European one far below (Mistral Medium 3.5 at 30.4). OpenAI’s open line, gpt-oss-120b, scores 24.1. So pillar two is not just “the open frontier”. It is, for practical purposes, “the Chinese open frontier”. There is no American leg to stand on if Beijing tightens export controls and Washington restricts the closed APIs in the same quarter.
Google fell off the frontier top. Gemini 3.6 Flash sits at 51.6, around rank twenty-three, behind Kimi K3, GLM-5.2, and every other lab in the table. For a long time the frontier story was a three-horse race between Anthropic, OpenAI, and Google. The scoreboard says it is a two-horse race at the top (Anthropic and OpenAI) and a crowded field below, with the open Chinese models closer to the frontier than Google is.
The scoreboard
Snapshot date: 2026-08-11. Scores from the Artificial Analysis Intelligence Index v4.1, pulled from their Data API (free tier) (Artificial Analysis 2026b, 2026a). Each row is the top reasoning variant of that lab’s frontier model. Treat every number as a point-in-time reading.
| Rank | Lab | Region | Latest frontier model | Open? | AA index | EU access | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | Anthropic | US | Claude Opus 5 | Proprietary | 63.1 | API | Revocable (ToS and political risk) |
| 2 | OpenAI | US | GPT-5.6 Sol | Proprietary | 60.9 | API | Revocable |
| 3 | Moonshot | China | Kimi K3 | Open weights | 59.7 | Self-host | Revocable (Beijing consulting export limits) |
| 4 | Alibaba | China | Qwen3.8 Max | Proprietary, weights promised | 58.1 | API-only today | Revocable (grace period ticking) |
| 5 | Meta | US | Muse Spark 1.2 | Proprietary | 56.8 | API | Revocable (Meta closed its frontier) |
| 6 | xAI | US | Grok 4.5 | Proprietary | 55.8 | API | Revocable |
| 7 | Z.ai | China | GLM-5.2 | Open weights (MIT) | 52.6 | Self-host | Revocable |
| 8 | US | Gemini 3.6 Flash | Proprietary | 51.6 | API | Revocable | |
| 9 | Mistral | EU | Mistral Medium 3.5 | Open weights | 30.4 | Self-host | EU pillar, under construction, far from frontier |
| 10 | OpenAI (open line) | US | gpt-oss-120b | Open weights | 24.1 | Self-host | Stable, but far from frontier |
| n/a | EUROPA / Domyn | EU | In training | TBD | TBD | EU native | Under construction, >400B params, 24 EU languages |
A few cells deserve footnotes. Kimi K3 is the leading open-weights model in the world right now, and it is Chinese (Moonshot AI 2026; Artificial Analysis 2026c). Qwen3.8 Max is proprietary today with weights promised (Alibaba 2026), which is why its row reads “Closed preview” in spirit even though the number exists. Meta’s row is the one that moved most since I last wrote: Muse Spark is the frontier, and it is behind an API (Meta 2026). The gpt-oss-120b row is there to make a point. An open-weights model scoring 24 is a research artifact. You cannot run a continent on it.
How to read it
A few things the table does not show on its own.
The AA index moves week to week as it re-aggregates nine benchmarks (GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR). A model can gain or lose a point or two without a new release, just because the index re-sampled. So a one-point gap between Opus 5 and Fable 5 is noise. The signal is which labs are within shouting distance of the frontier top, and which access regime they ship under. The gap between row 8 (Google, 51.6) and row 9 (Mistral, 30.4) is where the frontier stops and the rest of the field begins.
“Open?” is the column that matters most for Europe, and it is the one most likely to lie to you. Qwen3.8 Max carries a score of 58.1 today, and that score is for the proprietary preview endpoint. The weights are promised. Alibaba has fenced its Max tier across two generations now (Qwen3.7-Max proprietary, Qwen3.8-Max paid preview with no weights). The pattern is the signal. The promise is the noise.
“EU access” sounds like a yes/no question. It is not. API access means you can call the model today and you cannot run it yourself tomorrow if the provider decides you should not. Self-host means you have the weights on hardware you control, and the only thing that can take that away is a license change you can see coming or a jurisdiction deciding to restrict exports. The two are different kinds of access, even though both let you do inference right now.
Where Europe sits
Two rows at the bottom of the table are the ones I care about most, and the ones with the most question marks.
Mistral is the closest thing Europe has to a frontier lab, and the ASML-led 1.7 billion euro Series C put real compute behind that (CNBC 2025). But Mistral Medium 3.5 scores 30.4 on the Intelligence Index. That is not a frontier score. It is a competent mid-tier open-weights score, roughly half of Opus 5. The frontier-tier question (does Mistral ship a model that sits in the top three of the AA index, ever) is still open. Mistral is the European pillar under construction, and “under construction” is the generous reading.
The EUROPA consortium led by Domyn is the other one to watch. The EU Frontier AI Grand Challenge is funding a >400 billion parameter model trained on all 24 EU languages on a 6,000-chip Nvidia Blackwell cluster (European Commission 2026). That is a real project with real compute. It is also not a frontier model yet, and the gap between “we are training a 400B model” and “we have a model in the top three of the AA index” is exactly the gap Europe has been failing to close for two years.
There is no European row in the top half of the table. There is no European row above row 30 on the index. That is the scoreboard.
What would change the table
A few specific things would force an update beyond the normal score drift:
- Qwen3.8-Max weights actually ship. The row flips from “Proprietary, weights promised” to “Open weights” and the grace period resets. I will believe it when the model card lands on HuggingFace.
- Beijing formalizes the overseas-access restrictions it has been consulting on. Several Chinese rows flip their verdict from “Revocable” to “Restricted” and pillar two gets measurably weaker for Europe.
- A US provider changes terms in a way that breaks EU access (region lock, use-case restriction, government-mandated cutoff). The corresponding US row gets a date stamped on its verdict.
- Mistral or EUROPA publishes a model that lands in the AA index top ten. A European row leaves the bottom of the table.
- Meta re-opens a frontier-tier Muse Spark, or OpenAI ships an open-weights model above 40. The US regains an open-weights frontier leg, and the two-pillar story gets a third leg I did not expect.
- A new lab enters the frontier. The table grows a row.
If none of those happen, the table still moves, because the AA index moves. That is the point of a living scoreboard. The structure stays stable. The numbers keep it honest.
How this post will live
I will update this post in place. The snapshot date at the top of the scoreboard section is the version.
The argument will not change much. The two-pillars thesis has aged the way I feared, and the scoreboard is the receipt. What changes is the receipt. New models, new scores, new access regimes, new European rows (I hope).
If you spot a number that is wrong, or a question mark you can fill, reach out. I would rather have a scoreboard with real gaps than one with confident mistakes. The whole point of publishing this is that the picture moves, and I am one person with one reading of it. The numbers this round come straight from the Artificial Analysis Data API, so at least the arithmetic is not mine to get wrong.
Conclusion
I started writing about European AI sovereignty because I kept noticing the same thing every time a frontier model landed. The good news was always about someone else’s model. The bad news was always about someone else’s decision. Europe’s role in the story was to read the announcement and adjust its compliance documentation.
The scoreboard makes that visible in one place. Eight non-European rows hold the frontier. Two European rows hold the bottom, far off the index. Everything in between is the access regime we rent from people who owe us nothing. The one piece of good news in the structure is that the open-weights frontier is genuinely competitive right now, with a Chinese model (Kimi K3) sitting third in the world. The piece of bad news attached to it is that the entire open-weights frontier is now Chinese, because Meta closed Muse Spark.
The grace period on the open Chinese weights is real. It is also revocable, and we watched it wobble in July. The closed US frontier is available. It is also revocable, and we watched the political risk go concrete in March. The European pillar is under construction. It is the only one whose verdict we get to write ourselves.
If you are building one of the European rows, I would like to hear from you. If you think I have the framing wrong, tell me which row and why. The scoreboard is the point. The conversation is the point of the scoreboard.



