Introduction
I have written about Europe’s AI sovereignty problem four times now, and every time I land in the same uncomfortable place. The two pillars Europe leans on (the closed US frontier and the open Chinese weights frontier) are both rotting, and the continent I live in still has no pillar of its own.
Each post made me restate the same facts, updated. Who shipped what this week. Which weights slipped behind an API. Which score moved. After the third pass I noticed I was doing something a blog post is bad at: keeping a running tally.
So this is the running tally. A living scoreboard. I will update it every time a frontier model lands, an AA index version bumps, or a set of weights quietly moves behind an endpoint. The date at the top of the scoreboard section tells you when I last touched it.
If you want the long argument for why this matters, read the two-pillars posts first (Green 2026a, 2026b). This page is the scoreboard. The argument lives there.
The two pillars, restated in one paragraph
Europe runs its frontier AI on two pillars we do not control. Pillar one is the closed US frontier (Anthropic, OpenAI, Google, xAI, and now Meta) accessed through APIs whose terms can change overnight based on a government we have nothing to do with. Pillar two is the open Chinese weights frontier (Kimi, DeepSeek, the Qwen3.8 open line, and the GLM open line) which ships real weights today but ships them from a jurisdiction consulting on restricting overseas access to its most advanced models (The Next Web 2026). Both pillars can be pulled. We saw pillar two wobble in slow motion when Qwen3.8-Max launched API-only with weights “promised within days”, watched it partly reset on 2026-08-13 when the 2.4T-A95B weights landed on HuggingFace, watched it wobble again when Z.ai shipped GLM-5.3 API-first and delayed the weights over the model’s own cyber capability, and then watched the promise kept: the Flash tier landed open on 2026-08-26, with the full 5.3 weights due two weeks to the day (Z AI 2026a, 2026b). We saw pillar one’s political risk go from theoretical to concrete when Anthropic tried corporate self-regulation for military use and got punished so publicly that no rational actor will try it again (Anthropic 2026).
What changed
Before the table, three things in this snapshot moved the board, and two of them test the same promise: whether “weights promised” means anything.
The GLM-5.3 promise held. On 2026-08-26 Z.ai released GLM-5.3-Flash with open weights under MIT: 320B total parameters, 18B active, the first natively multimodal model in the GLM-5 series, built on a new hybrid of sparse and linear attention (Z AI 2026b). Before anyone knew whose it was, the model had spent six days at the top of OpenRouter under the anonymous name “Ox Alpha” (ToolScout 2026). Artificial Analysis grades it at 57.5, behind only Kimi K3 (59.7) and the open Qwen 2.4T (57.7) among open-weights models, and ahead of everything Meta and Google currently ship. The full GLM-5.3 weights are due on HuggingFace today, 2026-08-28, two weeks to the day after the API launch, which is the schedule Z.ai named when it delayed the release for safety hardening (Z AI 2026a). At snapshot time they had not landed. I will update the row the moment they do.
Qwen previewed Qwen4 in the open. Qwen3.8-Flash-Next (55.8, 2026-08-26) is an open-weights architecture preview: a 125B main model plus 51B n-gram embeddings and 4B multi-token prediction, 6B active per token, trained at roughly one-ninth the cost of Qwen3.7-Plus (Alibaba 2026a). Qwen3-Next played this role for Qwen3.5, and that hybrid design then spread across the whole Qwen3.8 line. Whatever ships as Qwen4 will be built on this architecture. The open frontier is no longer just keeping pace with the closed one; it is previewing the next generation first. One nuance for the “Open?” column: Flash-Next ships under a Qwen community license, not the Apache 2.0 of the 27B.
The Qwen Max hybrid is now four weeks past its launch-day promise of “within days”. Set the two Chinese labs side by side. Z.ai promised weights two weeks after launch, delayed them for safety hardening, and delivered the Flash tier on schedule. Alibaba promised the Max hybrid on launch day, shipped everything below the Max line (2.4T, Flash-Next, 27B), and has gone quiet on the top checkpoint. One promise kept, one promise stale. That difference is the whole “Open?” column in miniature.
Two entries round it out, neither from the US or China. Agnes 2.5 Pro (proprietary, 49.1, 2026-08-26) comes from Sapiens AI, a Singapore lab founded in 2025 (Sapiens AI 2026), and it is the first model from outside the US, China, and Europe to come within shouting distance of this board. It does not clear the frontier window, so it is not a row. Korea’s Motif Technologies shipped Motif 3 (314B total, 13.2B active, MIT) at 47.4 (Motif Technologies 2026), the strongest open-weights model yet from neither China nor the US and the first footnote against the “the open frontier is entirely Chinese” line. It sits below the board cutoff. For now.
One structural change, added 2026-08-24. The table is now generated straight from the Artificial Analysis Data API instead of a hand-kept list: every lab’s best model qualifies automatically when it lands within ten points of the top, and every European lab is on the board no matter what it scores. The Grok 4.6 miss cannot happen again. It also means the board grew two European rows that snapshot: Multiverse Computing’s HyperNova 60B (18.3) and the Swiss AI Initiative’s Apertus 70B (2.0). Neither is anywhere near the frontier. That is exactly why they are shown.
The scoreboard
Snapshot date: 2026-08-28. Scores from the Artificial Analysis Intelligence Index v4.1, pulled from their Data API (free tier) (Artificial Analysis 2026b, 2026a). Each row is the top reasoning variant of that lab’s frontier model, selected automatically: a lab enters when its best model is within ten points of the board’s top score, European labs enter regardless of score, and the curated open-line rows carry the open-weights argument. The Δ top column is each row’s distance from the board’s best score, because the comparison to the frontier is what the table is for. Treat every number as a point-in-time reading.
| Rank | Lab | Region | Latest frontier model | Open? | AA index | Δ top | EU access | Verdict |
|---|---|---|---|---|---|---|---|---|
| 1 | Anthropic | US | Claude Opus 5 | Proprietary | 63.1 | 0.0 | API | Revocable (ToS and political risk) |
| 2 | OpenAI | US | GPT-5.6 Sol | Proprietary | 60.9 | −2.2 | API | Revocable |
| 3 | xAI | US | Grok 4.6 | Proprietary | 60.9 | −2.2 | API | Revocable |
| 4 | Moonshot | China | Kimi K3 | Open weights | 59.7 | −3.4 | Self-host | Revocable (Beijing consulting export limits) |
| 5 | Z.ai | China | GLM-5.3 | API-first (Flash sibling open, MIT); full 5.3 weights due 2026-08-28 | 59.5 | −3.6 | API | Revocable (Flash sibling open; full 5.3 weights due 2026-08-28) |
| 6 | Alibaba | China | Qwen3.8 Max | Open weights (2.4T-A95B, 57.7); Max hybrid promised | 58.1 | −5.0 | Self-host (2.4T) · API (Max) | Revocable (2.4T open, Max hybrid pending) |
| 7 | Z.ai (open line) | China | GLM-5.3-Flash | Open weights (MIT) | 57.5 | −5.6 | Self-host | Revocable (Beijing consulting export limits) |
| 8 | Meta | US | Muse Spark 1.2 | Proprietary | 56.8 | −6.3 | API | Revocable (Meta closed its frontier) |
| 9 | US | Gemini 3.7 Flash | Proprietary | 56.0 | −7.1 | API | Revocable | |
| 10 | Alibaba (Flash-Next open) | China | Qwen3.8-Flash-Next | Open weights (Qwen community license) | 55.8 | −7.3 | Self-host | Revocable (Beijing consulting export limits) |
| 11 | DeepSeek | China | DeepSeek V4 Pro | Open weights | 53.2 | −9.9 | Self-host | Revocable (Beijing consulting export limits) |
| 12 | Alibaba (27B open) | China | Qwen3.8-27B | Open weights (Apache 2.0) | 52.0 | −11.1 | Self-host | Revocable (Beijing consulting export limits) |
| 13 | Mistral | EU | Mistral Medium 3.5 | Open weights | 30.4 | −32.7 | Self-host | EU pillar, under construction, far from frontier |
| 14 | OpenAI (open line) | US | gpt-oss-120b | Open weights | 24.1 | −39.0 | Self-host | Stable, but far from frontier |
| 15 | Multiverse Computing | EU | HyperNova 60B 2605 | Open weights | 18.3 | −44.8 | Self-host | EU pillar, under construction, far from frontier |
| 16 | Swiss AI Initiative | Europe (non-EU) | Apertus 70B Instruct | Open weights | 2.0 | −61.1 | Self-host | European, outside EU jurisdiction |
| n/a | EUROPA / Domyn | EU | In training | TBD | TBD | TBD | EU native | Under construction, >400B params, 24 EU languages |
A few cells deserve footnotes. Kimi K3 remains the leading open-weights model in the world, and it is Chinese (Moonshot AI 2026). The GLM-5.3 row’s openness cell carries a date: the Flash sibling (MIT, 320B total, 18B active) has been open since 2026-08-26 at 57.5, and the full 5.3 weights are due on HuggingFace today, two weeks after the API launch. They had not landed when this snapshot was pulled (Z AI 2026a, 2026b). The Qwen family now spans four openness states: a closed Max endpoint (58.1), an open 2.4T sibling (57.7), an open Flash-Next preview under a Qwen community license (55.8), and an Apache 2.0 consumer 27B (52.0), with the Max hybrid still just a promise (Alibaba 2026b, 2026a). The gpt-oss-120b row is there to make a point. An open-weights model scoring 24 is a research artifact. You cannot run a continent on it.
How to read it
A few things the table does not show on its own.
The AA index moves week to week as it re-aggregates nine benchmarks (GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR). A model can gain or lose a point or two without a new release, just because the index re-sampled. So a point or two between the rows at the top of the board is noise. The signal is which labs are within shouting distance of the frontier top, and which access regime they ship under. The gap between row 12 (Qwen3.8-27B, 52.0) and row 13 (Mistral, 30.4) is where the frontier stops and the rest of the field begins: 21.6 points, with no European model anywhere near the top of it.
“Open?” is the column that matters most for Europe, and it is the one most likely to lie to you. Two cells in it carry dates. GLM-5.3’s cell is a promise coming due: the Flash sibling has been open (MIT) since 2026-08-26, and the full 5.3 weights are due on HuggingFace today, the two-week safety delay Z.ai announced up front (Z AI 2026a, 2026b). The Qwen cell is a promise gone quiet: the Max hybrid was promised “within days” on 2026-08-03, and four weeks later the model card is still not on HuggingFace, while everything below the Max line (2.4T, Flash-Next, 27B) is open. Alibaba fenced its Max tier across two generations (Qwen3.7-Max proprietary, Qwen3.8-Max paid preview), and the 2026-08-13 open-weighting of the 2.4T-A95B broke the pattern for everything except the top. The 27B cell is the one with no promise and no asterisk. Apache 2.0, on HuggingFace, downloadable today.
“EU access” sounds like a yes/no question. It is not. API access means you can call the model today and you cannot run it yourself tomorrow if the provider decides you should not. Self-host means you have the weights on hardware you control, and the only thing that can take that away is a license change you can see coming or a jurisdiction deciding to restrict exports. The two are different kinds of access, even though both let you do inference right now.
Where Europe sits
The European rows at the bottom of the table are the ones I care about most, and the ones with the most question marks.
Mistral is the closest thing Europe has to a frontier lab, and the ASML-led 1.7 billion euro Series C put real compute behind that (CNBC 2025). But Mistral Medium 3.5 scores 30.4 on the Intelligence Index. That is a competent mid-tier open-weights score, roughly half of Opus 5. The frontier-tier question (does Mistral ship a model that sits in the top three of the AA index, ever) is still open. Mistral is the European pillar under construction, and “under construction” is the generous reading.
The EUROPA consortium led by Domyn is the other one to watch. The EU Frontier AI Grand Challenge is funding a >400 billion parameter model trained on all 24 EU languages on a 6,000-chip Nvidia Blackwell cluster (European Commission 2026). That is a real project with real compute. It is also not a frontier model yet, and the gap between “we are training a 400B model” and “we have a model in the top three of the AA index” is exactly the gap Europe has been failing to close for two years.
The two European rows below Mistral are there to show the depth of the field. HyperNova 60B (Multiverse Computing, 18.3) and Apertus 70B (Swiss AI Initiative, 2.0) are real, downloadable European open-weights models, and they sit 44.8 and 61.1 points off the frontier. Depth matters: Europe’s second-best graded lab is not close to Europe’s best, and Europe’s best is not close to anything.
There is no European row in the top half of the table. The best European lab sits 27th among the labs Artificial Analysis grades, ranked #174 once every graded variant is counted, 32.7 points off the top. That is the scoreboard.
What would change the table
A few specific things would force an update beyond the normal score drift:
- GLM-5.3 full weights land. The Flash tier landed open (MIT) on 2026-08-26, and the full 5.3 weights are due today, 2026-08-28, two weeks after the API launch. When they land, the Z.ai row flips from API-first to open and pillar two regains its strongest Chinese lab at full strength. If they slip again, a dated promise stops being a schedule.
- Qwen3.8-Max hybrid weights actually ship. The 2.4T-A95B shipped open on 2026-08-13, the 27B followed on 2026-08-14, Flash-Next on 2026-08-26, and the Max hybrid is now four weeks past the launch-day promise of “within days”. Z.ai has now kept a dated weights promise; Alibaba’s is the standing counterexample. I will believe the hybrid when the model card lands on HuggingFace.
- Qwen4 ships. Flash-Next is the open preview of the Qwen4 architecture, in the role Qwen3-Next played for Qwen3.5. If Qwen4 lands as a Max-class open drop, the top of the board is in play. If it lands API-only with open siblings below, the Qwen pattern repeats at the next generation.
- Beijing formalizes the overseas-access restrictions it has been consulting on. Several Chinese rows flip their verdict from “Revocable” to “Restricted” and pillar two gets measurably weaker for Europe.
- A US provider changes terms in a way that breaks EU access (region lock, use-case restriction, government-mandated cutoff). The corresponding US row gets a date stamped on its verdict.
- Mistral or EUROPA publishes a model that lands in the AA index top ten. A European row leaves the bottom of the table.
- Meta re-opens a frontier-tier Muse Spark, or OpenAI ships an open-weights model above 40. The US regains an open-weights frontier leg, and the two-pillar story gets a third leg I did not expect.
- A new lab enters the frontier. The table grows a row.
If none of those happen, the table still moves, because the AA index moves. That is the point of a living scoreboard. The structure stays stable. The numbers keep it honest.
How this post will live
I will update this post in place. The snapshot date at the top of the scoreboard section is the version.
The argument will not change much. The two-pillars thesis has aged the way I feared, and the scoreboard is the receipt. What changes is the receipt. New models, new scores, new access regimes, new European rows (I hope).
If you spot a number that is wrong, or a question mark you can fill, reach out. I would rather have a scoreboard with real gaps than one with confident mistakes. The whole point of publishing this is that the picture moves, and I am one person with one reading of it. The numbers this round come straight from the Artificial Analysis Data API, so at least the arithmetic is not mine to get wrong.
Conclusion
I started writing about European AI sovereignty because I kept noticing the same thing every time a frontier model landed. The good news was always about someone else’s model. The bad news was always about someone else’s decision. Europe’s role in the story was to read the announcement and adjust its compliance documentation.
The scoreboard makes that visible in one place. Twelve non-European rows hold the top of the field. The European rows start 32.7 points below the frontier and fall away from there. Everything in between is the access regime we rent from people who owe us nothing. The one piece of good news in the structure is that the open-weights frontier is genuinely competitive right now, with a Chinese model (Kimi K3) sitting fourth in the world, the open Qwen 2.4T within 0.4 of its closed sibling, and an open Chinese model (GLM-5.3-Flash) that now outscores every US flagship except the top three. The piece of bad news attached to it is that the open frontier near the top is still entirely Chinese, and the two exceptions are a research artifact and a footnote: OpenAI’s gpt-oss-120b at 24.1, and Korea’s Motif 3 at 47.4.
The grace period on the open Chinese weights is real. It is also revocable, and this snapshot is the first one where you can watch both outcomes at once: Z.ai delayed the GLM-5.3 weights for the model’s own offensive capability and then delivered the Flash tier on schedule, while Alibaba’s Max hybrid sits four weeks past a launch-day promise. A frontier you have to wait for, at the discretion of the lab that trained it, is a frontier you do not own. The closed US frontier is available. It is also revocable, and we watched the political risk go concrete in March. The European pillar is under construction. It is the only one whose verdict we get to write ourselves.
If you are building one of the European rows, I would like to hear from you. If you think I have the framing wrong, tell me which row and why.



