Introduction
A couple of weeks ago I wrote a piece called The Two Pillars Are Both Rotting. The argument was that most of the AI Europe actually runs sits on two pillars we don’t control: the closed frontier from the US (OpenAI, Anthropic, Google), and the open-weights frontier from China (Qwen, GLM, DeepSeek, Kimi). Both can be pulled, for different reasons, on different timelines, and nobody in European policy is planning for the day both wobble at once. I ended it with a plea for our own frontier labs.
I want to come back to it already, because two things happened in the space of ten days that changed half of what I said, and sharpened the other half into something I should have said the first time.
The first thing is that the closed-lab moat rotted faster than I argued. Kimi K3 shipped, and it shipped open weight, and on the only fair composite benchmark we have it sits three points behind the best closed model in the world. Three. The “frontier is a clear step ahead of open weights” line I drew in the last post is already stale.
The second thing is that nothing happened on the European side. Nothing. ASML wrote a €1.3 billion check for Mistral last September and Brussels still hasn’t decided whether it wants a frontier lab. The pillar-rotting risk I wrote about plays out over months and years, so two weeks of data doesn’t settle it either way. What this month does settle is the capability question: the closed-lab moat on the composite is down to three points, and Europe’s missing-pillar problem is exactly where I left it. We still don’t have a frontier lab of our own, and as I’ll argue below, that’s a choice we keep making.
The composite gap is three points now
The Artificial Analysis Intelligence Index, version 4.1, aggregates nine evaluations (GDPval, Terminal-Bench 2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, and a few others) into one composite on a 0-100 scale. As of mid-July 2026 the top of the board reads like this: Claude Fable 5 at 60, GPT 5.6 Sol at 59, Kimi K3 at 57, Claude Opus 4.8 at 56 (Artificial Analysis 2026).
Sit with that for a second. The largest open-weight model ever released, built by a Chinese lab (Moonshot AI), trained on Chinese compute, ships three composite points behind Anthropic’s most capable model (Fable 5 at 60), two behind OpenAI’s current flagship (GPT 5.6 Sol at 59), and a point above Claude Opus 4.8 (at 56). On a scale where the gap between frontier and a competent mid-tier model runs fifteen points or more, three points is a tie. These are peers on general intelligence.
In the last post I benchmarked GLM-5.2 against Claude Opus 4.8 and showed a real but narrow gap, and then stacked the actual frontier (Fable 5, GPT-5.5 Pro) against it to show the gap widens when you compare to the newest closed model rather than the comfortable one. That was fair to the data, and it was already a snapshot of a moving target. The target moved. The moving part is that the open-weight side shipped something that closed the composite gap to a margin of error, and the closed side shipped a new model (GPT 5.6 Sol, GA July 9 (OpenAI 2026)) that didn’t widen the gap back out. The moat is a three-point lead. That’s the whole moat, on the composite measure that tries to be fair about it.
What three points buys you
Ok fine, but so what? Three points is three points, and the people paying $50 per million output tokens for Fable 5 aren’t doing it for the composite. Let me show you what the three points actually are, because it matters for the Europe argument later.
The gap varies across the benchmark surface. It’s thin on general knowledge and reasoning, and real on long-horizon agentic work. On Humanity’s Last Exam and GPQA the open and closed frontier are within noise. On the coding-and-tool-use end, the picture splits by the shape of the work.
Folded out across twelve benchmarks, the picture is a lot less flat:
GPT 5.6 Sol leads DeepSWE at 73.0 and Terminal-Bench 2.1 at 88.8, with Kimi K3 essentially tied on the terminal test at 88.3. Fable 5 owns FrontierSWE at 86.6, five points ahead of K3 and fifteen ahead of Sol. Kimi K3 takes Program Bench at 77.8 and, the one I find most interesting, SWE Marathon at 42.0, a test of long sustained coding sessions where GPT 5.5 and GLM 5.2 both collapse into the low teens. On broad agentic evaluations, Fable 5 still holds the ceiling: GDPval v2 Elo at 1747 against Sol’s 1736 and K3’s 1686, JobBench at 57.4. When a task needs sustained judgment across many steps, Fable is still the front (Atoms 2026; Moonshot AI 2026a).
And to keep us grounded about the open side of the ledger: Artificial Analysis flagged that Kimi K3’s hallucination rate climbed from 39 to 51 percent even as its accuracy improved from 33 to 46 (The Decoder 2026). The composite peer still fabricates more than the closed frontier it’s chasing. Reliability is its own axis, and if your workload can’t tolerate a confident invention you’re paying the closed premium for a reason.
So the three composite points are real, and they live mostly in long-horizon agentic work, which is exactly the kind of work that production systems are starting to care about. The closed frontier still has a ceiling. But notice what kind of ceiling it is: a few points on the hard agentic end, at three to six times the price. Fable 5 charges $10/$50 per million tokens. Sol charges $5/$30. Kimi K3 charges $3/$15, with cached input at $0.30. On identical workloads Fable costs more than three times what K3 does, for a capability gap most tasks never surface (Moonshot AI 2026b; The Decoder 2026).
The moat that matters is a thin premium on judgment-heavy long-horizon work, sold at a premium price, revocable at someone else’s discretion. A year ago the closed labs were telling investors they had a data, compute, and talent wall. As of this month that wall is a three-point composite lead and a pricing argument.
Where the open-weights pillar stands now
I spent a lot of the last piece worried that the open-weights pillar would stop flowing at the frontier. That worry is a months-and-years question, and two weeks doesn’t move it. What I can do is show you where the pillar stands right now, because it moved a lot this month.
Kimi K3 launched July 16, and the full weights shipped July 27 under the Kimi K3 License, with vLLM, SGLang, and TokenSpeed support from day one. 2.8 trillion total parameters, 104 billion active, 16 of 896 experts per token, native image and video input, a million-token context window, open weight. That is the frontier tier, open. This month the open-weights pillar got stronger (Moonshot AI 2026b, 2026a).
Qwen3.8 Max is the messier data point. Alibaba previewed a 2.4-trillion-parameter multimodal model within days of the Kimi K3 release, made it callable through a paid preview endpoint on their AI Token Plan, and surfaced it in Qoder and QoderWork. As of writing, Alibaba has not published the weights, the license, the model card, the active-parameter count, or a single independent benchmark. The “second only to Fable 5” ranking is Alibaba’s own claim on Alibaba’s own evaluation. Bloomberg reports open weights are expected before August, but “expected” is not “shipped,” and self-hosting the thing at 4-bit takes roughly 1.2 terabytes of GPU memory, about eight H200s. So the conservative read is: Alibaba might open-source Qwen3.8, we don’t know when, and “preview” is not “release.” The pillar got stronger on the Kimi K3 side and stayed uncertain on the Qwen side (byteiota 2026).
And it’s worth knowing why two trillion-parameter-class open-weight models landed in the same week. Xi Jinping endorsed open-weight AI as national strategy on July 18, two days after Kimi K3 and one day before the Qwen3.8 preview (The Next Web 2026). The pillar got stronger because Beijing chose to strengthen it, which is also the reason it’s revocable. The same week, China is consulting on restricting overseas access to its most advanced models, open-weight releases included (Reuters 2026). Weights already downloaded are practically unenforceable to recall, but future releases are not. A pillar built on national strategy serves the nation that built it, and serves the rest of us only as long as serving us serves it.
So the picture is mixed. Some Chinese labs are pushing open weights all the way to the composite frontier (Moonshot, with Kimi K3). Some are fencing the frontier tier behind an API and keeping it there across generations (Alibaba: Qwen3.7-Max is proprietary per Alibaba’s own release (Alibaba Cloud / Qwen Team 2026), and Qwen3.8-Max landed as a paid preview with no published weights, license, or independent benchmarks). The same country can do both at once. The pillar moves unevenly, lab by lab and quarter by quarter, under a policy signal set in Beijing, and we still don’t get a vote in any of those choices.
The pillar we refuse to build
Last time I wrote “we need our own frontier labs”. That might be the understatement of the year, as not supporting frontier AI labs is a choice Europe keeps making, and the evidence for that is sitting in our own budget.
Europe already accepts strategic-scale spend when it considers a project important. CERN approved a 2026 budget of roughly €1.6 billion. ESA approved €8.26 billion. Together our leading particle-physics and space institutions spend about €9.9 billion a year, which is roughly what Epoch AI estimates Anthropic spent in 2025 ($9.7 billion) (Epoch AI 2026; Schuler 2026). One American AI company, one year, the budget of CERN plus ESA. We know how to spend at frontier-lab scale. We do it for particles and rockets. We don’t do it for the thing that is eating the rest of the economy.
We even have the legal machinery for preference. In May 2025 Europe adopted the SAFE defense instrument, €150 billion for defense procurement, with a rule that contractors must be based in the EU, EEA or Ukraine. The International Procurement Instrument already caps non-EU inputs at 50 percent in medical-device contracts above €5 million. We were willing to risk Washington’s displeasure over artillery (Schuler 2026). Our Tech Sovereignty Package stopped short of imposing a comparable preference for compute, largely to avoid the same confrontation over AI. That is a choice. We picked which strategic dependency we’d defend and which one we’d keep buying from abroad.
We have the compute too. We just spread it across three programs that embody three different theories of how to produce a frontier model, and we fund all of them without deciding which one is supposed to deliver. The Commission’s Frontier AI Grand Challenge went to the EUROPA consortium led by Domyn, with 2.5 percent of EuroHPC’s AI capacity for one year (European Commission 2026b, 2026a), and a cluster of 6,000 Nvidia Blackwell chips already in development (AI Weekly 2026). Germany’s SPRIND runs a separate €125 million tournament in rounds of €3 million, €8 million, and €15.5 million. EuroHPC spreads shared capacity across 19 AI Factories (Schuler 2026). Concentrate, narrow, or distribute. We financed all three and picked none. The one lab already at frontier scale, Mistral, is absent from each of them.
This is what I mean by choice rather than constraint. The money exists. The compute exists. The industrial base exists (ASML wrote the €1.3 billion check for an 11 percent Mistral stake in September 2025 (CNBC 2025), and Mistral is reported to be raising around €3 billion at close to a €20 billion valuation (Schuler 2026)). The legal machinery exists. The talent exists, though it keeps leaving for US labs because we won’t pay to keep it. What doesn’t exist is the decision to concentrate capital and compute behind frontier-scale runs and fund several labs to that standard. That’s a decision. We keep not making it.
Don’t put it all on Mistral
A single European frontier lab reproduces the fragility we already have with one US API, just domestically.
Washington runs multiple frontier labs. OpenAI, Anthropic, Google, X.ai (don’t get me started on that insanity), and the newer entrants. Beijing runs multiple. Alibaba, Zhipu, DeepSeek, Moonshot. Both sides field teams, plural, because redundancy is how you survive one lab going sideways (a bad release, a safety scandal, a government lean-on, a funding gap). Europe fielding one frontier lab and hoping nothing goes wrong with it is a gamble.
So the answer is to set a standard and fund every lab that clears it. The standard is the part we keep refusing to write: a frontier-scale training run, committed multi-year compute and power at frontier scale, conditioned on a public frontier-scale performance disclosure and an open-weight release under an EU-permissive license. Hit the standard, get the compute and the procurement preference, for years.
There’s a significant distinction that matters and that I think Brussels keeps blurring. The 19 AI Factories and the three Commission programs are shared compute and funding mechanisms, useful for SMEs and researchers, and none of them is a frontier-scale training run. Spreading compute across nineteen sites is a subsidy program for the broad ecosystem. Funding three competitions on conflicting theories is a procurement process that dodges the actual decision. The frontier needs a handful of actual labs, plural, each with committed multi-year compute and power at frontier scale, held to the frontier-scale standard I described above. Plural, because a single frontier lab is too fragile.
Mistral is one candidate, and at 0.5 composite points per month on the BenchLM trajectory I tracked in the last post, it’s a candidate that needs the compute to fix its rate. (Moonshot’s Kimi K2.6 to K3 jump is what that capital buys: GDPval v2 Elo went from 1190 to 1686 in a single generation (The Decoder 2026). Speed is a function of compute and capital.) But there should be a sovereign-backed lab alongside it, and room for one or two private challengers, each with multi-year compute and power reserved, each held to the same frontier-scale standard.
Arthur Mensch asked the French National Assembly in May for European-preference procurement, priority electricity, and regulatory proportionality (Schuler 2026). Those are the right asks, and they’re asks about state power. Private capital is moving on its own (ASML, the reported €3 billion round) (CNBC 2025; Schuler 2026); the bottleneck is the state meeting it with compute, power, and a multi-year horizon, at frontier scale, for several labs at once.
The grace period is for building
The open-weights pillar got stronger this month. Kimi K3 is a composite peer with the closed frontier, open weight, at a third of the price. Qwen3.8 is uncertain but pointing the same direction. If you wanted to read this as “the sovereignty panic was overblown, the open weights will keep us fine,” you could. You’d be wrong.
The grace period the open weights just bought us is the window to build. Right now, today, the deal is remarkably good: near-frontier capability, open weights, self-hostable, at a fraction of the closed price. That deal is a strategy choice by Moonshot and a permission choice by Beijing, and either can change without consulting us. The Reuters reporting from early July (Beijing consulting on tiered restrictions that could cover open-weight models too, future models first) is still live (Reuters 2026). The frontier tier landing behind a paid preview first, weights later, is still the pattern some labs will choose (Alibaba’s Qwen3.8-Max today: callable, no weights, no license, no independent benchmarks (byteiota 2026)). The closed pillar can still be turned off by a US export-control rewrite.
If the open weights keep flowing, Europe has breathing room. If they stop, and the closed side restricts at the same time, we are exactly where I said we were in the last post, with nothing of our own to lean on. The difference is that the breathing room is now, and the breathing room is the only chance we get to build the pillar. Spending it on complacency is the choice I’m asking us to stop making.
Conclusion
The moat rotted faster than I said. Three composite points, open weight, a third of the price. The closed-lab story that only US labs can do frontier is dead as of this month, killed by a Chinese lab shipping open weights to the top of the board.
The choice didn’t rot at all. Europe can compete. We have the money, the compute, the industrial base, the legal machinery, the talent. We keep choosing not to, by fragmenting the funding across three programs that can’t agree on a theory, by designing our AI policy to avoid a confrontation we already accepted over artillery, by fielding one slow lab and calling it a strategy.
How we fix it: concentrate the capital and the compute, distribute the funded labs, set a frontier-scale standard, fund several labs that clear it with multi-year compute and power and procurement preference, conditioned on open-weight release under an EU license. A pillar is a stack of labs, each backed to frontier scale.
The open weights just bought us a grace period. The only sane use of a grace period is to build the thing you were going to need anyway, before the grace period ends. We have the money. We have the compute. We have the legal tools. We are choosing not to. That’s the part that should keep us up at night, and it’s the part that’s entirely in our hands.





