← All articles

Mistral Large 4 vs Kimi K3, GLM-5.3 and Claude Opus 5

2026-10-10 — Michael Leung

Part 13 of 13 · AI Models & Agents 2026

Useful already? ☕ Buy me a coffee and keep these articles free and ad-free.

Mistral Large 4: a giant glowing crystal brain floating above a European city at dusk

Mistral Large 4 is a frontier-scale AI model that comes from neither the US nor China, and that is rare. Nearly every model this big in 2026 is made by an American or Chinese lab. Mistral, from France, released it as a public preview on 6 October 2026: 1 trillion parameters, open weights promised for late October, and a price of $1.36 per million input tokens and $4.18 per million output tokens. That is slightly cheaper than GLM-5.3 and far cheaper than Kimi K3 or Claude Opus 5. In a blind code-quality test that Mistral published, it narrowly beat GLM-5.3 and Kimi K3, but Claude Opus 5 was still clearly ahead. If your organisation needs a strong model that is not tied to the US or China, this is now the main option. If you just want the best code, Claude is still in front.

This makes three trillion-scale open-weight releases in three months: Kimi K3 in July and GLM-5.3 in August, both from China, and now Mistral Large 4 from Europe. In August I wrote about why you shouldn't lock into an annual AI subscription, and this release is one more reason. Below I put Mistral Large 4 next to the three models people are most likely to compare it with, using each maker's own published numbers.

What Is Mistral Large 4?

It is Mistral's new flagship, nicknamed "le Chonk". It is a mixture-of-experts model: it has 1 trillion parameters in total, but only 52 billion are used for each token. That is how a model this big can still be cheap to run per request. It can take images as well as text, and it can answer quickly or switch into a slower reasoning mode.

Mistral has not published the context window in the announcement. It says more details on the architecture and training will come with the open weights.

Why Does It Matter That Mistral Is Not From the US or China?

Look at who makes the biggest AI models in 2026. The closed frontier models, such as Claude, GPT and Gemini, come from US companies. The biggest open-weight models, such as Kimi, GLM, DeepSeek and Qwen, come from Chinese labs. Very few companies anywhere else have trained a model at this scale, and Mistral is the best-known of them. Mistral says Large 4 leads open-weight models developed outside China "by a wide margin" on cybersecurity tests.

A datacentre built inside an old European stone building, standing for Mistral

For many organisations, where a model comes from matters as much as how good it is. A bank, a hospital or a government department may not be allowed to send data to a Chinese provider. It may also not want to depend on a US company that could change its prices, terms or availability overnight. Until now, those organisations had to accept a weaker model or pick one of the two sides. Mistral's pitch is a third option: "Forged in Europe. Built for AI sovereignty."

Mistral backs this up with its own infrastructure. It trained the model from scratch on its own GPUs in European datacentres, serves the preview from the same place, and offers a European deployment that it runs end to end. Once the weights are released, organisations can also run it on their own servers and avoid any outside provider at all. For Australian developers like me, who mostly use American or Chinese models today, more choice is good news too.

How Does It Compare With Kimi K3, GLM-5.3 and Claude Opus 5?

On size, Mistral Large 4 sits in the middle of the open models. Kimi K3 is almost three times bigger, and GLM-5.3 is a little smaller. Anthropic does not publish the size of Claude Opus 5.

Four yachts racing at sunset, a picture of four frontier AI models competing
ModelMakerReleasedTotal / active parametersWeights
Mistral Large 4Mistral AI (France)6 Oct 2026 (preview)1T / 52BPromised for late October
GLM-5.3Z.ai (China)Aug 2026744B / 40BOpen, custom licence
Kimi K3Moonshot AI (China)Jul 20262.8T / 104BOpen, custom licence
Claude Opus 5Anthropic (US)24 Jul 2026Not publishedClosed, API only

The only test that puts all four side by side is a blind human review of code quality, run by Surge AI and published on Mistral's launch page. Reviewers rated code from each model on a 1 to 5 scale without knowing which model wrote it.

Blind code quality ratings: Claude Opus 5 4.22, Mistral Large 4 3.74, GLM-5.3 3.60, Kimi K3 3.59

Claude Opus 5 scored 4.22. Mistral Large 4 came second with 3.74, just ahead of GLM-5.3 (3.60) and Kimi K3 (3.59). Keep in mind that Mistral picked this test and published it. The gap between the three open models is small, while the gap to Opus 5 is large.

Mistral and Z.ai both report two of the same agent tests, although each company ran its own:

So GLM-5.3 is still better at fixing bugs, and Mistral Large 4 is better at automation work. Neither is a clear winner overall.

Is Mistral Large 4 Cheaper?

Yes, but only by a little compared with GLM-5.3. Compared with Kimi K3 and Claude Opus 5, the difference is large.

API price per million tokens: Mistral Large 4, GLM-5.3, Kimi K3 and Claude Opus 5
Price per 1M tokens (US$)Mistral Large 4GLM-5.3Kimi K3Claude Opus 5
Input$1.36$1.40$3$5
Output$4.18$4.40$15$25
Cost of the example job below$8.89$9.20$22.50$37.50

To make the prices concrete, I used the same example as in my Claude family comparison: summarising 1,000 documents, with 5,000 tokens in and 500 tokens out for each one.

Cost of one summarising job: Mistral Large 4 .89, GLM-5.3 .20, Kimi K3 .50, Claude Opus 5 .50

Two things to note. First, these are the makers' own API prices. Third-party hosts on OpenRouter often charge less for the open models. Second, Claude Opus 5 is no longer Anthropic's newest Opus. Opus 5.5 replaced it in September at $4 in and $20 out, which would make the same job cost $30.

What About Cybersecurity and Safety?

This is the part of the announcement I found most surprising. Mistral says Large 4 is one of the top five models in the world on the Artificial Analysis Cyber Index, and well ahead of every other open model built outside China.

A steel shield fused with a circuit board next to a padlock, standing for AI cybersecurity

This cuts both ways. A model that will reproduce vulnerabilities is useful for security teams, and it is also useful for attackers. Anthropic's models refuse these tasks on purpose. Which behaviour you want depends on who you are, and once the weights are public, anyone can use it.

Can I Run Mistral Large 4 at Home?

No, not on a normal PC. I run local models on an RTX 4060 Ti with 8 GB of video memory, and a 1 trillion parameter model is far beyond that. Even squeezed down to 4 bits, the weights alone would need roughly 500 GB of memory. For comparison, GLM-5.3, which is smaller, officially needs a server with eight 141 GB GPUs.

The open weights matter for companies with their own GPU servers, and for hosting providers who will offer it more cheaply. For the rest of us, it is an API model. If you want something that really runs on a home graphics card, see my guide to local LLMs versus cloud APIs.

Which One Should You Use?

My plan is to try Mistral Large 4 on my own website builds when the preview has settled, and compare it with Kimi K3 on the same tasks. Benchmarks published by the company that made the model are a starting point, not the answer.

Related reading: Claude models compared: Haiku, Sonnet, Opus, Fable · Why you should avoid annual AI subscriptions in 2026 · Local LLM vs cloud API in 2026

Sources

Figures checked on 10 October 2026. Mistral Large 4 is a public preview, so its price and details may change before the open weights are released. The DeepSWE and AutomationBench scores come from two different companies' own tests and may not have used identical settings. The cost example is my own arithmetic at standard rates and leaves out reasoning tokens. This article was researched and drafted with AI assistance and reviewed by me. The images were made locally with Qwen Image 2.1 on an 8 GB RTX 4060 Ti, and the charts were drawn from the published numbers.

Series: AI Models & Agents 2026

New frontier models, coding agents and what they cost, tested by a working developer.

  1. Why You Should Avoid Annual AI Subscriptions in 2026: Lessons from GLM-5.2 and Kimi K3
  2. One Night of Agentic Coding for Under $1: Meet GLM-5.3-Flash
  3. Why Claude's "Agentic Intelligence" Beats All-in-One AI Platforms
  4. Anthropic Introduces Claude Sonnet 5.5: Faster, Smarter, and More Efficient
  5. What Is Space Bunny Alpha? OpenRouter's Free 1M-Context Stealth Model
  6. The Era of AI Agents: From Chatbots to Autonomous Executors
  7. The Next Era of Frontier Intelligence: Why Gemini 4 Argon Changes Everything
  8. Beyond Reports: Meet Apodex 1.1 and the New Era of Agentic Workbenches
  9. Meet FrogNano: Microsoft's Compact Coding Agent Powered by Reinforcement Learning
  10. When AI Goes Rogue: OpenAI's Medicare Breach Hearing
  11. Claude Haiku 5.5: Cheaper, Faster, and Where It Fits
  12. Claude Models Compared: Haiku, Sonnet, Opus, Fable Prices
  13. Mistral Large 4 vs Kimi K3, GLM-5.3 and Claude Opus 5
Start the series from part 1 →

← All articles