🔍 Read the full analysis: A Practical Look At Mistral Large 4 Outside The US And China on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4, released as a research public preview, scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The score marks a sharp improvement over Mistral’s prior models, but the supplied analysis says leading US models and several Chinese models score higher, while cheaper alternatives perform well on the same benchmark. Its weights, licence and post-preview performance remain unsettled.
Mistral released Large 4 as a research public preview, with the model scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result is a steep improvement on the company’s previous models, but the supplied benchmark comparison places several US and Chinese models ahead, raising questions for buyers weighing performance, cost and reliability.
Large 4 is a 1-trillion-parameter model with 49 billion active parameters, according to the source material. It accepts text and images, produces text, and has a 512,000-token context window. Mistral is offering it through its API as a research public preview. The company says it plans to release the model weights at the end of October; until then, the model is proprietary, and its licence has not been published.
At standard API rates, the listed price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. The source says Mistral is offering a 50% discount for the first two weeks. Mistral also says reinforcement learning is still underway, meaning results could change as development continues.
On the same Artificial Analysis index version, Large 4 scored 38.4, compared with 9 for Mistral Large 3 and 14 for Medium 3.5. The supplied analysis describes the jump as a major gain, while noting that US leaders score above 52 and Chinese models in the comparison reach 44.8. These are benchmark results, not a guarantee of how a model will perform in any particular company’s workload.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Cost of the Capability Gap
The score matters because the Artificial Analysis index used here includes agentic and work-oriented tests, such as knowledge work, SaaS workflows and coding tasks. Those results may help buyers assess models intended to take multiple steps or operate tools, rather than only answer short prompts. Still, a composite benchmark cannot by itself predict performance in a specific production workflow.
The source’s comparison raises a procurement question: Large 4 costs $1.13 per Intelligence Index task, while the cited figures put GLM-5.3-Flash at $0.25 and DeepSeek V4.1 Flash at $0.27. Those two models score 41.8 and 39.5 respectively, above Large 4’s 38.4. The figures are from the supplied analysis; buyers would need to check current rates and test comparable tasks before treating them as a direct purchasing comparison.
The author also reports that Large 4 generated 200 million output tokens while completing the index, compared with a median of 81 million for comparable models. If that difference holds in a buyer’s use case, longer outputs could add cost and latency even when the per-token price appears manageable. The source does not provide enough detail to generalize that observation to all tasks.
As an affiliate, we earn on qualifying purchases.
A Sharp Rise, Not a Ranking Lead
The comparison is based on Artificial Analysis Intelligence Index v4.3.2, identified in the source as the current version. It lists US models such as Claude Opus 5.5 at 57.6 and Gemini 4 Argon at 52.6. Chinese models in the supplied table include GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. Large 4’s 38.4 places it below those named systems, but above DeepSeek V4 Pro 0813 at 36.0 and GLM-5.2 at 33.7.
The source frames Large 4 as the strongest model from outside the United States and China, while arguing that this distinction reflects a narrow competitive field. That is the author’s characterization, not a finding established by the score table alone. The supplied material does not present a full comparison across every model or region.
For Mistral, the change from 9 on Large 3 to 38.4 on Large 4 suggests a substantial benchmark improvement. But a higher score than the company’s earlier releases does not place it at the current frontier described in the same table. The model’s position depends on the comparison set: it is a large step forward for Mistral and still trails the listed leading US and Chinese systems.
“Reinforcement learning is still running.”
— Mistral
large language model API subscription
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights, Licence and Reliability
Several points remain unresolved. The weights have not yet been released, and the source says Mistral has not published the licence that would govern their use. The company’s stated end-of-October release plan could also change; the supplied material does not confirm a date beyond that commitment.
The source author reports seeing confident hallucinations in hands-on use, but provides no test protocol, sample size or independently verified rate for Large 4. The article distinguishes this observation from Artificial Analysis benchmark data. The supplied material also does not give a complete account of how the index’s task scores relate to performance across different business applications.
Pricing and benchmark comparisons can shift as providers update models, rates and evaluation methods. Mistral says reinforcement learning is ongoing, so the present score should be treated as a snapshot of a preview model rather than a settled evaluation of its eventual release.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the October Weights Release
The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. That release, if it proceeds as described, should clarify the model’s licence and allow developers to assess access and deployment options beyond the current API preview.
Before adopting Large 4 for multi-step work, buyers can compare it with alternatives on their own tasks, including output length, error rates, tool use, latency and total cost. Further benchmark results may also emerge as Mistral continues reinforcement learning and the preview develops. The source material does not confirm when an updated independent evaluation will be available.
text and image AI processing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is a 1-trillion-parameter model with 49 billion active parameters. It accepts text and images, produces text and has a 512,000-token context window, according to the supplied source.
How did Large 4 score?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The source lists several US and Chinese models with higher scores and says Large 4 improved on Mistral Large 3’s score of 9 on the same index version.
How much does the API cost?
The listed standard rates are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source says a 50% discount applies for the first two weeks.
Are the model weights available?
Not yet, according to the source. Mistral plans to release the weights at the end of October; until then, Large 4 is available as a proprietary API research preview. The licence has not been published in the supplied material.
Is Large 4 ready for agentic workflows?
The supplied analysis raises concerns about its benchmark position, output volume and reported hallucinations, but does not establish that it is unsuitable for every agentic task. Buyers should test it on their own workflows and compare accuracy, cost and latency with alternatives.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
