🔍 Read the full analysis: Mistral Large 4 Still Trails In The Push Toward The AI Frontier on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released an API preview of Large 4 on October 6, 2026, but Artificial Analysis rated it at 38 on its Intelligence Index, below several leading U.S. and Chinese models. The benchmark is a dated snapshot, not a direct test of every workload; the model’s weights are scheduled for release later in October.
Mistral AI released Mistral Large 4 in public API preview on October 6, but the model scored 38 on Artificial Analysis’s Intelligence Index, trailing leading U.S. models and several Chinese competitors in the October 7 snapshot. The result raises questions about its readiness for demanding agentic tasks, while the release remains a preview rather than the scheduled public release of its weights.
Artificial Analysis’s index places Mistral Large 4 Preview level with OpenAI’s GPT-6 Luna at maximum reasoning effort and one point below DeepSeek V4.1 Flash at maximum effort. The comparison lists Claude Opus 5.5 at 58, Gemini 4 Argon at 53, and GPT-6.1 Sol at 52. China’s GLM-5.3 scored 45 and Kimi K3 scored 44. These are index points, not percentages or direct predictions of success on an individual task.
Mistral described Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters, accepting text and images. Its preview is accessible through an API. The company says it trained the model on its own European infrastructure and is continuing to improve it. The weights are scheduled to be released later in October; they were not downloadable when the source report was published.
The source report’s author, Thorsten Meyer, said his own use of the preview included hallucinations and that he would not select it for demanding long-running agentic work when stronger alternatives are available. He identified that as a personal assessment, not the result of a controlled comparison. Artificial Analysis’s aggregate score offers one benchmark signal, but does not establish how the model will perform on a particular coding, research, or business workflow.
AI FRONTIER REPORT · OCTOBER 7, 2026
Mistral Large 4 Still Trails In The Push Toward The AI Frontier
Mistral’s API preview scores 38 on Artificial Analysis’s Intelligence Index, behind several leading U.S. and Chinese models. It is a dated benchmark snapshot; the weights are scheduled for release later in October.
Weights are pending.
Announced October 6, 2026. The score reported the next day offers a signal for comparison, not a verdict on every workload.
Artificial Analysis snapshot
Mixture-of-experts model
Reported by Artificial Analysis
Scheduled later this month
01 / BENCHMARK POSITION
A visible gap at the top
Large 4 Preview sits below several leading models in the October 7 index snapshot. These are index points, not percentages or predictions of success on an individual task.
Reasoning settings varied across models, and evaluations did not use identical compute budgets. The ranking is a snapshot from October 7, 2026, not a permanent leaderboard.
Large 4 is 14 to 20 index points behind the three listed leaders. Index points do not translate directly into a percentage gap or expected task success.
Cohere Command A+ scored 13 in the same table. The comparison shows variation across the field; it does not mean every competitor leads Mistral.
02 / MODEL PROFILE
Scale, access and timing
Mistral describes Large 4 as its largest model to date. The preview accepts text and images and is available through an API.
One trillion total parameters, with 49 billion active parameters.
Multimodal input is available in the API preview.
Introduced October 6. Weights were not downloadable in the reported snapshot.
Company statement: Mistral says the model was trained on its own infrastructure in Europe and that it is continuing to improve it. Weights are scheduled for release later in October; no specific date was given in the source report.
03 / WHAT A SCORE CAN’T TELL YOU
Benchmarks are a starting point
An aggregate index cannot settle how a model will behave on a particular coding, research or business workflow.
A score of 38 is not proof Large 4 will fail a specific task, and it does not directly measure long-horizon agent reliability.
A reported capacity of roughly 512,000 tokens says how much material can fit in a request. It does not establish reliable reasoning across that material.
Thorsten Meyer reported hallucinations in his own use and said he would not choose the preview for demanding long-running agentic work. This was a personal assessment, not a controlled comparison.
The source describes DeepSeek V4.1 Flash as having comparable benchmark intelligence at a much lower measured cost per task, but provides no underlying figures to detail that comparison.
Company claims need workload tests. Mistral advertises strengths in agentic coding and specialized professional tasks. The reported material does not include detailed results substantiating those claims.
04 / NEXT STEPS
Test the work you actually do
The next stated milestone is the planned weight release later in October. New versions and independent evaluations will need their own assessment.
Choose tasks
Use representative coding, research and business work.
Run the preview
Measure tool use, constraint-following and verification.
Check the chain
Look for early errors that affect later actions.
Re-evaluate
Compare again after weight release or model updates.
AT A GLANCE / KEY QUESTIONS
Quick answers
What did Mistral release?
An API preview of Large 4 on October 6, 2026. It accepts text and images; weights were scheduled for later in October.
How did it score?
Artificial Analysis rated Large 4 Preview at 38 on its Intelligence Index in the October 7 snapshot. It is an aggregate score, not a percentage.
Does that prove it cannot handle agentic work?
No. The index does not establish failure on a specific workflow or directly measure reliability across long tasks.
Can users download the weights?
Not in the snapshot covered by the report. Mistral had scheduled their release for later in October without a specific date.
The Gap in Frontier Benchmarks
The results matter because developers choosing a model for multi-step work need more than a high parameter count or a large context window. Agentic systems plan, use tools, interpret outputs, and carry decisions forward. Errors early in that chain can shape later actions, while a fluent final response may not reveal that the work went off course.
On the reported index, Large 4 is 14 to 20 points behind the three listed leading U.S. models. That is a comparison in index points, not a percentage gap or a measure of expected task success. The author also noted that GLM-5.3 and Kimi K3 scored higher, while DeepSeek V4.1 Flash had approximately comparable benchmark intelligence at a much lower measured cost per task, according to the report. Those findings make the preview’s position relevant to developers weighing capability and cost, though workload-specific testing remains necessary.
The result does not mean every competitor leads Mistral: Cohere Command A+ scored 13 in the same table. It does show why progress within Mistral’s own product line is not proof of parity with the strongest available models. A roughly 512,000-token context capacity, reported by Artificial Analysis, indicates how much material can fit into a request; it does not by itself establish reliable reasoning over that material.
As an affiliate, we earn on qualifying purchases.
A Preview Before Weight Release
The timing limits what can be concluded. Mistral introduced an API preview on October 6, and Artificial Analysis’s reported scores were a snapshot available on October 7. Mistral said it was continuing to improve the model, so future versions or subsequent evaluations could differ. At the time covered by the source, users could access the preview through the API, but not download its weights.
The comparison table records developer locations, not where individual API requests are processed. It also includes models evaluated at different reasoning settings, such as maximum effort, high, or default fallback. The source cautions that these are not evaluations under identical compute budgets. The scores therefore help show relative standing within that published snapshot, but should not be treated as a perfectly controlled contest or a permanent ranking.
Mistral’s stated European training infrastructure is relevant to the company’s role in developing AI capacity in Europe. That fact is separate from the question of how well the preview performs. The source report says Mistral advertises strengths in agentic coding and specialized professional tasks; those are company claims that require testing on the relevant workloads rather than inference from the benchmark score alone.
“The model was trained on the company’s own infrastructure in Europe, and Mistral continues to improve it.”
— Mistral AI, as described in its announcement
As an affiliate, we earn on qualifying purchases.
What the Scores Cannot Establish
The index does not answer whether Large 4 will perform well on a specific customer’s workload, and the reported reasoning settings were not evaluated with identical compute budgets. The score of 38 is a dated aggregate benchmark result, not proof that the model will fail a particular task or a direct measure of long-horizon agent reliability.
Meyer’s comments about hallucinations reflect his own use of the preview; the source provides no controlled comparative hallucination study. The frequency of unsupported output across different models and tasks is not established here. The report also refers to a cost comparison with DeepSeek V4.1 Flash but the supplied material ends before giving the underlying figures, so the amount and measurement basis of that cost difference cannot be independently detailed from this material.
It remains unclear how the model will score after further updates, how it will perform under independent workload-specific testing, and whether the planned weight release will arrive on schedule. Mistral’s claims about agentic coding and professional tasks have not been substantiated by detailed results in the source material.
As an affiliate, we earn on qualifying purchases.
Weight Release and Workload Tests
The next stated milestone is the release of Mistral Large 4’s weights later in October, according to the company’s announcement as reported in the source. Until that release, the immediate product being assessed is an API preview. The source does not give a specific release date beyond “later in October.”
For developers, the practical next step is to test the preview, and any later weight release, against their own representative tasks: tool use, coding, research, constraint-following, and verification. New independent benchmark results could alter the comparison, but the October 7 index should be read as a snapshot rather than a final verdict. Mistral’s stated plans to keep improving the model also mean later versions will need separate evaluation.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral release?
Mistral introduced an API preview of Large 4 on October 6, 2026. The source says the model accepts text and images and that its weights were scheduled for release later in October.
How did Mistral Large 4 score?
Artificial Analysis gave Large 4 Preview a score of 38 on its Intelligence Index in the snapshot reported on October 7, 2026. This is an aggregate benchmark score, not a percentage or a guarantee of performance on a particular task.
Does the score prove Large 4 cannot handle agentic work?
No. The index does not establish that the model will fail a specific workflow or directly measure reliability across long tasks. The source report’s recommendation against using it for demanding agentic work is the author’s judgment, based on the benchmark and personal experience.
Are Large 4’s weights available to download?
Not at the time covered by the October 7 report. Mistral said the weights were scheduled for release later in October, but the source does not provide a precise date.
Why is the benchmark comparison not definitive?
The table combines models evaluated at different reasoning settings, and the source says the comparisons were not made under identical compute budgets. Scores may also change as models are updated or new evaluations are published.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
