AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 Is A Notable AI Choice Beyond The US And China, With Caveats on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a sharp improvement over Mistral’s previous models and a result that places it among the cited leaders outside the United States and China. The same comparison shows leading US and Chinese models scoring higher, while two Chinese models cost less per benchmark task; the model’s weights and licence were not yet available in the source material.

Mistral has released Large 4, a French-developed AI model that scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The result makes it the highest-scoring model in the source report’s cited field outside the United States and China, but the same data places it below leading US and Chinese systems and shows two Chinese models completing benchmark tasks at a lower cost.

Artificial Analysis lists Large 4 at 38.4 points. That is a sharp rise from the same index version’s scores of 9 for Mistral Large 3 and 14 for Medium 3.5. The source report describes the gain as a major advance for a European lab, while cautioning that it does not put Mistral among the highest-scoring frontier models.

On the cited leaderboard, Anthropic’s Claude Opus 5.5 scores 57.6, and other leading US models score above 50. Chinese models GLM-5.3, Kimi K3, GLM-5.3-Flash and DeepSeek V4.1 Flash also exceed Large 4, with scores from 39.5 to 44.8. The source says Large 4 would rank eighth among open-weight models once its weights are released, but that ranking is prospective: at the time described, the weights had not shipped.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, image and text input, text output, and a 512,000-token context window. It was available through Mistral’s API as a Research Preview. The stated standard price was $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; the source reports a 50% introductory discount for the first two weeks.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a Research Preview, with benchmark data showing a substantial jump for the French lab but gaps in score and benchmark-task cost against several competitors.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A French Model With Clear Trade-Offs

Large 4 matters as evidence that Mistral has made substantial progress, and as a potential option for organisations seeking an AI provider based outside the United States and China. But the source’s comparison suggests buyers should distinguish that geographic choice from technical leadership: the model’s benchmark score trails several major competitors, including some open models.

Price comparisons also complicate the case. The source calculates a cost of $1.13 per Artificial Analysis Index task for Large 4, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models score 41.8 and 39.5 respectively. These figures are benchmark-specific, not a universal estimate of operating cost, but they indicate that Large 4 may be difficult to justify on price and score alone for workloads resembling the tests.

The report also says Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. If that difference carries over to a buyer’s tasks, longer outputs could add cost and latency. Real-world results will depend on task design, prompt length, model settings and the amount of human review.

Amazon

AI language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Benchmark Compares

The source bases its rankings on Artificial Analysis Intelligence Index v4.3.2, which includes tests of agentic work, coding and workflow tasks, rather than only short-form question answering. The report names AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0 among the index’s components. Its scores therefore reflect performance on a particular mix of evaluations; they do not establish how every model will perform for every user.

Large 4 is offered as a Research Preview, and Mistral says reinforcement learning is still under way, according to the source. That means its performance may change. The source also reports that Mistral promised to release model weights by the end of October, but it does not provide a year or confirm a completed release. Until weights and licence terms are available, the model’s status as an open-weight option cannot be treated as established.

Amazon

large language model for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Real-World Results

Several details were unresolved in the source material. Large 4’s weights had not yet been released, and its licence was described as unpublished. The promised end-of-October release was a future commitment in the report, not confirmation that it occurred. The source also does not specify the year for that deadline, so its current status cannot be established from the material provided.

The source author says hands-on tests found instances of confident hallucination, but this is a personal observation, not an Artificial Analysis benchmark result, and no test protocol or sample size is supplied. The reported 200-million-token benchmark output figure likewise does not by itself show how much a typical customer task would cost. Independent testing and production use would be needed to establish reliability, latency and cost across different workloads.

Finally, the leaderboard is a snapshot from a named index version. Mistral says training work continues, and model versions, prices and competitor scores can change. The provided comparison does not establish how Large 4 performs on tasks outside the index or under a buyer’s own evaluation.

Amazon

text input output AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Updated Scores

The next reported milestones are the release of Large 4’s model weights and licence terms, and any updated scores after Mistral’s ongoing reinforcement-learning work. Those details will clarify whether developers can run or adapt the model independently and under what conditions.

For prospective users, the practical next step is to test the preview against their own workloads, measuring accuracy, hallucination rates, output length, latency and total cost. Until such results and the promised release are confirmed, Large 4 is best understood as a significant Mistral benchmark improvement—not proof that it matches the leading US systems or offers better value than lower-cost alternatives.

Amazon

AI model token input output

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source report.

Is Mistral Large 4 the top-scoring AI model?

No. It leads the cited field outside the United States and China, but the listed US leaders score above 50, and several Chinese models also score higher.

Is Large 4 open source or open-weight now?

The source describes it as a proprietary Research Preview available through Mistral’s API. Weights were promised for a later release, while licence terms were unpublished in the report.

How does its benchmark-task cost compare?

The source estimates $1.13 per index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. These are benchmark-specific estimates, not guaranteed costs for other workloads.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bringing ChatGPT For Teachers To More U.S. School Districts

More U.S. school districts are adopting ChatGPT to support teachers, with pilot programs showing promising results in classroom integration and workload reduction.

The Growing Concern Over Grok’s Use Of Victims’ Media For AI Purposes

Survivors claim xAI’s Grok chatbot was trained on their images and videos without consent, raising legal and ethical concerns about AI data sourcing.

MiniMax H3: How Sound Is Integrated And The Meaning Of ‘Open’ Access

MiniMax H3 launches with integrated sound and claims of ‘open’ weights, but key details and limitations remain unclear. Here’s what is confirmed and what is not.

Elon Musk’s X Corp And SpaceXAI Just Moved To Dismiss Their Lawsuit Against Apple

X Corp and SpaceXAI have filed motions to dismiss their lawsuit against Apple, marking a significant legal development involving Musk’s companies and the tech giant.