Skip to content
Guideadvanced

Mistral Large 4 vs Aleph Alpha Kolibri: Which Open-Weight Model Fits an EU AI Act GPAI Stack

Published October 9, 2026 · by Pondero Platform

The short version

Kolibri's weights are downloadable today; Mistral Large 4 scores 38 on the Artificial Analysis index but stays API-only until end of October. Both labs signed the GPAI Code of Practice. Here is which one to run, per org profile, and what to do before October 31.

Table of Contents

Mistral Large 4 vs Aleph Alpha Kolibri: Which Open-Weight Model Fits an EU AI Act GPAI Stack

If your platform team has to self-host a European model this month, Kolibri is the only one of the two you can download. Mistral Large 4 has the stronger third-party evidence, a 38 on the Artificial Analysis Intelligence Index where Kolibri has no score yet (Artificial Analysis), but its weights are promised for "end of this month" and today it is API-only (Mistral). On compliance, the gap is smaller than the marketing suggests: Aleph Alpha and Mistral AI are both on the European Commission's list of GPAI Code of Practice signatories (European Commission). Decide on three things: whether you need weights this month, what GPUs you have, and whether you need more than English and German.

AxisAleph Alpha KolibriMistral Large 4 (preview)Source
ReleasedOctober 3, 2026October 6, 2026, Research Public PreviewAleph Alpha, Mistral
Total / active parameters78.1B / 3.46B1T / 49B per Artificial Analysis; 1.05T / 52B per Mistral's docsAleph Alpha, Artificial Analysis, Mistral docs
Context window262,144 native; 1,048,576 by extrapolation524k per Artificial Analysis; 1M per Mistral's docsKolibri model card, AA model page, Mistral docs
Weights todayYes, Hugging Face, Apache 2.0No; promised by end of OctoberKolibri model card, Mistral
Minimum self-host hardwareAbout 78 GB FP8; e.g. 2x H100 SXM5 or 1x B200Not publishedKolibri model card
AA Intelligence IndexNo third-party score at launch38Tech Jacks Solutions, Artificial Analysis
API price per 1M tokens (in / out)None; self-host$1.36 / $4.18, 50% off for the first two weeksArtificial Analysis
LanguagesEnglish and German only160+ languages, all official EU languagesKolibri model card, Mistral
GPAI Code of Practice signatoryYesYesEuropean Commission

Kolibri: small enough to run on hardware you already own

Kolibri's case is operational. The FP8 weights need about 78 GB, and Aleph Alpha's minimum configurations include two H100 SXM5s or a single B200 (model card). Aleph Alpha says it picked 78B over a 123B candidate because, on two H100s, the 78B model serves 18 concurrent 256k-token requests where the 123B served three (Aleph Alpha). With a GPU allocation already approved, you can schedule that deployment this sprint.

Aleph Alpha also trained Kolibri with its Merlin-Arthur protocol to say "I don't know" when the context lacks the answer. On AA-Omniscience it abstains instead of answering wrong on 44% of items, versus 15% for its predecessor, per Aleph Alpha's own harness (Aleph Alpha). All of those numbers are self-reported; Tech Jacks Solutions noted on launch day that Epoch AI and LMSYS evaluations were still pending (Tech Jacks Solutions).

What breaks first: the 1M context. Kolibri was trained up to 262,144 tokens; beyond that it extrapolates, needs --max-model-len and --hf-overrides flags to serve, and the card recommends staying at or under 262,144 for complex tasks (model card). Plan your RAG chunking around 256k, not 1M. And if your users write in French, Polish, or Italian, stop here: the card calls the two-language scope "a deliberate choice of depth over breadth."

Mistral Large 4: frontier-adjacent, but you are renting it until the weights land

Mistral Large 4 scores 38 on the Artificial Analysis Intelligence Index, level with GPT-6 Luna (max, 38) and one point behind DeepSeek V4.1 Flash (max, 39) (Artificial Analysis). Mistral Large 3 scored 9 on the same index, and Simon Willison puts the new model "about 6 months behind the frontier" (Simon Willison). It also takes image input; Kolibri's Hugging Face card lists it as a text-generation model (model card).

Cost is where it loses. Artificial Analysis prices a single Intelligence Index task at $1.13, or $0.57 during the launch discount, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash at similar scores (Artificial Analysis). Part of that is verbosity: the model generated 200M tokens running the index, against a median of 81M (AA model page).

What breaks first: the spec sheet. Artificial Analysis and Simon Willison report 49B active parameters; Mistral's launch post and docs say 52B. Artificial Analysis lists a 524k context; Mistral's docs page shows 1M (Mistral docs, AA model page). Do not write either into a capacity plan until the weights ship with a model card. Hardware is the other gap. Large 4 carries roughly 13x Kolibri's total parameters, and an MoE model keeps every expert in memory (a tradeoff Kolibri's own model card spells out), so expect a far larger node. Mistral has not published requirements.

What your compliance team will ask

For a deployer, GPAI compliance mostly sits with the provider: Articles 53 and 55 obligations are theirs, and both providers here signed the Code, whose Transparency and Copyright chapters cover Article 53 (European Commission). Your Article 50 duties do not change with the model; our GPAI obligations map covers them.

The difference is documentation you can read today. Kolibri's card publishes training compute (392k GPU-hours of pre-training on 768 B200s), the data mix, PII redaction, and the Code signature (model card). Mistral's launch post does not address the Code. For a model at 1T parameters, ask Mistral directly whether it falls under Article 55: Article 51 presumes systemic risk above 10^25 training FLOPs (EU AI Act, Art. 51). Treat Aleph Alpha's "compliance by design" pitch as Aleph Alpha's own claim, pending third-party evaluation.

Verdict per org profile

Your situationPickFlip condition
German-language public sector or regulated industry, on-prem, production this quarterKolibriThird-party evals come back well below Aleph Alpha's numbers
Multilingual EU workloads or image input, can wait three weeksMistral Large 4: pilot via API now, self-host after weightsWeights slip past October, or the license is not permissive
EU-hosted is enough, self-hosted is not requiredMistral Large 4 API, served from Mistral's own European datacenters (Mistral)Your DPA review requires on-prem
Cost per task decides and model origin does notNeither: GLM-5.3-Flash or DeepSeek V4.1 Flash at about a quarter of the cost per task (Artificial Analysis)Procurement adds an EU-origin requirement

What to do before October 31

  1. Stand Kolibri up on two H100s with the aleph-alpha-inference vLLM plugin, capped at 262,144 tokens, and run your own eval set, including questions the context cannot answer.
  2. Run the same eval set against the Mistral Large 4 preview API on non-sensitive data while the 50% launch discount lasts; it covers the first two weeks after October 6 (Artificial Analysis).
  3. Ask both vendors for their completed Model Documentation Form, the template in the Code's Transparency chapter (European Commission), and ask Mistral for its Article 55 position.
  4. When Mistral's weights drop, read three things before anything else: the license, the active-parameter count, and the hardware floor. If they have not shipped by October 31, make Kolibri the default for German and English workloads and put multilingual teams on the EU-hosted API until they do.