Aleph Alpha Kolibri: Open-Weight Specs, Benchmarks and Setup
Aleph Alpha Kolibri is a 78B open-weight MoE, 3.46B active, Apache 2.0, English and German, up to 1M tokens. Not on RouterPlex yet.

Aleph Alpha released Kolibri on 3 October 2026, the Day of German Unity. It is a bilingual English-German mixture-of-experts model: 78.1B parameters, 3.46B active per token, a validated context of 1,048,576 tokens, and Apache 2.0 weights on Hugging Face as Aleph-Alpha/Kolibri-1.
Kolibri is not on RouterPlex yet. The announcement and the model card publish weights and a vLLM command. They do not publish a hosted API price or a RouterPlex model ID. Create a RouterPlex account if you want a prepaid key ready the day a catalog row exists. Until then, the way to run Kolibri is the open weights.

Sources: Aleph Alpha's Kolibri announcement (3 October 2026), the Kolibri-1 model card, and the tech report. Checked 4 October 2026. Every score on this page is Aleph Alpha's. RouterPlex has not re-run these evals. The launch image is Aleph Alpha's. The charts are RouterPlex drawings of their published numbers.
Kolibri at a glance #
| Kolibri | |
|---|---|
| Released | 3 October 2026 |
| Maker | Aleph Alpha Research GmbH |
| Hugging Face | Aleph-Alpha/Kolibri-1 (FP8). Parent: Aleph-Alpha/Kolibri-1-BF16 |
| License | Apache 2.0 |
| Total / active | 78,103,074,560 / 3,457,573,120 (78.1B / 3.46B) |
| Experts | 384 per layer, 1 shared, 6 routed |
| Context | 1,048,576 validated. Trained to 262,144. Card recommends ≤262,144 for serving |
| Languages | German and English |
| Knowledge cutoff | 18 June 2026, both languages |
| Weights on disk | FP8, about 78 GB. Router, embeddings, head, and norms stay BF16 |
| Reasoning | none, low, medium, high |
| On RouterPlex | Not yet. No model ID, no list price |
There is no public per-token Kolibri API price. Aleph Alpha's contact for enterprise deployment is their sales page. On-prem GPU time is the bill until a hosted row exists.
What the open-weight model actually is #
Kolibri is the second model through Aleph Alpha's Model Factory. The first, Kolibri Origin, finished pre-training on 11 June 2026 at 30.6B total and 3.27B active, with a 65,536-token trained context. Origin was not released. Kolibri finished pre-training on 11 September and shipped three weeks later.
The size choice is a serving decision, not a "bigger is better" one. Aleph Alpha says a 123B variant could serve only 3 concurrent 256k queries on two H100s, while 78B handled 18 and decoded 28% faster. They also preferred 384 narrower experts over fewer wide ones. Of 50 layers, 10 use full attention (every 5th layer) and 40 use a 512-token sliding window, so decode cost in those layers does not grow with context.
Pre-training ran on 768 B200 GPUs for 21 days and 20T tokens at a 16k sequence length, then 3.44T tokens of mid-training at 64k, then 201B tokens of long-context adaptation at 256k (the blog rounds that last stage to 200B). The model card puts pre-training at 392,000 GPU-hours, mid-training at 90,000, and long context at 10,000. Estimated training energy, including data-center overhead, is 950 MWh. That figure excludes supervised fine-tuning and reinforcement learning.

Routing during training uses exact quantile balancing. Aleph Alpha contrasts that with Kimi K3, which estimates the global quantile from histograms. They say the exact version has a fixed communication cost and improved both load balance and quality. The optimizer is Muon, as it was for Origin. They report no loss spikes on either run.
Post-training is supervised fine-tuning, then reinforcement learning. The SFT mix is 268B tokens after filtering, including 174B of synthetic tokens before that filter. RL used more than 1.2 million tasks. Reasoning effort is a trained control: none, low, medium, or high. The chat template treats an omitted effort as high.
Kolibri vs Kolibri Origin #
| Kolibri Origin | Kolibri | |
|---|---|---|
| Pre-training finished | 11 June 2026 | 11 September 2026 |
| Public weights | No | 3 October 2026, Apache 2.0 |
| Total / active | 30.6B / 3.27B | 78.1B / 3.46B |
| Pre-training tokens | 7.51T | 20T |
| Layers | 50 (2 dense + 48 MoE) | 50, all MoE, 1 shared expert |
| Experts, total / active | 128 / 8 | 384 / 6 |
| Attention | Full, every layer | 512-token window, full every 5th layer |
| Longest trained context | 65,536 | 262,144 |
| Vocab | 96,000 | 128,000 |
| Reasoning | One mode | none, low, medium, high |
| Knowledge cutoff | EN 1 Sep 2024, DE 1 Aug 2025 | EN and DE 18 Jun 2026 |
Active parameters barely moved. Total parameters, expert count, context, and the German data share are what changed.
Benchmarks Aleph Alpha published #
These numbers come from Aleph Alpha's own harness, with each model at its highest reasoning effort. They are not an independent re-run.

The announcement's comparison set, on a shared 0–100 scale. Higher is better. A dash means Aleph Alpha did not publish a score.
| Benchmark | Kolibri | Origin | Qwen3.6-35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|---|
| AIME 2025 | 96.9 | 81.9 | 84.6 | 91.7 | 79.8 |
| AIME 2025 (DE) | 87.5 | 73.5 | 82.9 | 85.6 | 72.3 |
| AIME 2026 | 96.0 | 81.5 | 91.0 | 90.4 | 83.1 |
| AIME 2026 (DE) | 90.0 | 75.2 | 84.4 | 87.5 | 78.5 |
| GPQA Diamond | 84.3 | 68.1 | 83.4 | 78.0 | 74.7 |
| GPQA Diamond (DE) | 81.3 | 58.5 | 80.6 | 76.6 | 72.9 |
| AA-Omniscience Index | −32.8 | −64.0 | −15.3 | −36.5 | −24.0 |
| BrowseComp | 29.4 | 4.4 | 26.9 | 29.1 | — |
| τ³-bench banking | 38.1 | 5.7 | 10.6 | 15.5 | 5.7 |
| τ²-bench retail | 69.9 | 58.5 | 71.6 | 67.5 | 62.9 |
| τ²-bench airline | 76.7 | 58.7 | 70.7 | 72.7 | 40.0 |
| τ²-bench telecom | 94.7 | 67.5 | 99.1 | 68.1 | 41.5 |
| BFCL v4 overall | 61.4 | 36.4 | 67.2 | 61.0 | 58.0 |
| LiveCodeBench v6 | 85.9 | 59.2 | 82.5 | 82.0 | 71.2 |
| HumanEval+ | 92.7 | 76.8 | 92.8 | 94.7 | 92.8 |
| LongBench Pro | 64.5 | — | 70.8 | 62.9 | 56.4 |
| AA-LCR | 68.3 | — | 69.7 | 67.0 | 52.3 |
Qwen3.6-35B-A3B is the fair peer: about 3B active, same order of magnitude as Kolibri's 3.46B. Nemotron 3 Super is the "up to four times the active parameters" model Aleph Alpha names (about 12B active). Mistral Small 4 is about 6B active.
Where Kolibri leads that peer set
- Math, including German. AIME 2025 is 96.9, against 91.7 for Nemotron 3 Super and 84.6 for Qwen3.6. German AIME 2025 is 87.5, against 85.6 and 82.9.
- Science questions. GPQA Diamond is 84.3 in English and 81.3 in German. The German drop from the English score is 3.0 points. Origin dropped 9.6.
- One agent bench by a wide margin. τ³-bench banking is 38.1, against 15.5 for Nemotron 3 Super and 10.6 for Qwen3.6. That is still a low absolute score.
- LiveCodeBench v6. 85.9, against 82.5 and 82.0.
Aleph Alpha's overall row, which averages a wider set than the table above, puts Kolibri at 75.5 English and 70.8 German. Among the mixture-of-experts models in that row, those are the high scores. Qwen3.5-35B-A3B is close on English at 74.7. The dense Qwen3.8 27B is higher still, at 80.2 and 79.9, and it activates about 27B parameters per token rather than 3.46B. Aleph Alpha greys dense models out of the "best MoE" marks for that reason.
Where Kolibri trails
- SWE-bench Verified: 66.4. Qwen3.6 is 73.8. Qwen3.5 is 71.6. Qwen3.8 27B is 72.6. Nemotron 3 Super is 60.2.
- Tool calling. BFCL v4 overall is 61.4, behind Qwen3.6 at 67.2. Kolibri is level with Nemotron 3 Super at 61.0.
- Long context. LongBench Pro is 64.5, behind Qwen3.6 at 70.8 and ahead of Nemotron 3 Super at 62.9.
- Terminal work. TerminalBench 2.1 is 27.7. Qwen3.5 and Nemotron 3 Super are both 39.7. Qwen3.8 27B is 76.8. Qwen3.6 has no published cell in this row.
- The omniscience index. −32.8 is better than Origin's −64.0 and worse than Qwen3.6's −15.3. Higher is better, and a negative index means the model is still punished for wrong answers.
If the job is a software-engineering agent or a tool-calling loop, this table does not make Kolibri the default. If the job is German and English math, science questions, or the banking tool bench they published, it does.
German, on purpose #
The announcement says 21.3% of the 20T pre-training tokens are German, about 4.3T, against roughly 62% English and 14% code. The model card describes that same 20T mix as about 62.5% English, 23.9% German, and 13.6% code. Both put German in the low twenties, not in the leftover bin. Each unique German token was seen about 1.8 times.
Open German datasets were not enough. After dedup and filtering, Aleph Alpha had 390B German tokens against a target near 4T. They filled the gap with a German Common Crawl pipeline (1.3T unique tokens; English word-length filters were dropping administrative German), LLM rephrases of German documents (about 1T; the source stays German), and a small translated share. Translation was a Kolibri Origin tool. They argue it imports English-web geography and institutions along with the words.
The tokenizer is their own 128k English-German UniBPE. It is the part of the release that changes your bill if you serve the weights yourself: fewer German tokens means less decode for the same document.

On German web text (FineWeb-2), Kolibri scores 4.90 bytes per token, the best figure in Aleph Alpha's comparison. Their own plain BPE 128k control, trained on the same data, scores 4.89. The compression gain from the training method is a hundredth of a byte. The gain against other tokenizers is the real one: GLM 5.3 is 3.93, DeepSeek V4 is 3.72, Kimi K3 is 3.28. A German document that costs 1,000 Kolibri tokens costs roughly 1,490 Kimi K3 tokens at those rates.
English web is a different story. Kolibri is 4.58 bytes per token. Plain BPE is 4.59. GPT-5 is 4.67, which is better. "Best German compression in this table" is accurate. "Better English compression than the frontier tokenizers" is not.
The word pictures are where UniBPE shows up. Kolibri splits *Bundessozialgerichtes* into Bundes / sozial / gericht / es. GPT-5, Qwen3.8, and Gemini split the middle into ess / oz / ial. *Protokolldaten* is Protokoll / daten for Kolibri and a five-piece cut for GPT-5. Cleaner morphemes are the claim. A large bytes-per-token win over their own BPE is not.
Grounding: it is trained to abstain #
Aleph Alpha's regulated-customer pitch is a model that says it does not know. Training includes "I don't know" targets plus their Merlin-Arthur procedure: one player adds evidence, one strips it, and the model has to answer only when the context still supports an answer. The paper is arXiv:2512.11614.
| Benchmark | Kolibri | Origin | Qwen3.6-35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|---|
| AA-Omniscience non-hallucination | 44.0 | 14.8 | 56.7 | 13.9 | 34.7 |
| RGB: holds back | 85.6 | 73.9 | 79.6 | 74.6 | 82.3 |
| RGB: invents nothing | 87.3 | 75.6 | 84.3 | 86.0 | 87.0 |
| FRAMES | 71.2 | 65.7 | 74.7 | 74.9 | 71.9 |
| M/A grounding score | 0.23 | 0.00 | 0.12 | 0.00 | 0.06 |
Kolibri is much harder to bait than Origin. It is not the most reluctant model in the set: Qwen3.6's non-hallucination rate is 56.7 against Kolibri's 44.0. The M/A score, which Aleph Alpha describes as a lower bound on how much of an answer came from the document, is where Kolibri leads (0.23 vs 0.12).
Sector scores, on their own suites #
Public benchmarks miss German public administration, aerospace, and factory language, so Aleph Alpha built private suites and synthetic training environments. They say no customer data went into training. These are still vendor evals of vendor tasks.
| Benchmark | Kolibri | Qwen3.6-35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| Honeypot (agentic RAG) | 80.8 | 74.3 | 68.8 | 68.1 |
| MuSiQue, cleaned | 77.3 | 61.2 | 79.1 | 66.8 |
| Semiconductors | 80.4 | 79.4 | 69.6 | 62.7 |
| German public sector | 75.0 | 72.0 | 78.0 | 50.0 |
| Aerospace | 58.9 | 59.0 | 54.9 | 47.0 |
| Automotive supplier | 99.0 | 92.6 | 91.0 | 87.1 |
| Industrial drive technology | 60.0 | 59.5 | 37.3 | 56.8 |
Kolibri leads four of the seven. Nemotron 3 Super leads cleaned MuSiQue and the German public-sector suite (78.0 vs 75.0). Aerospace is a tie with Qwen3.6, 58.9 against 59.0. A 99 on the automotive-supplier suite is a vendor number on a vendor task. Treat it as a clue about the training target, not as a score you can quote to a customer.
Run Kolibri with vLLM today #
This is Aleph Alpha's command. It serves weights you download. It does not call api.routerplex.com.
Install the plugin, which also installs the vLLM build it supports, or use the container ghcr.io/aleph-alpha/aleph-alpha-inference:
pip install "aleph-alpha-inference>=1.0"vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \--reasoning-parser kolibri1 \--tool-call-parser kolibri1 \--enable-auto-tool-choice
Contexts past 262,144 tokens need an override. Aleph Alpha's line is:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \--reasoning-parser kolibri1 \--tool-call-parser kolibri1 \--enable-auto-tool-choice \--max-model-len 1048576 \--hf-overrides '{"max_position_embeddings": 1048576}'
Recommended sampling is temperature=1.0, top_p=0.97, top_k=128. Pass reasoning effort as none, low, medium, or high through the chat template. Leaving it unset selects high.
The FP8 checkpoint is about 78 GB. The model card's minimum is 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200, or 1× B300. The recommended set drops the A100s: 2× H100 SXM5, 2× H200, 1× B200, or 1× B300. That is the whole model in memory. Active parameters only reduce the math, not the footprint.
Not on RouterPlex, and what is #
RouterPlex: Kolibri will be on this catalog once it is listed. Hugging Face availability is not that bar. There is no model ID to paste into Cursor, Claude Code, or the OpenAI SDK, and there is no price to quote.
When a row exists it will be vendor list price, 0% markup, prepaid, with a hard per-key budget. The request stops at $0. The OpenAI-compatible chat API is the contract those rows use. Aleph Alpha has not published a list price, so this page will not invent one.
The comparison models that already have a hosted route are live here:
- Qwen3.8-Max at $2 / $6 per 1M, the dense model that leads Aleph Alpha's overall row. See Qwen3.8-Max API pricing.
- Qwen3.8 Flash at $0.15 / $0.47 per 1M when you want the cheap sibling.
- GLM-5.3 at $1.40 / $4.40 per 1M. The tokenizer table above is one reason German text can cost more on that route than on Kolibri's own tokenizer.
curl https://api.routerplex.com/v1/chat/completions \-H "Authorization: Bearer $ROUTERPLEX_API_KEY" \-H "Content-Type: application/json" \-d '{"model": "qwen3.8-max","messages": [{"role": "user", "content": "Fasse die Risiken dieses Migrationsplans auf Deutsch zusammen."}]}'
That ID is Qwen3.8-Max. It is not Kolibri. Another open-weight model waiting on a catalog row is Aikido Altar-1, a security prune of GLM-5.3.
Weights this week, a route when it is listed #
Download Aleph-Alpha/Kolibri-1 and serve it with the vLLM command above if you need the open-weight model this week. Create a RouterPlex account if you want the prepaid key ready when Kolibri is listed, and use the models already on the marketplace until then. This page will name the RouterPlex model ID on the day the catalog row exists.
Common questions
Frequently asked questions
What is Aleph Alpha Kolibri?
Kolibri is Aleph Alpha's open-weight mixture-of-experts model, released 3 October 2026. It has 78.1 billion parameters, activates 3.46 billion per token, and is trained for English and German. Weights are on Hugging Face as Aleph-Alpha/Kolibri-1 under Apache 2.0.
Is Kolibri an open-weight model?
Yes. Aleph-Alpha/Kolibri-1 is public on Hugging Face under the Apache 2.0 license. The checkpoint on that repo is FP8. The BF16 parent is Aleph-Alpha/Kolibri-1-BF16. Kolibri Origin, the earlier 30.6B model, was not publicly released.
What is Kolibri's context window?
Aleph Alpha validated quality up to 1,048,576 tokens. The model was trained out to 262,144 tokens, and the model card recommends staying at or below that length for serving efficiency and complex tasks. Longer contexts need an explicit vLLM max-model-len override.
How do I run Kolibri with vLLM?
Install aleph-alpha-inference 1.0 or newer, then run vllm serve Aleph-Alpha/Kolibri-1 with --kv-cache-dtype fp8, --reasoning-parser kolibri1, --tool-call-parser kolibri1, and --enable-auto-tool-choice. Recommended sampling is temperature 1.0, top_p 0.97, top_k 128. The FP8 weights need about 78 GB.
Is Kolibri on RouterPlex?
Not yet. There is no RouterPlex model ID and no published per-token API price. The weights are the way to run Kolibri today. Create a RouterPlex account if you want a prepaid key ready when a catalog row exists. Qwen3.8 and GLM-5.3 are already live on the same key.
How does Kolibri compare with Qwen3.6-35B-A3B?
In Aleph Alpha's 3 October 2026 table, both are about 3B active. Kolibri leads AIME 2025 (96.9 vs 84.6), German AIME 2025 (87.5 vs 82.9), and tau3-bench banking (38.1 vs 10.6). Qwen3.6 leads SWE-bench Verified (73.8 vs 66.4), BFCL v4 (67.2 vs 61.4), and LongBench Pro (70.8 vs 64.5). RouterPlex has not re-run these evals.
Start from a model that is listed.
This release is not on RouterPlex yet. Open a live catalog row, or browse every listed model and its price.



