The takeaway
Aleph Alpha chose a 78B sparse model over a stronger 123B one because it serves six times as many concurrent long-context requests: the sovereign-AI pitch now rests on serving economics, not raw scale.
Why it matters for builders
Kolibri needs roughly 78 GB of FP8 weights, so the minimum self-hosted footprint is two A100 80GB cards, two H100s, or a single H200, B200 or B300, served through a vLLM plugin and container image. For teams that cannot send data to a US API, a 78B model with about 3.5B active parameters, tool calling, a 1M-token context ceiling and Apache 2.0 terms is a workable middle layer between small local models and hosted frontier APIs, especially where German-language quality or EU data residency are hard requirements.
Aleph Alpha Releases Kolibri, a Sovereign Open-Weight Model
Aleph Alpha released Kolibri on 3 October, an English-German mixture-of-experts reasoning model with 78B total parameters published under the Apache 2.0 licence. The German lab is positioning it as sovereign infrastructure: weights you download, inspect and run on hardware you control.
What Aleph Alpha shipped
Kolibri carries 78.1B parameters but activates only 3.46B per token, according to the model card. It supports a context window of 1,048,576 tokens, though Aleph Alpha recommends staying at or below 262,144 for serving efficiency and complex tasks. Weights ship in FP8 with bfloat16 embeddings, norms and router, and the model covers tool calling plus an explicit reasoning mode.
Pre-training ran on 20T tokens of a bilingual corpus, roughly 62.5 percent English, 23.9 percent German and 13.6 percent code, followed by 3.44T mid-training tokens and 201B more for long-context extension. The knowledge cutoff is 18 June 2026.
The efficiency trade-off behind the size
Aleph Alpha explains that it built a 123B model first and walked away from it. The larger version could serve only three concurrent 256k-token queries across two H100s, while the 78B model handles 18 and decodes 28 percent faster. The team used 384 small experts instead of fewer wide ones, and only 10 of 50 layers process the full context; the remaining 40 keep a tight 512-token sliding window, which holds decode cost and memory flat as context grows.

What it means for builders
Kolibri needs about 78 GB of FP8 weights, so the minimum footprint is two A100 80GB cards, two H100s, or a single H200, B200 or B300. That is not a laptop model, but it is a realistic self-hosted tier for teams that cannot route sensitive data to a US API. Serving runs through a vLLM plugin and a container image, and Aleph Alpha is a signatory of the EU GPAI Code of Practice.
The practical read is that open-weight releases are converging on sparse architectures that keep long context affordable. A 78B model with roughly 3.5B active parameters, tool calling and a permissive licence is a workable middle layer between small local models and hosted frontier APIs, especially where German-language quality or EU data residency are hard requirements. The same open-weight pull is visible lower in the stack, where Amazon shipped an open-source decision model for agent loops earlier this week.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
3 October 2026
3 October 2026
Sources
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.




