Skip to main content
Back to News
news/AI Models

Aleph Alpha Releases Kolibri, a Sovereign Open-Weight Model

Aleph Alpha released Kolibri, a 78B open-weight MoE reasoning model under Apache 2.0, with a 1M-token context, tool calling and FP8 weights.

Stefan Trbojevic

Stefan Trbojevic

3 October 20262 min read
LinkedIn
Abstract grid of glowing routing nodes with a sparse cluster of bright active nodes

The takeaway

Aleph Alpha chose a 78B sparse model over a stronger 123B one because it serves six times as many concurrent long-context requests: the sovereign-AI pitch now rests on serving economics, not raw scale.

Why it matters for builders

Kolibri needs roughly 78 GB of FP8 weights, so the minimum self-hosted footprint is two A100 80GB cards, two H100s, or a single H200, B200 or B300, served through a vLLM plugin and container image. For teams that cannot send data to a US API, a 78B model with about 3.5B active parameters, tool calling, a 1M-token context ceiling and Apache 2.0 terms is a workable middle layer between small local models and hosted frontier APIs, especially where German-language quality or EU data residency are hard requirements.

Aleph Alpha Releases Kolibri, a Sovereign Open-Weight Model

Aleph Alpha released Kolibri on 3 October, an English-German mixture-of-experts reasoning model with 78B total parameters published under the Apache 2.0 licence. The German lab is positioning it as sovereign infrastructure: weights you download, inspect and run on hardware you control.

What Aleph Alpha shipped

Kolibri carries 78.1B parameters but activates only 3.46B per token, according to the model card. It supports a context window of 1,048,576 tokens, though Aleph Alpha recommends staying at or below 262,144 for serving efficiency and complex tasks. Weights ship in FP8 with bfloat16 embeddings, norms and router, and the model covers tool calling plus an explicit reasoning mode.

Pre-training ran on 20T tokens of a bilingual corpus, roughly 62.5 percent English, 23.9 percent German and 13.6 percent code, followed by 3.44T mid-training tokens and 201B more for long-context extension. The knowledge cutoff is 18 June 2026.

The efficiency trade-off behind the size

Aleph Alpha explains that it built a 123B model first and walked away from it. The larger version could serve only three concurrent 256k-token queries across two H100s, while the 78B model handles 18 and decodes 28 percent faster. The team used 384 small experts instead of fewer wide ones, and only 10 of 50 layers process the full context; the remaining 40 keep a tight 512-token sliding window, which holds decode cost and memory flat as context grows.

Diagram of sparse expert routing with a few bright active nodes among many dim ones across a circuit grid

What it means for builders

Kolibri needs about 78 GB of FP8 weights, so the minimum footprint is two A100 80GB cards, two H100s, or a single H200, B200 or B300. That is not a laptop model, but it is a realistic self-hosted tier for teams that cannot route sensitive data to a US API. Serving runs through a vLLM plugin and a container image, and Aleph Alpha is a signatory of the EU GPAI Code of Practice.

The practical read is that open-weight releases are converging on sparse architectures that keep long context affordable. A 78B model with roughly 3.5B active parameters, tool calling and a permissive licence is a workable middle layer between small local models and hosted frontier APIs, especially where German-language quality or EU data residency are hard requirements. The same open-weight pull is visible lower in the stack, where Amazon shipped an open-source decision model for agent loops earlier this week.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

3 October 2026

Updated

3 October 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.