Germany's Sovereign Kolibri Model Lands With Open Weights on Reunification Day
Aleph Alpha released Kolibri on October 3, a 78B-parameter bilingual English-German MoE with open Apache 2.0 weights, a 1M token context window, and a compliance story built entirely on German infrastructure and EU law.
Aleph Alpha released Kolibri on October 3, timed to the Day of German Reunification, and the timing is the argument: a 78B-parameter open-weights model (about 3.46B active per token as a mixture-of-experts) built entirely in Germany, trained on German and Finnish infrastructure “with no foreign control,” and positioned for the European public administration, aerospace, and industrial sectors. The weights are on Hugging Face under Apache 2.0, served through their vLLM plugin. After the April merger in which Cohere acquired Aleph Alpha left questions about what Aleph Alpha’s sovereign-AI identity would even mean, Kolibri is the answer: Europe now has a frontier-adjacent model whose entire supply chain sits inside its own jurisdiction.
The Specs: Small Active Footprint, Long Memory
The architecture follows the efficiency frontier this season standardized on, with German-specific choices layered in. 78B total parameters with 384 experts and 6 active per token; context up to 1 million tokens (trained to 256k); sliding-window attention with full attention every fifth layer; and the Muon optimizer. Training ran on 768 B200 GPUs over roughly 24 trillion tokens, surviving 38 unplanned interruptions with automatic recovery over 21 days. The bilingual design is deliberate: 21.3% of pre-training tokens are German (about 4.3 trillion), with a custom bilingual tokenizer that Aleph Alpha claims gives best-in-comparison German text compression. Benchmarks: AIME 2025 at 96.9, GPQA Diamond at 84.3, SWE-Bench Verified at 66.4, with the claim that it matches models carrying up to 4 times its active parameters on math and code.
The Abstention Training Is the Compliance Feature
The most European thing about Kolibri is not where it was trained but what it was trained to say: on the AA-Omniscience benchmark, it answers “I don’t know” on 44% of items, up from its predecessor’s 15%, via a training protocol called Merlin-Arthur. In public administration, an AI that confidently invents an answer is a legal liability under GDPR and the EU AI Act; an AI that abstains is a feature. Kolibri was designed against the EU AI Act, GDPR, and the GPAI Code of Practice from the start, which makes its regulatory posture architectural rather than a policy page. Combined with on-premise deployment freedom, this is the model as a compliance artifact, and for regulated European buyers that may matter more than any leaderboard.
Sovereignty as a Product Category
The context is the shift this blog has tracked all month: Gemini 4 Argon shipping to vetted defenders first, the FTC probing AI product claims, and governments worldwide discovering that agent incidents do not respect vendor jurisdictions. Sovereignty used to be a procurement slogan; Kolibri makes it a shipping artifact with published weights. The competitive read is equally direct: American open-weights dominance from MiMo and DeepSeek has been the story of the season, and Kolibri is Europe’s entry, differentiating on jurisdiction and language rather than raw scale.
What to Watch
Three things. First, independent German-language evaluation, since the compression and quality claims are the model’s core differentiator and come from Aleph Alpha itself. Second, public-sector adoption: the first German Land or ministry deploying Kolibri on-premise would validate the sovereignty thesis far better than benchmarks. Third, the Cohere integration question, since a merged company now spans two model families and two regulatory philosophies, and how Kolibri and Cohere’s offerings coexist will define the combined company’s actual strategy.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Deploy Mistral Small 4 for Multimodal Reasoning and Coding
Learn how to deploy Mistral Small 4 with reasoning controls, multimodal input, and optimized serving on API, Hugging Face, or NVIDIA.
Xiaomi's MiMo v2.6 Takes the Top Open-Weights Spot on the Intelligence Index
Xiaomi released MiMo v2.6 on September 21, a 1-trillion-parameter open-weights MoE model that debuts at number one among open models on Artificial Analysis' Intelligence Index, with an MIT license and aggressive API pricing.
Cloudflare Enters the Decision-Model War With Open-Weights Clef
Cloudflare announced Clef and Clef-flash on October 1, open-weight decision models that beat TypeSafe's Jev on classification benchmarks at lower latency, Jev-API-compatible, and fine-tunable through a new RL platform.
Gemini 4 Argon Lands With 1M-Token Output and a Safety-Gated Rollout
Google announced Gemini 4 Argon on September 30 with a 1 million token output limit, a DeepSWE state of the art at 77.9%, and defensive-cyber focus, but most users cannot touch it yet: rollout runs through a trusted-defenders program first.
GPT-6.1 Sol Arrives With Near-Astra Performance at a Fifth of the Price
Announced at DevDay on September 29, GPT-6.1 Sol claims near-GPT-6 Astra performance at $2/$10 per million tokens, with a 1.05 million token context window and a new Ultrafast tier at 300 tokens per second, one week after GPT-6 Sol launched.