On October 3, Aleph Alpha, a German company, unveiled Kolibri, an Anglo-German AI model with open weights aimed at governmental organizations and regulated industries. This release coincided with Germany's Unity Day, as noted in the developer's blog.
Kolibri is based on the Mixture-of-Experts architecture and features 78.1 billion parameters, utilizing approximately 3.46 billion parameters for processing each token.
The model is designed to support reasoning and invoke external tools, with a maximum context window of 1,048,576 tokens.
The complete weights for Kolibri have been released on Hugging Face under the Apache 2.0 license, with versions available in BF16 and FP8 formats.
"Sovereign" Model
Aleph Alpha describes Kolibri as a "sovereign" AI model tailored for critical tasks in public administration, industry, and aerospace sectors.
The development teams worked in Germany, and the training was conducted using German and Finnish infrastructure while adhering to EU and German regulations.
Organizations can deploy Kolibri on their own infrastructure, ensuring that internal data remains private and is not shared with third-party cloud AI service providers. Developers emphasized that the model was created in compliance with the European AI Act, the Code of Practice for General Purpose AI Models, and GDPR regulations.
Kolibri has been programmed to refrain from providing answers when the context is insufficient to verify information, utilizing Aleph Alpha's proprietary Merlin-Arthur approach.
Training Involved Nearly 24 Trillion Tokens
Kolibri underwent pre-training on 768 Nvidia B200 GPUs, with the primary phase lasting 21 days and encompassing 20 trillion tokens, each with a sequence length of 16,384.
Subsequently, the model was fine-tuned using 3.44 trillion tokens with a context length of 65,536 and adapted for extended context using 201 billion tokens with a length of 262,144. In total, around 23.64 trillion tokens were utilized across three training phases.
Of the original dataset, approximately 21.3% (around 4.3 trillion tokens) was in German, while English accounted for about 62%, and programming code made up about 14%.
Aleph Alpha highlighted that to maintain linguistic and cultural context, the German portion of the dataset was not primarily built on machine translations from English, with only about 6% of the content being translated.
A separate Anglo-German tokenizer was also developed for Kolibri, which considers the morphology of both languages, including the complex compound words typical in German.
The model supports four reasoning modes: none, low, medium, and high.
Aleph Alpha Compares Kolibri with Competitors
In its own evaluations, Kolibri scored 96.7% on AIME 2026 and 90.2% on the German version of the benchmark. Its scores on GPQA Diamond and LiveCodeBench v6 were 84.8% and 85.1%, respectively.
Aleph Alpha conducted comparisons of Kolibri with models such as Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B, and Mistral Small 4 119B-A6B.
According to the company, for certain tasks involving mathematics, programming, long-context handling, and tool utilization, Kolibri performs comparably to models that utilize up to four times as many active parameters.
Notably, the FP8 version of Kolibri requires about 78 GB of memory, with Aleph Alpha recommending a minimum configuration of one Nvidia H200, B200, or B300, or two A100 80 GB/H100 SXM5 units. The BF16 weights necessitate approximately 156 GB of memory.
As a reminder, since August 2, the European Commission has had the authority to impose fines on providers of general-purpose AI models for non-compliance with the AI Act, with maximum penalties reaching €15 million or 3% of the global annual turnover.
