Aleph Alpha relies on 3.46 billion active parameters for its sovereign model
Aleph Alpha releases Kolibri, an open-weights bilingual German-English model with 78 billion parameters, of which approximately 3 billion are active per token. Under the Apache 2.0 license, it combines a Mixture-of-Experts architecture, a context window of up to one million tokens, and multiple levels of reasoning.
A 78-billion-parameter model that activates only a fraction
Kolibri contains 78.1 billion parameters, while its Mixture-of-Experts architecture activates roughly 3.46 billion per token. Aleph Alpha designed the model around inference efficiency for enterprise, public administration, industrial and aerospace workloads.
The model's weights are available on Hugging Face under the Apache 2.0 license. Organizations can therefore deploy it on their own infrastructure without sending internal data to a third-party inference service.
Kolibri is natively German-English and supports four reasoning effort levels: none, low, medium and high. Its context window can reach one million tokens, compared with 65,536 for Kolibri Origin, the earlier 30.6-billion-parameter experimental model. 384 experts behind the model
All 50 layers in Kolibri use a MoE architecture with 384 experts, six of which are active. Only ten layers process the entire context, while the remaining 40 use a 512-token sliding window.
Aleph Alpha says it selected the 78B configuration after experimenting with architectures up to 123B. In its internal tests, two H100s could handle 18 concurrent 256,000-token requests with the 78B model, compared with three for the 123B variant. The 78B configuration also decoded 28% faster in that comparison.
The main training run used 768 NVIDIA B200 GPUs. It included 20 trillion pre-training tokens over 21 days, followed by 3.44 trillion tokens of mid-training and another 200 billion tokens for long-context adaptation. German accounts for more than one fifth of pre-training
The bilingual design does not primarily depend on material translated from English. 21.3% of Kolibri's pre-training tokens are German, representing roughly 4.3 trillion tokens across the 20-trillion-token first stage.
Aleph Alpha built part of this dataset directly from the German web, adapting its filtering pipeline to the language. Existing German documents were also rephrased into different forms while preserving their original cultural context.
A 128,000-entry bilingual tokenizer called UniBPE accompanies the model. It combines elements of BPE and Unigram to better preserve language morphology, particularly German compound words. In Aleph Alpha's own comparison, Kolibri achieves the highest German compression among the evaluated tokenizers. Learning when not to answer
Part of the post-training process explicitly focuses on grounding. Training examples where the correct response is "I don't know" are included so that missing information does not automatically lead to a guess.
Aleph Alpha also uses its Merlin-Arthur protocol. Arthur is the model being trained. Merlin creates contexts containing the evidence required to answer, while Morgana removes the relevant evidence in an attempt to trigger a hallucination. Arthur must learn to distinguish between the two situations.
On the public AA-Omniscience set, Aleph Alpha reports that Kolibri abstains rather than answering incorrectly on 44% of relevant items, compared with 15% for Kolibri Origin. These results come from evaluations conducted by the company. An automated training pipeline between two generations
Kolibri follows Kolibri Origin after three months between their respective pre-training runs. Total parameters increased from 30.6 billion to 78.1 billion, pre-training data from 7.51 trillion to 20 trillion tokens, and the longest trained context from 65,536 to 262,144 tokens.
During Kolibri's 21-day pre-training run, Aleph Alpha reports 38 unplanned interruptions, roughly one every 10,000 GPU hours. Its pipeline automatically moved training jobs to other machines and resumed from checkpoints without manual intervention.
Kolibri is available as open weights under Apache 2.0. Running it through vLLM currently requires the `aleph-alpha-inference` package, with reasoning and tool calling supported. Contexts beyond 262,144 tokens require an explicit configuration to extend the model length to 1,048,576 tokens.