Aleph Alpha Kolibri: What It Costs to Run the German AI Model and When It Pays Off
On October 3, 2026, Aleph Alpha released Kolibri, a German-English AI model that companies may download for free and use commercially under the Apache 2.0 license. Running it takes data center GPUs, for example two with 80 GB each, rented from OVHcloud starting at €2,618 per month. That pays off if your data can't go to an external AI service; at typical small and midsize business volumes, an AI API is far cheaper.
What is Aleph Alpha Kolibri?
Kolibri (German for hummingbird) is an open-weight language model: you download it and run it on your own or rented hardware instead of sending text to someone else's service. The German AI company Aleph Alpha released it as Kolibri-1 on Hugging Face on October 3, 2026, German Unity Day. The Apache 2.0 license lets you use and modify it without license fees, including commercially, and it can't be revoked.
The key specs, according to Aleph Alpha and the model card:
- 78.1 billion parameters, of which only 3.46 billion are active per token, a token being a word fragment (mixture of experts)
- German and English, with more than a fifth of the training data in German
- trained on a context length of 262,144 tokens, usable up to just over a million tokens according to the company
- a reasoning mode with four levels, plus tool calling for agent workflows
- developed in Germany, trained in Germany and Finland
- built for regulated sectors such as public administration, industry, and aerospace
How good is Kolibri in German?
Good, but it's not the strongest open model. In Aleph Alpha's own comparison, Kolibri reaches an overall German score of 70.8, narrowly the best result among the mixture-of-experts models it was compared with. Clearly ahead is the dense Qwen3.8 27B from Alibaba Cloud, which uses all of its parameters for every token. That's why Aleph Alpha grays it out in the table.
| Model | Active parameters (class per Aleph Alpha) | Overall score, German | Overall score, English |
|---|---|---|---|
| Kolibri | 3B | 70.8 | 75.5 |
| GPT-OSS 120B | 4 to 6B | 70.2 | 72.3 |
| Nemotron 3 Super | 12B | 67.9 | 73.0 |
| Mistral Small 4 | 4 to 6B | 61.4 | 63.1 |
| Qwen3.8 27B (dense) | 27B | 79.9 | 80.2 |
One clear outlier is an agent benchmark set in banking (Tau3-Bench): Kolibri scores 38.1 points, the second-best mixture-of-experts model 16.0. Claude, ChatGPT, and Gemini are not part of the comparison, and all scores come from the vendor.
Our take: Kolibri's selling point is performance per unit of compute, not peak performance. By Aleph Alpha's calculation, two H100 cards have room for 18 concurrent requests with 256,000 tokens of context, compared with only three for a larger 123B variant. China isn't entirely out of the picture, though: according to the technical report, several of the teacher models that generated training data were built in China. Aleph Alpha says it filters out content that follows Chinese Communist Party narratives.
What hardware does Kolibri need?
Data center GPUs, not an office PC. In FP8 format, the weights take up about 78 GB. The model card lists the minimum as two A100s with 80 GB each, two H100 SXM5s, or a single H200, B200, or B300.
You run Kolibri with the serving software vLLM plus an add-on package from Aleph Alpha. vLLM provides an OpenAI-compatible API, so existing applications and automations can be switched over in principle.
There is no ready-made service with pay-per-request billing yet: as of October 5, Hugging Face lists no inference provider for Kolibri.
What does self-hosting Kolibri cost compared with an AI API?
At the volumes typical for small and midsize businesses, many times more. The German price list of the French provider OVHcloud includes two suitable servers, with prices including German VAT: two A100s with 80 GB each, the official minimum, for about €6.55 per hour or €2,618 on the monthly plan, and two H100s with 80 GB each for about €6.66 per hour or €4,617 per month. The H100 is the newer generation, but OVH offers it as a PCIe variant, while the model card names the SXM5.
Keep in mind: this is hosting with a European provider, so your data sits with OVHcloud. If you have to keep it in your own data center, you buy the hardware yourself, and its price isn't included here.
For comparison, here is a sample calculation with Claude Sonnet 5.5, which Anthropic introduced on September 28. According to Anthropic's price list, it costs $2 per million input tokens and $10 per million output tokens.
- Assumption: 500 emails per day, each with 3,000 input tokens and 800 output tokens including reasoning, over 30 days
- API: 45 million input tokens ($90) plus 12 million output tokens ($120), for a total of $210 per month
- GPU server: €2,200 to €3,880 per month excluding VAT, a good 10 to 18 times as much even at a one-to-one exchange rate
On paper, the balance only tips at around 5,200 to 9,200 of these emails per day, provided the server can handle that volume. If you don't need answers right away, Anthropic's Batch API cuts the price in half. Self-hosting also adds staff time for updates and monitoring. An eight-hour test day costs €52 to €53: trying it out is cheap, running it permanently is not.
When does Kolibri pay off for your company?
When data sovereignty matters more than price. Our decision guide:
| Situation | Our recommendation |
|---|---|
| Data must not go to an AI provider, a server at a European host is allowed | Test Kolibri on a rented server |
| Data must not leave your own data center at all | Run the numbers for Kolibri on your own hardware, GPU purchase price included |
| Thousands of similar tasks per day, around the clock | Run the numbers for Kolibri against an API |
| A few hundred tasks per day, no GPU experience in-house | Stick with an AI API |
| Demanding texts where quality comes first | Top model via API, keep an eye on Kolibri |
The company behind it is worth watching, too: according to its press release of October 5, Aleph Alpha has announced an agreement with Cohere that still needs regulatory approval; until closing, the company continues to operate independently. The published weights remain usable thanks to the Apache license; what happens with future versions and support is an open question.
Our take: For most automations at small and midsize businesses, an API remains cheaper and simpler for now. Kolibri becomes interesting for them once a European provider offers it with pay-per-request billing, because then the fixed monthly cost goes away.
What does this mean for you?
A small test tells you whether Kolibri fits your data better than any debate about principles:
- Pick a process where data privacy concerns have held back the use of AI so far, for example contracts, HR files, or technical specifications.
- Collect 30 real, anonymized examples along with the correct answers.
- Rent a GPU server for a day, start Kolibri with vLLM, and run all the examples through it.
- Test the same examples with an API model and, if European origin isn't a requirement, with Qwen3.8 27B. Compare accuracy, response time, and cost per case.
- Only then decide: run it yourself, wait, or stay with the API.
Our use cases show what automated AI workflows look like in practice. If you'd rather not set up the comparison yourself, we can help as part of our AI automation services: get in touch.
Frequently asked questions
Why does Kolibri need so much memory if only 3.46 billion parameters are active per token?
Kolibri is a mixture-of-experts model: in each layer, it picks 6 of 384 specialized subnetworks for every token. That saves compute, but all 78 billion parameters still have to sit in memory, because different tokens call on different experts. Aleph Alpha itself describes the memory requirement as the trade-off of this architecture.
Can I modify Kolibri and build it into my own products?
The Apache 2.0 license allows that, including commercially and free of charge. If you distribute the model or a modified version, you have to include the license text, keep the copyright notices, and mark any files you changed. For your specific case, it is still worth having your legal team take a look.
Does Kolibri know about current events?
No, its built-in knowledge ends on June 18, 2026. It only gets newer or company-internal information through tools such as a search or connected documents. Aleph Alpha also positions Kolibri for workflows in which a person reviews the results before anyone acts on them.
Is a self-hosted AI model automatically GDPR compliant?
No. Whether your use complies with data protection law depends on how you run it: where the data is processed, who has access, and what gets logged. Aleph Alpha says it built Kolibri with the GDPR and the EU AI Act in mind from the start, and it has signed the EU General-Purpose AI Code of Practice. That does not replace a review with your data protection officer.
Does Kolibri run on smaller hardware?
Not officially. Hugging Face already lists more than a dozen quantized, meaning compressed, community versions that are supposed to need less memory (as of October 5). Since compression can cost quality, we would only use them in production after a direct comparison with the official version.
Sources
Tell me what you're planning. You'll get an honest assessment within 24 hours.
More articles