PuppyIP Resource Center
AI Tool Tutorials 8 min Published 2026-10-04

Deploying Kolibri 78B: FP8 and BF16 GPU memory and vLLM plugin requirements

Choose hardware for the complete weights and runtime overhead before installing a compatible plugin. Starting with short input and a local API, this guide helps developers assess deployment feasibility and verify that the service actually works.

Kolibri Aleph Alpha vLLM FP8 Local models

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • 3.46B is the number of active parameters per token, not a basis for estimating the full model's GPU-memory requirements.
  • Distinguish weight format, plugin version, and runtime environment. Complete a small initial check before increasing context and concurrency.
  • This article is based on official deployment material and does not report GPU benchmarks. Successful startup does not establish that real business quality requirements have been met.

First question: can it be deployed like an ordinary 3B model?

Aleph Alpha released Kolibri on October 3, 2026, with open weights and inference integration material. Its MoE architecture has approximately 78B total parameters and activates about 3.46B per token. Active parameters affect computation, but the model still needs space for the complete weights. This does not mean hardware that normally runs a 3B model is sufficient.

For a self-hosted application, confirm that hardware, weight format, and inference software match before downloading files. The initial deployment below assumes a Linux GPU environment. These steps are drawn from public documentation, not tests of a particular graphics card by this site.

FP8 and BF16: account for the whole model and leave runtime headroom

The model card for the FP8 repository Aleph-Alpha/Kolibri-1 lists approximately 78GB of weights. Official minimum-configuration examples include one H200, B200, or B300, or two A100 80GB GPUs; consult that card for the full list. “Approximately 78GB of weights” does not mean a single 80GB GPU is guaranteed to run it.

The BF16 repository Aleph-Alpha/Kolibri-1-BF16 has approximately 156GB of weights. Its minimum-configuration examples include four H100 SXM5 GPUs or two H200 GPUs; recommended configurations also include two B200 GPUs. Check the two formats separately instead of applying FP8 hardware conclusions to BF16.

Before selecting a machine, record GPU models, count, available memory per card, and other processes using them. Beyond weights, caches and runtime overhead need headroom that changes with request settings. Treat fitting the model and supporting the planned requests as separate acceptance checks, rather than discovering an invalid capacity plan after downloading.

Prepare Linux and compatible plugin versions

The official aleph-alpha-inference plugin is currently 1.0.0, with dependencies restricted to vLLM ≥0.29.0 and <0.30.0. The project declares Python ≥3.10 and Linux. This guide uses Python 3.12 to reduce version-interaction problems; that does not mean every later Python release works without adjustment.

The vLLM GPU installation guide recommends a fresh environment to avoid compatibility issues with existing PyTorch and CUDA build combinations. Check the driver and installation requirements for this version and confirm that GPUs are visible to the process. In an isolated test directory, run python3.12 -m venv .venv, followed by source .venv/bin/activate.

In the activated environment, run python -m pip install 'aleph-alpha-inference==1.0.0' 'vllm==0.29.0'. A normal installation brings in a supported vLLM version; do not force a vLLM upgrade and ignore dependency warnings. Afterwards, run python -m pip show aleph-alpha-inference vllm and check dependencies with python -m pip check. Retain the version information for troubleshooting.

First startup: use a local address and small context

The following verification example narrows the official startup parameters for a single-GPU environment meeting the FP8 model-card requirements, such as one B200 from the official recommended list. Confirm available resources and software compatibility first. Multi-GPU deployment requires parallelism configured for the actual topology and cannot simply copy this single-GPU example.

Run this complete command in the environment prepared above: vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice --host 127.0.0.1 --port 8000 --max-model-len 8192 --max-num-seqs 1

Here, 8192 is this article's recommended initial verification limit, covering input and output, not the model's maximum capability. --max-num-seqs 1 limits simultaneously processed sequences for the first check. The address 127.0.0.1 restricts this example to access from the machine hosting the service. Wait for weight loading and service readiness. If startup exits, resolve the first logged error instead of repeatedly sending requests.

For BF16, use Aleph-Alpha/Kolibri-1-BF16, remove --kv-cache-dtype fp8 as in the official example, and allocate resources that meet BF16 requirements. Changing only the model name without reassessing hardware does not complete a format migration.

Send a short request to verify that the client reaches the service

In another terminal on the same Linux machine, run: curl -sS http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"Aleph-Alpha/Kolibri-1","messages":[{"role":"user","content":"Reply with one short greeting."}],"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'

This example disables thinking to check the request path with a simple output. Confirm a normal JSON response and a nonempty message.content inside choices. If an error is returned, retain its type and the server logs from the same time. A listening port or completed download cannot replace this model request.

OpenAI-compatible describes the API format; this example sends the request to your own local service. When integrating a client, set the local base URL and exact model name separately. If the client runs on another computer, its 127.0.0.1 refers to that computer itself. Plan controlled access separately instead of copying the address unchanged.

Get the basic request working before increasing context and business load

The model card lists a maximum of 1,048,576 tokens and recommends no more than 262,144 tokens for efficiency-sensitive deployments and complex tasks. Do not start by targeting the maximum. Save a configuration that completes short requests, then adjust gradually to real document lengths.

Increase one variable at a time: input length first, output budget next, and simultaneous requests last. Record success rate, waiting time, and GPU-memory use, and use your own samples to assess requirements. This article gives no throughput or latency guarantee.

To verify business quality, prepare small samples with known answers and check correctness, formatting, and omissions separately. Before enabling tool calls, use test tools that cannot change real business data and verify parameter parsing and result return. Passing ordinary question-and-answer requests establishes only that the basic request path works.

Troubleshoot by stage instead of changing every setting at once

For dependency errors during installation or before startup, first identify the environment activated in the current terminal, verify plugin and vLLM versions, and address pip check results. Installation records from an old environment do not establish that the new one is ready.

For insufficient GPU memory while loading weights or initializing, check whether FP8 or BF16 is being loaded, which GPUs the process actually sees, and memory used by other tasks. If the full weights do not fit, merely shortening requests cannot supply the missing capacity. If the error occurs during cache planning, check context and concurrency settings.

For a refused connection, check whether the service is still running, the ports match, and the request originates from the same machine. For a missing model or invalid parameters, verify the repository name and JSON in the request. Restore this guide's short request first, then add your own parameters one by one.

When reporting a problem, include plugin and vLLM versions, GPU information, the startup command, the stage of failure, and redacted logs. Also retain the difference between successful and failing configurations; that is more useful than only saying that it will not run.

Sources

Frequently Asked Questions

Should I choose FP8 simply because its files are smaller?

First check whether your hardware meets official requirements and whether the inference software supports the required format, then compare with identical business samples. File size is one selection factor and does not replace task-quality and resource checks.

Can it be launched directly as a general Chinese-language assistant?

Its official positioning primarily targets German and English. Prepare separate samples for Chinese-language work, check answers, terminology, and formatting, then decide whether to integrate it. This article does not infer Chinese quality from English or German results.