Instructions to use velaia/Kolibri-1-MLX-3bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use velaia/Kolibri-1-MLX-3bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("velaia/Kolibri-1-MLX-3bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use velaia/Kolibri-1-MLX-3bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "velaia/Kolibri-1-MLX-3bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "velaia/Kolibri-1-MLX-3bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use velaia/Kolibri-1-MLX-3bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "velaia/Kolibri-1-MLX-3bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "velaia/Kolibri-1-MLX-3bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "velaia/Kolibri-1-MLX-3bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use velaia/Kolibri-1-MLX-3bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "velaia/Kolibri-1-MLX-3bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default velaia/Kolibri-1-MLX-3bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use velaia/Kolibri-1-MLX-3bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "velaia/Kolibri-1-MLX-3bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "velaia/Kolibri-1-MLX-3bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Kolibri-1 — MLX 3-bit
Unofficial MLX conversion of Aleph-Alpha/Kolibri-1 by Aleph Alpha GmbH. Not affiliated with or endorsed by Aleph Alpha.
Modifications: the original FP8 (e4m3, 128×128 block-scaled) weights were dequantized and re-quantized to MLX affine quantization (group size 64). Routed experts use the bit width below; attention and the shared expert are 4-bit, embeddings and LM head 8-bit, and the MoE router is unquantized. Expect quality differences vs. the original, growing at lower bit widths.
| Variant | Bits/weight | Size | Peak memory (short / 4k-token prompt) | Perplexity DE / EN* | Fits |
|---|---|---|---|---|---|
| 4-bit | 4.54 | 41 GiB | 44 / 48 GB | 12.77 / 16.47 | 64 GB+ Macs |
| 3-bit (this repo) | 3.57 | 33 GiB | 35 / 39 GB | 13.17 / 16.31 | 48 GB+ Macs |
| 2-bit | 2.61 | 24 GiB | 26 / 30 GB | 14.01 / 17.46 | 36 GB+ Macs |
*First 4,096 tokens of the German "Kolibris" / English "Hummingbird" Wikipedia articles; lower is better; a small sample, so only roughly indicative. Generation runs at ~52–56 tok/s on an M1 Max for all three.
On a 48 GB Mac, raise the GPU memory limit for the 3-bit variant:
sudo sysctl -w iogpu.wired_limit_mb=40960 (resets on reboot). Don't run
another large model at the same time.
Relative weight-reconstruction RMSE vs. the dequantized FP8 weights: routed experts 0.194 (3-bit) / 0.400 (2-bit) / ~0.10 (4-bit); attention + shared expert 0.098.
Without reasoning_effort, the chat template defaults to high reasoning.
Usage
Native mlx-lm support
Support for the kolibri1 architecture is proposed upstream in
ml-explore/mlx-lm#1945; the badge
above shows its current state. The PR adds the same kolibri1.py that ships in this
repo, so these weights load unchanged. Once it is merged and included in an
mlx-lm release, the workaround below is no longer needed:
pip install -U mlx-lm
mlx_lm.chat --model velaia/Kolibri-1-MLX-3bit --temp 1.0 --top-p 0.97
mlx_lm.generate --model velaia/Kolibri-1-MLX-3bit --prompt "Hallo!" \
--chat-template-config '{"reasoning_effort": "low"}'
(Before a release, you can install mlx-lm from GitHub main once the PR is merged:
pip install -U git+https://github.com/ml-explore/mlx-lm.)
Until then: bundled launcher
mlx-lm does not ship the kolibri1 architecture yet. This repo includes it as
kolibri1.py (ported from Aleph Alpha's Apache-2.0
vLLM plugin), plus
run.py, a small launcher that registers it and then runs the normal mlx-lm
commands (generate, chat, server, ...).
Setup
pip install -U mlx-lm huggingface_hub
hf download velaia/Kolibri-1-MLX-3bit --local-dir Kolibri-1-MLX-3bit
One-off prompt
python Kolibri-1-MLX-3bit/run.py generate --model Kolibri-1-MLX-3bit \
--prompt "Erkläre kurz den Unterschied zwischen Wetter und Klima." \
--chat-template-config '{"reasoning_effort": "low"}' \
--temp 1.0 --top-p 0.97 --top-k 128 --max-tokens 1024
reasoning_effort:none|low|medium|high. Without it, the chat template uses high, which thinks at length.--temp 0gives deterministic output, useful for comparing variants.- Tokens/s and peak memory are printed at the end.
Interactive chat
python Kolibri-1-MLX-3bit/run.py chat --model Kolibri-1-MLX-3bit --temp 1.0 --top-p 0.97 --max-tokens 2048
Type q to quit, h for help; --system-prompt "..." sets a system prompt.
chat has no reasoning-effort or top-k option, so it always reasons at
high. Use generate to test other effort levels.
OpenAI-compatible server
python Kolibri-1-MLX-3bit/run.py server --model Kolibri-1-MLX-3bit \
--temp 1.0 --top-p 0.97 --top-k 128 \
--chat-template-args '{"reasoning_effort": "medium"}'
Listens on http://localhost:8080/v1. Clients can also pass
chat_template_kwargs per request.
Python
import sys, importlib.util
from huggingface_hub import snapshot_download
path = snapshot_download("velaia/Kolibri-1-MLX-3bit")
spec = importlib.util.spec_from_file_location("mlx_lm.models.kolibri1", f"{path}/kolibri1.py")
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)
sys.modules["mlx_lm.models.kolibri1"] = mod # register before load()
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tok = load(path)
prompt = tok.apply_chat_template([{"role": "user", "content": "Hallo! Wer bist du?"}],
add_generation_prompt=True, tokenize=False, reasoning_effort="low")
print(generate(model, tok, prompt, max_tokens=1024, verbose=True,
sampler=make_sampler(temp=1.0, top_p=0.97, top_k=128)))
Notes: a [transformers] ... model of type kolibri1 warning on load is
harmless. Recommended sampling (from the original release): temperature 1.0,
top_p 0.97, top_k 128. Don't load two large models at the same time.
Provenance
- Source:
Aleph-Alpha/Kolibri-1@e52eb4627d11516b0c01de49210ab5a4e4061444(FP8 release) - mlx 0.32.3, mlx-lm 0.32.0; converted on an Apple M1 Max (64 GB)
- Recipe:
kolibri1.py(architecture),run.py(CLI launcher) andconvert.py(mixed-precision predicate):python -m kolibri_mlx.convert --hf-path Kolibri-1 --mlx-path out --expert-bits 3 --dense-bits 4
License & responsible use
Weights: Apache 2.0, see LICENSE (copied unchanged from the original repo).
Per the original model card, the license applies only to the weights and
configuration files; Aleph Alpha retains all rights to its code, architecture
and training methods. Please follow the original card's responsible-use
guidance (no unlawful use, no practices prohibited by Art. 5 EU AI Act, no
military/nuclear applications) and see it for evaluations, limitations and
training details.
- Downloads last month
- 484
3-bit