Kolibri-1 — MLX 3-bit

Unofficial MLX conversion of Aleph-Alpha/Kolibri-1 by Aleph Alpha GmbH. Not affiliated with or endorsed by Aleph Alpha.

Modifications: the original FP8 (e4m3, 128×128 block-scaled) weights were dequantized and re-quantized to MLX affine quantization (group size 64). Routed experts use the bit width below; attention and the shared expert are 4-bit, embeddings and LM head 8-bit, and the MoE router is unquantized. Expect quality differences vs. the original, growing at lower bit widths.

Variant Bits/weight Size Peak memory (short / 4k-token prompt) Perplexity DE / EN* Fits
4-bit 4.54 41 GiB 44 / 48 GB 12.77 / 16.47 64 GB+ Macs
3-bit (this repo) 3.57 33 GiB 35 / 39 GB 13.17 / 16.31 48 GB+ Macs
2-bit 2.61 24 GiB 26 / 30 GB 14.01 / 17.46 36 GB+ Macs

*First 4,096 tokens of the German "Kolibris" / English "Hummingbird" Wikipedia articles; lower is better; a small sample, so only roughly indicative. Generation runs at ~52–56 tok/s on an M1 Max for all three.

On a 48 GB Mac, raise the GPU memory limit for the 3-bit variant: sudo sysctl -w iogpu.wired_limit_mb=40960 (resets on reboot). Don't run another large model at the same time.

Relative weight-reconstruction RMSE vs. the dequantized FP8 weights: routed experts 0.194 (3-bit) / 0.400 (2-bit) / ~0.10 (4-bit); attention + shared expert 0.098.

Without reasoning_effort, the chat template defaults to high reasoning.

Usage

Native mlx-lm support

mlx-lm PR #1945

Support for the kolibri1 architecture is proposed upstream in ml-explore/mlx-lm#1945; the badge above shows its current state. The PR adds the same kolibri1.py that ships in this repo, so these weights load unchanged. Once it is merged and included in an mlx-lm release, the workaround below is no longer needed:

pip install -U mlx-lm
mlx_lm.chat --model velaia/Kolibri-1-MLX-3bit --temp 1.0 --top-p 0.97
mlx_lm.generate --model velaia/Kolibri-1-MLX-3bit --prompt "Hallo!" \
  --chat-template-config '{"reasoning_effort": "low"}'

(Before a release, you can install mlx-lm from GitHub main once the PR is merged: pip install -U git+https://github.com/ml-explore/mlx-lm.)

Until then: bundled launcher

mlx-lm does not ship the kolibri1 architecture yet. This repo includes it as kolibri1.py (ported from Aleph Alpha's Apache-2.0 vLLM plugin), plus run.py, a small launcher that registers it and then runs the normal mlx-lm commands (generate, chat, server, ...).

Setup

pip install -U mlx-lm huggingface_hub
hf download velaia/Kolibri-1-MLX-3bit --local-dir Kolibri-1-MLX-3bit

One-off prompt

python Kolibri-1-MLX-3bit/run.py generate --model Kolibri-1-MLX-3bit \
  --prompt "Erkläre kurz den Unterschied zwischen Wetter und Klima." \
  --chat-template-config '{"reasoning_effort": "low"}' \
  --temp 1.0 --top-p 0.97 --top-k 128 --max-tokens 1024
  • reasoning_effort: none | low | medium | high. Without it, the chat template uses high, which thinks at length.
  • --temp 0 gives deterministic output, useful for comparing variants.
  • Tokens/s and peak memory are printed at the end.

Interactive chat

python Kolibri-1-MLX-3bit/run.py chat --model Kolibri-1-MLX-3bit --temp 1.0 --top-p 0.97 --max-tokens 2048

Type q to quit, h for help; --system-prompt "..." sets a system prompt. chat has no reasoning-effort or top-k option, so it always reasons at high. Use generate to test other effort levels.

OpenAI-compatible server

python Kolibri-1-MLX-3bit/run.py server --model Kolibri-1-MLX-3bit \
  --temp 1.0 --top-p 0.97 --top-k 128 \
  --chat-template-args '{"reasoning_effort": "medium"}'

Listens on http://localhost:8080/v1. Clients can also pass chat_template_kwargs per request.

Python

import sys, importlib.util
from huggingface_hub import snapshot_download

path = snapshot_download("velaia/Kolibri-1-MLX-3bit")
spec = importlib.util.spec_from_file_location("mlx_lm.models.kolibri1", f"{path}/kolibri1.py")
mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod)
sys.modules["mlx_lm.models.kolibri1"] = mod  # register before load()

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tok = load(path)
prompt = tok.apply_chat_template([{"role": "user", "content": "Hallo! Wer bist du?"}],
    add_generation_prompt=True, tokenize=False, reasoning_effort="low")
print(generate(model, tok, prompt, max_tokens=1024, verbose=True,
               sampler=make_sampler(temp=1.0, top_p=0.97, top_k=128)))

Notes: a [transformers] ... model of type kolibri1 warning on load is harmless. Recommended sampling (from the original release): temperature 1.0, top_p 0.97, top_k 128. Don't load two large models at the same time.

Provenance

  • Source: Aleph-Alpha/Kolibri-1 @ e52eb4627d11516b0c01de49210ab5a4e4061444 (FP8 release)
  • mlx 0.32.3, mlx-lm 0.32.0; converted on an Apple M1 Max (64 GB)
  • Recipe: kolibri1.py (architecture), run.py (CLI launcher) and convert.py (mixed-precision predicate): python -m kolibri_mlx.convert --hf-path Kolibri-1 --mlx-path out --expert-bits 3 --dense-bits 4

License & responsible use

Weights: Apache 2.0, see LICENSE (copied unchanged from the original repo). Per the original model card, the license applies only to the weights and configuration files; Aleph Alpha retains all rights to its code, architecture and training methods. Please follow the original card's responsible-use guidance (no unlawful use, no practices prohibited by Art. 5 EU AI Act, no military/nuclear applications) and see it for evaluations, limitations and training details.

Downloads last month
484
Safetensors
Model size
78B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for velaia/Kolibri-1-MLX-3bit

Quantized
(18)
this model