PLaMo 2 Translate — MLX 8bit

A standalone MLX conversion of pfnet/plamo-2-translate for Apple Silicon. Affine 8-bit quantization, group size 64, with BF16 floating parameters and activations. The recurrent SSM state and the exponentiation of A_log use FP32. This model specializes in translation.

October 7, 2026 update

Regenerated from source revision cae8da342a3e051ed69f90ce24c23eacff908732, using MLX 0.31.1 and mlx-lm 0.31.2. The prior repository revision remains available at 84f3a775383c115fa91d82c36fc813de85b755d7.

The bundled modeling_mlx_plamo2.py is selected by config.json's model_file. It applies the checkpoint's RoPE bases (1,000,000 here) and PLaMo's unbounded softplus(dt + bias) SSM time steps. The prompt includes BOS, and EOS IDs include both 1 and the operation delimiter 4. It uses the upstream convolution; no experimental fusion or all-FP16 conversion is enabled.

The template supports ordinary user/assistant messages and explicit PLaMo input lang=... / output lang=... messages. Both forms produce the same tested prompt when language labels are specified. The vocabulary and source training are unchanged.

Usage

Tested with Python 3.13 on Apple Silicon:

pip install 'mlx==0.31.1' 'mlx-lm==0.31.2' 'transformers>=4.46,<5' numba
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load(
    "mlx-community/plamo-2-translate-8bit",
    tokenizer_config={"trust_remote_code": True},
)
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Write the text to be translated here."}],
    tokenize=True,
    add_generation_prompt=True,
    source_language="English",
    target_language="Japanese",
)
print(generate(
    model, tokenizer, prompt=prompt, sampler=make_sampler(temp=0),
    max_tokens=2048, prefill_step_size=512,
))

Omitting language labels retains English/Japanese automatic direction. Explicit labels are recommended for reproducible comparisons. Inference corrections require a runtime that honors config.model_file, as mlx-lm 0.31.2 does. Runtimes that instantiate their own generic PLaMo 2 implementation may differ. This repository targets MLX; the included original Transformers source does not establish PyTorch compatibility for these MLX parameter layouts.

Validation

The full supplied English-to-Japanese corporate-profile example was validated on an Apple M1 Max with 32 GPU cores and 64 GB memory. Greedy decoding, prompt length 364, prefill chunk 512, generation limit 2048, two warmed standalone runs.

Check Result
Full output versus corrected direct inference at the same precision Exact byte match in both runs
Structured versus natural-message prompt Exact token match
Completion EOS; all 8 paragraphs present
Required names and phrase Preferred Networks, both founders, Learn or Die retained
chrF against the supplied Japanese reference 73.94
Weight files 10.12 GB
Tensor dtypes BF16, U32

See evaluation.json for the exact runs, versions and source revision. The machine was shared with another workload; elapsed times and throughput are observational, not a controlled speed comparison. Complete streaming and non-streaming CLI outputs were also checked against the same direct-inference output before publication.

This is one example, not a broad quality benchmark. The output differs from the supplied Japanese reference and is not guaranteed to match BF16. Output equality is with the corresponding same-precision corrected implementation; it does not imply equality across quantization levels. The previous public 8bit weights were not included in this comparison, so these checks do not establish superiority over that release.

Source, license, and limitations

PLaMo Translation Model was developed by Preferred Networks. See the technical announcement and press release.

The original PLaMo community license applies. License files are included in LICENSE; a Japanese version is also available. Consult the source model's license and commercial-use contact form as applicable.

This model is not instruction-tuned for general dialogue. It may produce inaccurate, biased, or otherwise unsuitable translations; the source model's documented limitations continue to apply.

The source model was trained under “Research and Development Project of the Enhanced Infrastructures for Post 5G Information and Communication System” (JPNP 20017), subsidized by NEDO.

Downloads last month
43
Safetensors
Model size
10B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/plamo-2-translate-8bit

Base model

pfnet/plamo-2-8b
Quantized
(5)
this model

Collection including mlx-community/plamo-2-translate-8bit