Instructions to use mlx-community/plamo-2-translate-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/plamo-2-translate-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/plamo-2-translate-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use mlx-community/plamo-2-translate-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/plamo-2-translate-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/plamo-2-translate-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/plamo-2-translate-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Atomic Chat
PLaMo 2 Translate — MLX 8bit
A standalone MLX conversion of pfnet/plamo-2-translate for Apple Silicon. Affine 8-bit quantization, group size 64, with BF16 floating parameters and activations. The recurrent SSM state and the exponentiation of A_log use FP32. This model specializes in translation.
October 7, 2026 update
Regenerated from source revision cae8da342a3e051ed69f90ce24c23eacff908732, using MLX 0.31.1 and mlx-lm 0.31.2. The prior repository revision remains available at 84f3a775383c115fa91d82c36fc813de85b755d7.
The bundled modeling_mlx_plamo2.py is selected by config.json's model_file. It applies the checkpoint's RoPE bases (1,000,000 here) and PLaMo's unbounded softplus(dt + bias) SSM time steps. The prompt includes BOS, and EOS IDs include both 1 and the operation delimiter 4. It uses the upstream convolution; no experimental fusion or all-FP16 conversion is enabled.
The template supports ordinary user/assistant messages and explicit PLaMo input lang=... / output lang=... messages. Both forms produce the same tested prompt when language labels are specified. The vocabulary and source training are unchanged.
Usage
Tested with Python 3.13 on Apple Silicon:
pip install 'mlx==0.31.1' 'mlx-lm==0.31.2' 'transformers>=4.46,<5' numba
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load(
"mlx-community/plamo-2-translate-8bit",
tokenizer_config={"trust_remote_code": True},
)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Write the text to be translated here."}],
tokenize=True,
add_generation_prompt=True,
source_language="English",
target_language="Japanese",
)
print(generate(
model, tokenizer, prompt=prompt, sampler=make_sampler(temp=0),
max_tokens=2048, prefill_step_size=512,
))
Omitting language labels retains English/Japanese automatic direction. Explicit labels are recommended for reproducible comparisons. Inference corrections require a runtime that honors config.model_file, as mlx-lm 0.31.2 does. Runtimes that instantiate their own generic PLaMo 2 implementation may differ. This repository targets MLX; the included original Transformers source does not establish PyTorch compatibility for these MLX parameter layouts.
Validation
The full supplied English-to-Japanese corporate-profile example was validated on an Apple M1 Max with 32 GPU cores and 64 GB memory. Greedy decoding, prompt length 364, prefill chunk 512, generation limit 2048, two warmed standalone runs.
| Check | Result |
|---|---|
| Full output versus corrected direct inference at the same precision | Exact byte match in both runs |
| Structured versus natural-message prompt | Exact token match |
| Completion | EOS; all 8 paragraphs present |
| Required names and phrase | Preferred Networks, both founders, Learn or Die retained |
| chrF against the supplied Japanese reference | 73.94 |
| Weight files | 10.12 GB |
| Tensor dtypes | BF16, U32 |
See evaluation.json for the exact runs, versions and source revision. The machine was shared with another workload; elapsed times and throughput are observational, not a controlled speed comparison. Complete streaming and non-streaming CLI outputs were also checked against the same direct-inference output before publication.
This is one example, not a broad quality benchmark. The output differs from the supplied Japanese reference and is not guaranteed to match BF16. Output equality is with the corresponding same-precision corrected implementation; it does not imply equality across quantization levels. The previous public 8bit weights were not included in this comparison, so these checks do not establish superiority over that release.
Source, license, and limitations
PLaMo Translation Model was developed by Preferred Networks. See the technical announcement and press release.
The original PLaMo community license applies. License files are included in LICENSE; a Japanese version is also available. Consult the source model's license and commercial-use contact form as applicable.
This model is not instruction-tuned for general dialogue. It may produce inaccurate, biased, or otherwise unsuitable translations; the source model's documented limitations continue to apply.
The source model was trained under “Research and Development Project of the Enhanced Infrastructures for Post 5G Information and Communication System” (JPNP 20017), subsidized by NEDO.
- Downloads last month
- 43
8-bit