Qwen3.8-27B-MLA
Qwen3.8-27B with its attention KV cache compressed 2× by retrofitting Multi-head Latent Attention (MLA), without any training. In paired runs on SWE-bench Verified it resolved 0.421 against the original model's 0.395, while generating 12% fewer tokens.
Serving: this checkpoint uses a custom attention architecture. It runs in vLLM with the qwen-mla-vllm plugin (experimental; see Serving). It does not load in
transformersor stock vLLM. Includes the base model's vision tower and MTP head. FP8 build: Qwen3.8-27B-MLA-FP8.
| Qwen3.8-27B | Qwen3.8-27B-MLA | |
|---|---|---|
| KV cache per token (bf16, TP=1) | 64 KiB | 32 KiB (2.0×) |
| SWE-bench Verified (4 / 3 runs) | 0.395 | 0.421 |
| HELMET @64k (5 tasks, mean) | 77.43 | 77.93 |
| Output tokens on SWE-bench | 1.00× | 0.88× |
| Training | — | none |
Serving (vLLM plugin, experimental)
This checkpoint is served by qwen-mla-vllm, a vLLM plugin. Install it and vLLM finds the architecture automatically. No fork or extra flags needed.
pip install "git+https://github.com/sootaugur/qwen-mla-vllm" # installs vllm==0.27.1
vllm serve TelperionAI/Qwen3.8-27B-MLA --reasoning-parser qwen3 \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'
--speculative-config turns on MTP speculative decoding (optional). Images work as with the
base model; for text-only serving add --limit-mm-per-prompt '{"image": 0, "video": 0}'.
The plugin is experimental. This version has been verified on RTX PRO 6000 (Blackwell) at
TP=1 and TP=2, with vLLM 0.27.1 only (earlier versions also ran on H100), and has seen little
real-world use so far. The first start JIT-compiles the fast decode kernel, which needs nvcc
and ninja. Without them the plugin falls back to Triton kernels: correct, but slower. See the
plugin README for requirements, configuration and verification.
What this is
Qwen3.8-27B is a hybrid model: 48 Gated DeltaNet (linear attention) layers and 16 full-attention layers with grouped-query attention (24 query heads, 4 KV heads, head dim 256). Only the 16 full-attention layers keep a KV cache that grows with context, and those are what this release changes.
In each full-attention layer, the separate K and V projections are replaced by a shared low-rank latent that all 24 heads read, plus a small decoupled RoPE key. At inference the cache stores the latent instead of per-head keys and values:
- latent rank per layer: 256 / 768 / 1792 (mean 768), allocated by sensitivity;
- RoPE: 58 of 64 rotary dimensions kept per KV head. The three lowest-frequency rotary pairs turn only 3.4°, 2.1° and 1.2° across 131k tokens, so they carry content rather than position. They are not discarded: they stop rotating and move into the latent. This makes each cache row exactly 512 / 1024 / 2048 values with no memory-page padding (2.000× served instead of 1.882×), at a measured cost of 0.49% of attention logits at 131k tokens;
- cache row per layer: 512 / 1024 / 2048 values (vs 2048 for the original), 2.0× smaller on average.
The projections are fit training-free: whitened SVD of the original K/V weights under the covariance of each layer's inputs, measured on a calibration corpus. Nothing else in the model changes. Embeddings, MLPs, all 48 linear-attention layers, and the query projections (up to a fixed permutation) are byte-identical to the base model, as are the vision tower and the MTP head.
Files. Full bf16 weights (the complete model, not a delta, including the base model's vision
tower and MTP head), tokenizer, chat template, preprocessor configs and the base model's
generation_config.json. config.json declares Qwen3_5MLAForConditionalGeneration, the
latent-cache architecture the vLLM plugin serves; the mla_* fields record the per-layer ranks and RoPE layout.
Calibration data (this matters: it is the only thing the fit sees):
- long-context QA: HotpotQA train passages (~62k tokens each) with the base model's own answers;
- agentic coding traces: community sessions from Qwen3.8-27B;
- software engineering: the base model's own reasoning and patches on SWE-bench train (22 repositories, none shared with SWE-bench Verified).
Moving the calibration to the model's own outputs (on-policy QA and SWE-bench train traces, replacing plain retrieval passages) cut the student's excess NLL on SWE-bench conversations by a third.
Is SWE-bench data in calibration a contamination risk? No training happens: no gradients and no labels. Calibration measures one statistic, the covariance of each attention layer's inputs, and the latent keeps the directions with the most variance under it. The SWE-bench traces come from the train split, and their 22 repositories are disjoint from the 12 in SWE-bench Verified. No Verified task, issue, or repository is seen. What calibration can do is specialise the latent toward a domain's activations, which is the intended effect.
Results
All comparisons are paired: same prompts, same harness, same serving settings, run on the same machine as the base model.
SWE-bench Verified
evalscope, BM25-retrieved context (princeton-nlp/SWE-bench_bm25_40K), thinking on,
temperature 1.0, 500 instances, repeated runs.
| runs | resolved | vs base (paired, 95% CI) | output tokens | |
|---|---|---|---|---|
| Qwen3.8-27B | 4 | 0.395 | — | 1.00× |
| Qwen3.8-27B-MLA | 3 | 0.421 | +2.6 pts [+0.8, +4.5] | 0.88× |
The difference comes from fewer patches that fail to apply and fewer missing patches; the harness error rate is identical. Results cover SWE-bench Verified and HELMET (64k) only; we make no claims about other domains.
HELMET long-context (64k tokens)
KILT RAG (NQ, TriviaQA, HotpotQA, PopQA) and JSON key-value retrieval, 100 questions per task, chat template, thinking off.
| NQ | TriviaQA | HotpotQA | PopQA | JSON KV | mean | |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | 58.8 | 93.0 | 68.7 | 66.7 | 100 | 77.43 |
| Qwen3.8-27B-MLA | 63.2 | 92.2 | 68.7 | 65.7 | 100 | 77.93 |
Within noise of the base model at n=100 per task.
Fidelity to the base model
Teacher-forced NLL on text the base model generated, excess over the base model (lower is closer):
| held-out set | excess NLL (nats/token) |
|---|---|
| agentic coding traces | +0.019 |
| SWE-bench Verified conversations | +0.023 |
Generation behaviour: none of 1,500 SWE-bench generations hit the length limit; 1 had more than 20% repeated text (base model: 0).
KV cache and tensor parallelism
MLA's latent is shared by every head, so under tensor parallelism it is replicated on each GPU, while the base model's 4 KV heads are split across GPUs. Per-GPU KV per token:
| TP=1 | TP=2 | TP=4 | TP=8 | |
|---|---|---|---|---|
| Qwen3.8-27B | 64 KiB | 32 KiB | 16 KiB | 16 KiB |
| Qwen3.8-27B-MLA | 32 KiB (2.0×) | 28 KiB (1.14×) | 26 KiB (0.62×) | 26 KiB (0.62×) |
(The base model's 4 KV heads cannot be split beyond TP=4, so its per-GPU cost stops falling.)
This model is for single-GPU (TP=1) serving, where it doubles how many long-context requests fit in memory (measured: 1.75× served KV capacity in vLLM at TP=1, including the linear-attention state). For TP=2 see Qwen3.8-27B-GLA-g2, which splits its latent across GPUs.
Speed
RTX PRO 6000 (Blackwell), bf16 weights and KV, vLLM 0.27.1 with the plugin, 16k-token prompts, 256 decode steps. "Max" is each model's own largest batch that fits in KV memory.
| TP=1 | max concurrent 16k sequences | batch 1 (ms/token) | batch 16 (ms/step) | throughput at max batch |
|---|---|---|---|---|
| Qwen3.8-27B | 24 | 38.4 | 54.4 | 378 tok/s |
| Qwen3.8-27B-MLA | 40 | 40.0 | 56.1 | 505 tok/s |
Single-stream latency is within 5% of the base model, and the larger batch gives 1.34× the peak throughput. Prefill matches the base model to within 3%. At TP=2 the replicated latent leaves this model slightly behind the base model (82 vs 92 sequences, 988 vs 1,043 tok/s). Use GLA-g2 there.
Speculative decoding (MTP)
The base model's MTP head is included unchanged, and it drafts for the retrofit as well as for the base model: mean accepted tokens per step at k=3 are 1.86 vs 1.82 (retrofit vs base). Decode throughput, tokens/s at 1 / 8 / 32 concurrent requests (64 mixed prompts, thinking on, Qwen's recommended sampling, bf16, RTX PRO 6000):
| no MTP | MTP, k=3 | speedup | |
|---|---|---|---|
| Qwen3.8-27B, TP=1 | 26 / 155 / 547 | 57 / 326 / 972 | 2.16× / 2.10× / 1.78× |
| Qwen3.8-27B-MLA, TP=1 | 25 / 129 / 548 | 47 / 227 / 858 | 1.85× / 1.76× / 1.57× |
The retrofit's speedup is smaller because a verify step currently reads the latent cache once per draft token in the plugin; a single causal pass per request is planned.
Vision
The vision tower is the base model's, unchanged. Paired against the base model on the same items, greedy, thinking off (300 ChartQA test questions, relaxed accuracy; 300 DocVQA validation questions, ANLS):
| ChartQA | DocVQA | |
|---|---|---|
| Qwen3.8-27B | 89.7 | 97.1 |
| Qwen3.8-27B-MLA | 88.0 (−1.7 [−4.3, +1.0]) | 97.3 (+0.1 [−1.0, +1.3]) |
Neither difference is significant; 90–95% of answers are identical to the base model's.
Roadmap
- vLLM plugin: qwen-mla-vllm (experimental)
- Optimized decode kernels (patched FlashInfer MLA, two-pass Triton for the widest latent)
- FP8: Qwen3.8-27B-MLA-FP8
- Vision tower and MTP head (speculative decoding) included
- 4-bit variant
- Method and code release: calibration corpus builders, covariance tooling, the exporter, and evaluation harness
Limitations
- Serving requires the experimental vLLM plugin (vLLM 0.27.1 only, for now).
- Evaluated on SWE-bench Verified and HELMET (64k); other domains and much longer contexts are untested.
- Results are from one harness and sampling configuration; absolute scores differ across harnesses. The paired differences are the meaningful numbers.
- The compression only helps at TP=1 (see above).
License and attribution
Released under Apache-2.0, the license of the base model. Derived from Qwen/Qwen3.8-27B by the Qwen team. This is an independent project, not affiliated with or endorsed by the Qwen team.
Citation
@misc{telperion2026qwen38mla,
title = {Qwen3.8-27B-MLA: training-free Multi-head Latent Attention retrofit},
author = {TelperionAI},
year = {2026},
url = {https://huggingface.co/TelperionAI/Qwen3.8-27B-MLA}
}
- Downloads last month
- 636