Qwen3.8-Flash-Next, INT4/INT6 mixed (AutoRound), for vLLM on 24x2GB, 32x2GB, 24x3GB, 24x4GB GPU setup
Update: A new, better quant is now available on new-main. It uses exactly the same model structure and bit allocation as main. See Model Evaluations.
Update: Quant for 4-GPU users is now uploaded. See Have four 24 GB GPUs?
Update: Two 24 GB GPUs and two 32GB GPUs are now supported (vllm-patch/2x24GB/), with some experts in host RAM. On two 24 GB cards: 262,144 tokens x 2 requests, or 524,288 x 2 with YaRN. The build directories were renamed from 3x3090 to 3x24GB and so on. See vllm-patch/README.md.
This repository has mixed-precision AutoRound quantizations of Qwen3.8-Flash-Next (main, its recalibrated version new-main, and 4x3090 for four GPUs), and the vLLM patch that runs it on three 24 GB cards with 262,144 tokens of context x 4 requests with text and images. Two cards also work, with some experts in host RAM.
new-main uses the same bit allocation as main, with new calibration data generated by a higher-precision version of this model. It works with the builds for main, including 3x24GB and 4x24GB/main/. The branch contains weights only; download the patches and build files from main.
hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision new-main --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-new-main
Measured on 3x RTX 3090 (PCIe 4.0 x16) with 125 GiB of RAM. 2x24GB ran on two of the three cards:
3x24GB |
3x24GB+MTP |
2x24GB |
|
|---|---|---|---|
| VRAM, weights per GPU (vision tower included) | 21.20 / 20.44 / 21.05 GiB | 21.20 / 21.71 / 21.18 GiB | 18.8 / 18.6 GiB, plus 13 GiB per card of experts in host RAM |
| Context | 262,144 tokens x 4 requests | 262,144 tokens x 2 requests | 262,144 tokens x 2 requests, or 524,288 x 2 (YaRN) |
| Host RAM | ~85 GiB (67.8 GiB pinned K/V pool) | 41.2 GiB pinned K/V pool | ~77 GiB (40.0 GiB pinned K/V pool, 26 GiB experts) |
| Disk | 95.37 GiB table file on NVMe | same | same |
| Decode, 1 request, short prompt | 95-99 tok/s | 117-141 tok/s | 81-90 tok/s |
| Decode, 1 request, long context (8k-160k) | 83-96 tok/s | 110-116 tok/s | 84-87 tok/s |
| Decode, 2 requests | 188 tok/s total | 190 tok/s total | 150 tok/s total (135 with two 224k-token prompts) |
| Decode, 4 requests | 243-245 tok/s total | 188 tok/s total | 2 at a time as shipped |
| Prefill | ~3,700 tok/s at 248k, 5,913-5,947 tok/s at 39k | 5,613-5,809 tok/s at 39k | ~2,000 tok/s at 516k, ~2,300 at 256k, 2,494-2,499 at 39k |
All three builds support prefix caching and image input.
Model Evaluations
KLD
Lower KLD means closer agreement with BF16 on this dataset.
| Quant | Body bits per weight | Mean KLD | Median KLD |
|---|---|---|---|
AutoRound main |
4.178 | 0.014705 | 0.000872 |
AutoRound new-main |
4.178 | 0.011680 | 0.000620 |
The AutoRound points were measured with qbench against the same BF16 reference, scoring response positions on the same English trace (14,496 input + 45,588 output tokens).
The Unsloth GGUF points are reference results from Turboderp's published qbench chart, which uses the same trace. These points compare closeness to BF16 on this trace and they are not a comparison of task benchmark scores.
Plot rendered with Turboderp's qbench plotting code.
Benchmarks
The task benchmark scores below are for main.
| Benchmark | This quant | Official (BF16) | ± 1 SE |
|---|---|---|---|
| IFBench, prompt-level loose | 81.0 | 81.3 | 2.3 |
| GPQA Diamond | 90.4 | 91.7 | 2.1 |
| LiveCodeBench v6, pass@1 | 92.4 | 91.9 | 2.3 |
The official figures come from the official model card.
- IFBench: 300 single-turn prompts. Scoring used the official
evaluation_lib, applied to the answer after the reasoning was removed. Strict scores were 73.3 (prompt level) and 76.2 (instruction level). Loose instruction level was 83.4. - GPQA Diamond: 198 questions with the simple-evals prompt. The answer options were shuffled with a fixed seed, and the grader read the last
Answer: Xline. Per domain: physics 95.3, chemistry 87.1, biology 84.2. - LiveCodeBench v6: 131 problems dated 2025-02-01 to 2025-04-06 (31 easy, 39 medium, 61 hard). The official label is "25.02-25.05", but
livecodebench/code_generation_litehas no problems after 2025-04-06. The prompts used the genericlcb_runnertemplate, and the code was taken from the last fenced block. Per difficulty: easy 100, medium 94.9, hard 86.9.
Pick a build
| Directory | GPUs | Context | Use it for |
|---|---|---|---|
vllm-patch/3x24GB/ |
3 x 24 GB | 262,144 x 4 | Default, no MTP. Good for throughput |
vllm-patch/3x24GB+MTP/ |
3 x 24 GB | 262,144 x 2 | For one user. One stream is 18-43% faster, 4 streams are 22% slower. Good for latency |
vllm-patch/2x24GB/ |
2 x 24 GB | 262,144 x 2, or 1 x 524,288 (YaRN). Up to x 4 with more host RAM (see vllm-patch/README.md) |
For two cards. Decode 81-90 tok/s, prefill ~2,500 tok/s, ~77 GB free host RAM, PCIe 4.0 x16 |
vllm-patch/2x32GB/ |
2 x 32 GB | 262,144 x 1 as shipped, up to x 4 with more host RAM (see vllm-patch/README.md) |
For RTX 5090 and the like, 3 GiB of experts per card in host RAM. Untested on 32 GB cards |
vllm-patch/4x24GB/ |
4 x 24 GB | 262,144 x 4 (calculated) | For four cards, with MTP. Maybe more throughput than three. Start it from 4x24GB/4x3090/ with the weights on the 4x3090 branch (recommended), or from 4x24GB/main/ with the main weights |
The 24 GB builds were tested on RTX 3090 only.
Have four 24 GB GPUs?
Download the 4x3090 branch instead of main, build the image in vllm-patch/4x24GB/, and start it from vllm-patch/4x24GB/4x3090/. It is a different quantization made for four cards, which have more room for weights. Routed experts are INT4 gs32 (INT8 gs64 in layers 0, 1, 16, 30, 45, 46, 47), and linear_attn, QSA attention and the shared expert are INT8 gs64 instead of INT6. Everything else is the same as main. The numbers on this page are for the main weights on 3x RTX 3090.
hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision 4x3090 --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090
If the 4x3090 weights do not work for you, the main weights also run on four cards: start the same image from vllm-patch/4x24GB/main/ instead. Its layer split is set for main.
Requirements
- 2, 3 or 4 x 24 GB NVIDIA GPUs (tested on RTX 3090), or 2 x 32 GB (untested).
- ~85 GB free host RAM at the shipped settings, 67.8 GiB of it a pinned pool (
3x24GB+MTP: 41.2 GiB,4x24GB: 96.7 GiB, ~115 GB free in total;2x24GB: ~77 GB free in total). Less RAM works with less context, see "Settings". - 96 GB free on a local NVMe drive for the table file. Not a hard disk, not NFS.
- podman or docker with the NVIDIA container toolkit.
- The vLLM image
vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96(0.29.1rc1.dev47+gdc36fcce9). The patches may not apply to other nightlies. - PCIe 4.0 x16 for each card with
2x24GBand2x32GB: they read experts from host RAM over PCIe for every token. The 3- and 4-card builds do not need it. They read less over PCIe (the K/V cache on long prompts), so a slower link costs less; how much was not measured.
Quick start (3x24GB)
cd vllm-patch/3x24GB
podman build -t flash-next-vllm:local -f Dockerfile \
--build-arg BASE=docker.io/vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 .
MODEL_DIR=/srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
PLE_TABLE=/srv/nvme/flash-next/ngram_table.bin \
podman compose up -d
MODEL_DIR: this repository.PLE_TABLE: a file path on NVMe. It does not have to exist. The first start writes it (95.37 GiB) and is slow. Later starts are fast.
The server listens on 127.0.0.1:8000 (OpenAI API), model name Qwen3.8-Flash-Next.
For the other directories, see vllm-patch/README.md.
Settings
Change these in compose.yaml. The comments there give the numbers.
| Setting | Shipped | If you change it |
|---|---|---|
--kv-cache-memory-bytes |
783000000 |
Sets context and host RAM together. 550000000: 262,144 x 2.8, 47.5 GiB pinned. |
--max-num-seqs |
4 |
Requests that run at once. Raise it only with --kv-cache-memory-bytes |
--max-model-len |
262144 |
Maximum context per request. Do not go above 262144 |
--enable-prefix-caching |
on | Off (--no-enable-prefix-caching): 45% more context for the same RAM, but repeated prompts are prefilled again |
--limit-mm-per-prompt, --mm-processor-kwargs |
4 images, 2 MP | Replace both with --language-model-only for a text-only server. Larger images than 2 MP were not tested |
--max-num-batched-tokens |
512 |
1568: 470 MiB more VRAM, same speed. 256: prefill takes 32% longer |
--host, --port |
127.0.0.1, 8000 |
Where the server listens |
Do not remove PYTORCH_CUDA_ALLOC_CONF, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS, ulimits, or --mamba-cache-mode align. Long prompts fail or the server does not start without them. The environment switches that turn single patches off are in vllm-patch/README.md.
Check that it started correctly
podman logs flash-next | grep "GPU KV cache size" # 1,048,576 tokens, 4.00x
podman logs flash-next | grep "QSA host KV" # 244 blocks x 12144 tokens ... 67.8 GiB, 12 times
podman logs flash-next | grep "QSA staging arena" # no warning. A warning means slow prefill
podman logs flash-next | grep "PLE page prefetch" # "process_madvise". "per-range madvise": add cap_add SYS_PTRACE
podman logs flash-next | grep "fused MoE decode" # "enabled"
Setting attention block size to 1568 in the log means the image is not patched correctly.
2x24GB shows 808,901 tokens, 1.54x (counted per 524,288-token request) and 216 blocks x 8096 tokens ... 40.0 GiB 12 times. It must not show Expert table over VRAM and host RAM unavailable: with that warning, decode runs at ~60 tok/s and the start takes ~30 min. Take a larger expert file or a smaller prefill chunk (table in 2x24GB/compose.yaml).
3x24GB+MTP shows 170 blocks x 9776 tokens ... 41.2 GiB 13 times. To see that MTP works, check that spec_decode_num_accepted_tokens_total in curl -s localhost:8000/metrics goes up. Do not sum all spec_decode_* lines: the *_created lines are timestamps.
Things you should know
- 262,144 tokens is
max_position_embeddings. Above it the model needs YaRN, or it fails.2x24GBships with YaRN factor 2.0 (one request of up to 524,288 tokens). With it, a needle at 10, 50 and 90% of 302k-, 406k-, 505k- and 516k-token prompts was found, and short texts scored the same (average log-likelihood -0.14%, within run-to-run noise). The other builds ship without YaRN; to add it, see "Context on two cards" invllm-patch/README.md. Up to 1M works the same way, but I don't recommend it unless you really need it. - The
3x24GBbuild doesn't support speculative decoding. Use3x24GB+MTPfor MTP. VLLM_QSA_KV_OFFLOAD=1requires TP=1, and the mmap PLE backend requires ETP=1.- TP is not implemented. I tried, but I didn't see any improvement.
What is quantized
Everything below uses compressed-tensors, pack-quantized, and symmetric group quantization.
| Group | Scheme | What it covers |
|---|---|---|
| A | INT4, gs128 | MoE routed experts (512 per layer, top-10). 58.0 GiB, 92% of the body |
| B | INT6, gs64 | linear_attn in and out projections, QSA q/k/v/o_proj, shared expert |
| C | INT8, gs64 | hyper-connection low-rank mixers |
| D | INT8, gs128 | lm_head, embed_tokens, PLE key_proj and value_proj, indexer index_qk_proj |
| E | BF16 | router, vision tower, PLE n-gram table, MTP module |
3x24GB+MTP converts the MTP experts to INT4 in a separate directory (make_mtp_int4.py).
License
The model weights inherit the upstream Qwen Community License 1.0. See LICENSE in this directory.
The patches in every directory under vllm-patch/ are derivative works of vLLM, so they stay under the Apache License, Version 2.0, like vLLM itself. vllm-patch/LICENSE holds the full text, and vllm-patch/NOTICE names the vLLM files that each of them changes.
The Dockerfile and compose.yaml in each directory, and make_mtp_int4.py, contain no vLLM code. They are under the MIT license, in vllm-patch/LICENSE.MIT.
If you find these useful, make sure to like my repo.
- Downloads last month
- 1,138
