Qwen3.8-Flash-Next, INT4/INT6 mixed (AutoRound), for vLLM on 24x2GB, 32x2GB, 24x3GB, 24x4GB GPU setup

Article, Article #2

Update: A new, better quant is now available on new-main. It uses exactly the same model structure and bit allocation as main. See Model Evaluations.

Update: Quant for 4-GPU users is now uploaded. See Have four 24 GB GPUs?

Update: Two 24 GB GPUs and two 32GB GPUs are now supported (vllm-patch/2x24GB/), with some experts in host RAM. On two 24 GB cards: 262,144 tokens x 2 requests, or 524,288 x 2 with YaRN. The build directories were renamed from 3x3090 to 3x24GB and so on. See vllm-patch/README.md.

This repository has mixed-precision AutoRound quantizations of Qwen3.8-Flash-Next (main, its recalibrated version new-main, and 4x3090 for four GPUs), and the vLLM patch that runs it on three 24 GB cards with 262,144 tokens of context x 4 requests with text and images. Two cards also work, with some experts in host RAM.

new-main uses the same bit allocation as main, with new calibration data generated by a higher-precision version of this model. It works with the builds for main, including 3x24GB and 4x24GB/main/. The branch contains weights only; download the patches and build files from main.

hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision new-main --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-new-main

Measured on 3x RTX 3090 (PCIe 4.0 x16) with 125 GiB of RAM. 2x24GB ran on two of the three cards:

3x24GB 3x24GB+MTP 2x24GB
VRAM, weights per GPU (vision tower included) 21.20 / 20.44 / 21.05 GiB 21.20 / 21.71 / 21.18 GiB 18.8 / 18.6 GiB, plus 13 GiB per card of experts in host RAM
Context 262,144 tokens x 4 requests 262,144 tokens x 2 requests 262,144 tokens x 2 requests, or 524,288 x 2 (YaRN)
Host RAM ~85 GiB (67.8 GiB pinned K/V pool) 41.2 GiB pinned K/V pool ~77 GiB (40.0 GiB pinned K/V pool, 26 GiB experts)
Disk 95.37 GiB table file on NVMe same same
Decode, 1 request, short prompt 95-99 tok/s 117-141 tok/s 81-90 tok/s
Decode, 1 request, long context (8k-160k) 83-96 tok/s 110-116 tok/s 84-87 tok/s
Decode, 2 requests 188 tok/s total 190 tok/s total 150 tok/s total (135 with two 224k-token prompts)
Decode, 4 requests 243-245 tok/s total 188 tok/s total 2 at a time as shipped
Prefill ~3,700 tok/s at 248k, 5,913-5,947 tok/s at 39k 5,613-5,809 tok/s at 39k ~2,000 tok/s at 516k, ~2,300 at 256k, 2,494-2,499 at 39k

All three builds support prefix caching and image input.

Model Evaluations

KLD

Lower KLD means closer agreement with BF16 on this dataset.

qbench mean KLD: main, new-main and Unsloth GGUF

Quant Body bits per weight Mean KLD Median KLD
AutoRound main 4.178 0.014705 0.000872
AutoRound new-main 4.178 0.011680 0.000620

The AutoRound points were measured with qbench against the same BF16 reference, scoring response positions on the same English trace (14,496 input + 45,588 output tokens).

The Unsloth GGUF points are reference results from Turboderp's published qbench chart, which uses the same trace. These points compare closeness to BF16 on this trace and they are not a comparison of task benchmark scores.

Plot rendered with Turboderp's qbench plotting code.

Benchmarks

The task benchmark scores below are for main.

Benchmark This quant Official (BF16) ± 1 SE
IFBench, prompt-level loose 81.0 81.3 2.3
GPQA Diamond 90.4 91.7 2.1
LiveCodeBench v6, pass@1 92.4 91.9 2.3

The official figures come from the official model card.

  • IFBench: 300 single-turn prompts. Scoring used the official evaluation_lib, applied to the answer after the reasoning was removed. Strict scores were 73.3 (prompt level) and 76.2 (instruction level). Loose instruction level was 83.4.
  • GPQA Diamond: 198 questions with the simple-evals prompt. The answer options were shuffled with a fixed seed, and the grader read the last Answer: X line. Per domain: physics 95.3, chemistry 87.1, biology 84.2.
  • LiveCodeBench v6: 131 problems dated 2025-02-01 to 2025-04-06 (31 easy, 39 medium, 61 hard). The official label is "25.02-25.05", but livecodebench/code_generation_lite has no problems after 2025-04-06. The prompts used the generic lcb_runner template, and the code was taken from the last fenced block. Per difficulty: easy 100, medium 94.9, hard 86.9.

Pick a build

Directory GPUs Context Use it for
vllm-patch/3x24GB/ 3 x 24 GB 262,144 x 4 Default, no MTP. Good for throughput
vllm-patch/3x24GB+MTP/ 3 x 24 GB 262,144 x 2 For one user. One stream is 18-43% faster, 4 streams are 22% slower. Good for latency
vllm-patch/2x24GB/ 2 x 24 GB 262,144 x 2, or 1 x 524,288 (YaRN). Up to x 4 with more host RAM (see vllm-patch/README.md) For two cards. Decode 81-90 tok/s, prefill ~2,500 tok/s, ~77 GB free host RAM, PCIe 4.0 x16
vllm-patch/2x32GB/ 2 x 32 GB 262,144 x 1 as shipped, up to x 4 with more host RAM (see vllm-patch/README.md) For RTX 5090 and the like, 3 GiB of experts per card in host RAM. Untested on 32 GB cards
vllm-patch/4x24GB/ 4 x 24 GB 262,144 x 4 (calculated) For four cards, with MTP. Maybe more throughput than three. Start it from 4x24GB/4x3090/ with the weights on the 4x3090 branch (recommended), or from 4x24GB/main/ with the main weights

The 24 GB builds were tested on RTX 3090 only.

Have four 24 GB GPUs?

Download the 4x3090 branch instead of main, build the image in vllm-patch/4x24GB/, and start it from vllm-patch/4x24GB/4x3090/. It is a different quantization made for four cards, which have more room for weights. Routed experts are INT4 gs32 (INT8 gs64 in layers 0, 1, 16, 30, 45, 46, 47), and linear_attn, QSA attention and the shared expert are INT8 gs64 instead of INT6. Everything else is the same as main. The numbers on this page are for the main weights on 3x RTX 3090.

hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound --revision 4x3090 --local-dir /srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090

If the 4x3090 weights do not work for you, the main weights also run on four cards: start the same image from vllm-patch/4x24GB/main/ instead. Its layer split is set for main.

Requirements

  • 2, 3 or 4 x 24 GB NVIDIA GPUs (tested on RTX 3090), or 2 x 32 GB (untested).
  • ~85 GB free host RAM at the shipped settings, 67.8 GiB of it a pinned pool (3x24GB+MTP: 41.2 GiB, 4x24GB: 96.7 GiB, ~115 GB free in total; 2x24GB: ~77 GB free in total). Less RAM works with less context, see "Settings".
  • 96 GB free on a local NVMe drive for the table file. Not a hard disk, not NFS.
  • podman or docker with the NVIDIA container toolkit.
  • The vLLM image vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 (0.29.1rc1.dev47+gdc36fcce9). The patches may not apply to other nightlies.
  • PCIe 4.0 x16 for each card with 2x24GB and 2x32GB: they read experts from host RAM over PCIe for every token. The 3- and 4-card builds do not need it. They read less over PCIe (the K/V cache on long prompts), so a slower link costs less; how much was not measured.

Quick start (3x24GB)

cd vllm-patch/3x24GB
podman build -t flash-next-vllm:local -f Dockerfile \
  --build-arg BASE=docker.io/vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 .
MODEL_DIR=/srv/models/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
PLE_TABLE=/srv/nvme/flash-next/ngram_table.bin \
podman compose up -d
  • MODEL_DIR: this repository.
  • PLE_TABLE: a file path on NVMe. It does not have to exist. The first start writes it (95.37 GiB) and is slow. Later starts are fast.

The server listens on 127.0.0.1:8000 (OpenAI API), model name Qwen3.8-Flash-Next.

For the other directories, see vllm-patch/README.md.

Settings

Change these in compose.yaml. The comments there give the numbers.

Setting Shipped If you change it
--kv-cache-memory-bytes 783000000 Sets context and host RAM together. 550000000: 262,144 x 2.8, 47.5 GiB pinned.
--max-num-seqs 4 Requests that run at once. Raise it only with --kv-cache-memory-bytes
--max-model-len 262144 Maximum context per request. Do not go above 262144
--enable-prefix-caching on Off (--no-enable-prefix-caching): 45% more context for the same RAM, but repeated prompts are prefilled again
--limit-mm-per-prompt, --mm-processor-kwargs 4 images, 2 MP Replace both with --language-model-only for a text-only server. Larger images than 2 MP were not tested
--max-num-batched-tokens 512 1568: 470 MiB more VRAM, same speed. 256: prefill takes 32% longer
--host, --port 127.0.0.1, 8000 Where the server listens

Do not remove PYTORCH_CUDA_ALLOC_CONF, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS, ulimits, or --mamba-cache-mode align. Long prompts fail or the server does not start without them. The environment switches that turn single patches off are in vllm-patch/README.md.

Check that it started correctly

podman logs flash-next | grep "GPU KV cache size"     # 1,048,576 tokens, 4.00x
podman logs flash-next | grep "QSA host KV"           # 244 blocks x 12144 tokens ... 67.8 GiB, 12 times
podman logs flash-next | grep "QSA staging arena"     # no warning. A warning means slow prefill
podman logs flash-next | grep "PLE page prefetch"     # "process_madvise". "per-range madvise": add cap_add SYS_PTRACE
podman logs flash-next | grep "fused MoE decode"      # "enabled"

Setting attention block size to 1568 in the log means the image is not patched correctly.

2x24GB shows 808,901 tokens, 1.54x (counted per 524,288-token request) and 216 blocks x 8096 tokens ... 40.0 GiB 12 times. It must not show Expert table over VRAM and host RAM unavailable: with that warning, decode runs at ~60 tok/s and the start takes ~30 min. Take a larger expert file or a smaller prefill chunk (table in 2x24GB/compose.yaml).

3x24GB+MTP shows 170 blocks x 9776 tokens ... 41.2 GiB 13 times. To see that MTP works, check that spec_decode_num_accepted_tokens_total in curl -s localhost:8000/metrics goes up. Do not sum all spec_decode_* lines: the *_created lines are timestamps.

Things you should know

  • 262,144 tokens is max_position_embeddings. Above it the model needs YaRN, or it fails. 2x24GB ships with YaRN factor 2.0 (one request of up to 524,288 tokens). With it, a needle at 10, 50 and 90% of 302k-, 406k-, 505k- and 516k-token prompts was found, and short texts scored the same (average log-likelihood -0.14%, within run-to-run noise). The other builds ship without YaRN; to add it, see "Context on two cards" in vllm-patch/README.md. Up to 1M works the same way, but I don't recommend it unless you really need it.
  • The 3x24GB build doesn't support speculative decoding. Use 3x24GB+MTP for MTP.
  • VLLM_QSA_KV_OFFLOAD=1 requires TP=1, and the mmap PLE backend requires ETP=1.
  • TP is not implemented. I tried, but I didn't see any improvement.

What is quantized

Everything below uses compressed-tensors, pack-quantized, and symmetric group quantization.

Group Scheme What it covers
A INT4, gs128 MoE routed experts (512 per layer, top-10). 58.0 GiB, 92% of the body
B INT6, gs64 linear_attn in and out projections, QSA q/k/v/o_proj, shared expert
C INT8, gs64 hyper-connection low-rank mixers
D INT8, gs128 lm_head, embed_tokens, PLE key_proj and value_proj, indexer index_qk_proj
E BF16 router, vision tower, PLE n-gram table, MTP module

3x24GB+MTP converts the MTP experts to INT4 in a separate directory (make_mtp_int4.py).

License

The model weights inherit the upstream Qwen Community License 1.0. See LICENSE in this directory.

The patches in every directory under vllm-patch/ are derivative works of vLLM, so they stay under the Apache License, Version 2.0, like vLLM itself. vllm-patch/LICENSE holds the full text, and vllm-patch/NOTICE names the vLLM files that each of them changes.

The Dockerfile and compose.yaml in each directory, and make_mtp_int4.py, contain no vLLM code. They are under the MIT license, in vllm-patch/LICENSE.MIT.

If you find these useful, make sure to like my repo.

Downloads last month
1,138
Safetensors
Model size
140B params
Tensor type
I32
·
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound

Quantized
(394)
this model
Quantizations
1 model

Collection including Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound