Qwopus3.5-4B-Coder-oQ4e-mtp

This model was quantized using oQ (oMLX v0.5.1) mixed-precision quantization.

Quantization details

  • Model type: qwen3_5
  • Bits: 4
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.1
  • Settings:
    • Thinking: Disabled
    • Chat template parameter: enable_thinking=false (forced)
    • TurboQuant KV Cache: Disabled (if enable: reduces speed)
    • Native MTP: Enabled (key speed improvement)

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 868.8 15.05 1178.7 tok/s 67.0 tok/s 2.795 412.1 tok/s 3.51 GB
pp4096/tg128 3546.2 14.97 1155.1 tok/s 67.3 tok/s 5.470 772.2 tok/s 4.12 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 67.0 tok/s 1.00x 1178.7 tok/s 1178.7 tok/s 868.8 2.795
2x 82.9 tok/s 1.24x 1058.3 tok/s 529.1 tok/s 1935.2 5.025
4x 134.4 tok/s 2.01x 1053.2 tok/s 263.3 tok/s 3758.2 7.698

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 63.3% 19 30 26.0 No
TRUTHFULQA 63.3% 19 30 10.3 No
GSM8K 96.7% 29 30 121.4 No
MATHQA 53.3% 16 30 15.5 No
HUMANEVAL 76.7% 23 30 123.7 No

Summary

Benchmark Accuracy
MMLU 63.3%
TRUTHFULQA 63.3%
GSM8K 96.7%
MATHQA 53.3%
HUMANEVAL 76.7%
Average 70.7%
Downloads last month
121
Safetensors
Model size
1B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Qwopus3.5-4B-Coder-oQ4e-mtp