Q4 · 4K context standard

Qwen3.8-Flash-Next: 125B total, 6B per token activated in MoE model · Q4 execution memory approximately 108GB

Compare prefill, token generation speed, quantization, and 262,144 token context length required for Qwen3.8-Flash-Next Q4 execution, which needs approximately 108GB of memory.

Model specifications

Q4 Memory Required
about 108GB
total parameters
125B
active parameter
6B
native context
262,144 tokens

Approximately 180B resident weight MoE model consisting of 125B core LMs (6B active) with 51B n-gram/PLE embeddings and 4B MTPs built into the architecture. The built-in 51B PLE embedding table is RAM offloadable and is independent of the Prompt Lookup speculative decoding that looks for repeated prompt phrases.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

Qwen3.8-Flash-Next: a single DGX Spark SGLang recipe

Verify prefill and token generation speed of this model