Q4 · 4K context standard
Qwen3.8-Flash-Next: 125B total, 6B per token activated in MoE model · Q4 execution memory approximately 108GB
Compare prefill, token generation speed, quantization, and 262,144 token context length required for Qwen3.8-Flash-Next Q4 execution, which needs approximately 108GB of memory.
Model specifications
- Q4 Memory Required
- about 108GB
- total parameters
- 125B
- active parameter
- 6B
- native context
- 262,144 tokens
Approximately 180B resident weight MoE model consisting of 125B core LMs (6B active) with 51B n-gram/PLE embeddings and 4B MTPs built into the architecture. The built-in 51B PLE embedding table is RAM offloadable and is independent of the Prompt Lookup speculative decoding that looks for repeated prompt phrases.
Main equipment
- RTX 3090 24GB — available 22.5GB, 936GB/s Feel the speed · Price, power, and sales composition
- RTX PRO 6000 96GB — available 93GB, 1792GB/s Feel the speed · Price, power, and sales composition
- MBP M3 Max 128GB — available 114GB, 400GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.