Q4 · 4K context standard
Qwen3.8 27BTotal 124.848460496B 중 토큰당 5.1B가 활성화되는 MoE 모델 · MXFP4 예상 메모리 약 71GB
An image/video-input MoE model with about 124.85B total and 5.1B active parameters. Native context is 128K. Only official DGX Spark INT4/MXFP4 measurements anchor the speed experience; rates are not transferred to other hardware.
- MXFP4 예상 메모리
- About 71GB
- total parameters
- 27B
- active parameter
- 27B
- native context
- 256K
Representative equipment to run this model
Decode
42.8 ~ 50.3 tok/s
First token
6.7 to 13.3 seconds
Running Guide
Qwen3.8 27B Local Serving Run Recipe
Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.