Q4 · 4K context standard
Qwen3.8 27B총 552B 중 토큰당 16B가 활성화되는 MoE 모델 · Q2_K 264.5GB 저장 파일 · 작업 메모리 미확인
An MoE model with a 552B backbone and separate Engram tables. Published runs on Spark, RTX 5090, and high-memory Macs are separated by hardware and format, including experimental paths that require SSD streaming and dedicated runtimes.
- Q2_K 저장 파일
- 264.5GB · mmap 페이징
- total parameters
- 27B
- active parameter
- 27B
- native context
- 256K
Representative equipment to run this model
Decode
42.8 ~ 50.3 tok/s
First token
6.7 to 13.3 seconds
Running Guide
Qwen3.8 27B Local Serving Run Recipe
Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.