Q4 · 4K context standard

Qwen3.8 27B총 552B 중 토큰당 16B가 활성화되는 MoE 모델 · Q2_K 264.5GB 저장 파일 · 작업 메모리 미확인

An MoE model with a 552B backbone and separate Engram tables. Published runs on Spark, RTX 5090, and high-memory Macs are separated by hardware and format, including experimental paths that require SSD streaming and dedicated runtimes.

Q2_K 저장 파일
264.5GB · mmap 페이징
total parameters
27B
active parameter
27B
native context
256K

Representative equipment to run this model

DGX Spark 128GB

available 116GB · 273GB/s

SGLang · NVFP4 · DFlash2

Decode

42.8 ~ 50.3 tok/s

First token

6.7 to 13.3 seconds

Running Guide

Qwen3.8 27B Local Serving Run Recipe

Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.

Speed tuning for each device

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.