Q4 · 4K context standard

Qwen3.8 27B: 27B density model · Q4 execution memory approximately 18GB

Compare memory required for Qwen3.8 27B Q4 execution, approximately 18GB, and device-specific prefill, token generation speed, quantization, and 262,144 token context length.

Model specifications

Q4 Memory Required
about 18GB
total parameters
27B
active parameter
27B
native context
262,144 tokens

27B dense open weight model with built-in multi-stage multi-token prediction (MTP) learning head. Supports native 262,144 (256K) token contexts.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

Qwen3.8-27B setup: oMLX on Mac and llama.cpp on RTX

Verify prefill and token generation speed of this model