Q4 · 4K context standard

Qwen3.8 27B2.51675648B 밀도형 모델 · Q4_K_M 예상 메모리 약 3GB

A dense open-weight model with 2.51675648B parameters and a native 131,072-token context. Displayed memory and generation speeds are estimates based on quantization and hardware specifications, not measured results.

Q4_K_M 예상 메모리
About 3GB
total parameters
27B
active parameter
27B
native context
256K

Representative equipment to run this model

RTX 5060 Ti 16GB

available 15GB · 448GB/s

llama.cpp CUDA · Q4_K_M

Decode

5.6 ~ 5.7 tok/s

First token

6.9 to 12.2 seconds

2× RTX 5090 64GB (PCIe)

available 61GB · 3584GB/s

llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces

Decode

87.6 ~ 135.3 tok/s

First token

1.4 to 3.1 seconds

MBA M4 16GB

available 13GB · 120GB/s

MLX 4-bit

Decode

0.0 tok/s

First token

0.0 seconds (OOM)

Running Guide

Qwen3.8 27B Local Serving Run Recipe

Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.

Speed tuning for each device

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.