Q4 · 4K context standard

GLM-5.3-Flash (MoE): MoE model with 18B activated per token out of a total of 320B · Q4 execution memory approximately 191GB

Compares prefill, token generation speed, quantization, and 1,048,576 token context length required for GLM-5.3-Flash (MoE) Q4 execution, needing approximately 191GB of memory per device.

Model specifications

Q4 Memory Required
about 191GB
total parameters
320B
active parameter
18B
native context
1,048,576 tokens

Hybrid sparse+linear attention MoE model with a total of 320B parameters (18B active). Supports the official 1,048,576 (1M) token context (text_config.max_position_embeddings) and loads in large memory environments of 256GB or more.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

GLM-5.3-Flash: serving a server-scale 320B MoE

Verify prefill and token generation speed of this model