Q4 · 4K context standard
GLM-5.3-Flash (MoE): MoE model with 18B activated per token out of a total of 320B · Q4 execution memory approximately 191GB
Compares prefill, token generation speed, quantization, and 1,048,576 token context length required for GLM-5.3-Flash (MoE) Q4 execution, needing approximately 191GB of memory per device.
Model specifications
- Q4 Memory Required
- about 191GB
- total parameters
- 320B
- active parameter
- 18B
- native context
- 1,048,576 tokens
Hybrid sparse+linear attention MoE model with a total of 320B parameters (18B active). Supports the official 1,048,576 (1M) token context (text_config.max_position_embeddings) and loads in large memory environments of 256GB or more.
Main equipment
- 2× RTX 5090 64GB (PCIe) — available 61GB, 3584GB/s Feel the speed · Price, power, and sales composition
- Studio M5 Ultra 256GB — available 236GB, 1200GB/s Feel the speed · Price, power, and sales composition
- Studio M3 Ultra 512GB — available 480GB, 819GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.