Q4 · 4K context standard
DeepSeek-V4-Flash-0731 (MoE): total 304B with 13B activated per token MoE model · Q4 execution memory approximately 181GB
DeepSeek-V4-Flash-0731 (MoE) Q4 execution requires approximately 181GB of memory, and compares device-specific prefill, token generation speed, quantization, and 1,048,576 token context length.
Model specifications
- Q4 Memory Required
- about 181GB
- total parameters
- 304B
- active parameter
- 13B
- native context
- 1,048,576 tokens
Large MoE model with a total of 304B parameters (13B active) supporting 1,048,576 (1M) token contexts (max_position_embeddings). The official example uses 4×GB300 nodes, and the community compressed version must be checked separately for actual size and runtime compatibility for each artifact.
Main equipment
- 2× RTX 5090 64GB (PCIe) — available 61GB, 3584GB/s Feel the speed · Price, power, and sales composition
- Studio M5 Ultra 256GB — available 236GB, 1200GB/s Feel the speed · Price, power, and sales composition
- Studio M3 Ultra 512GB — available 480GB, 819GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.