Q4 · 4K context standard

DeepSeek-V4-Flash-0731 (MoE): total 304B with 13B activated per token MoE model · Q4 execution memory approximately 181GB

DeepSeek-V4-Flash-0731 (MoE) Q4 execution requires approximately 181GB of memory, and compares device-specific prefill, token generation speed, quantization, and 1,048,576 token context length.

Model specifications

Q4 Memory Required
about 181GB
total parameters
304B
active parameter
13B
native context
1,048,576 tokens

Large MoE model with a total of 304B parameters (13B active) supporting 1,048,576 (1M) token contexts (max_position_embeddings). The official example uses 4×GB300 nodes, and the community compressed version must be checked separately for actual size and runtime compatibility for each artifact.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

DeepSeek-V4-Flash: distributed serving for a 304B model

Verify prefill and token generation speed of this model