Q4 · 4K context standard

Gemma 4 26B-A4B (MoE): MoE model with 3.8B activated per token out of total 25.2B · Q4 execution memory approximately 16GB

Compare memory required for prefill, token generation speed, quantization, and 262,144 token context length for Gemma 4 26B-A4B (MoE) Q4 execution across different hardware.

Model specifications

Q4 Memory Required
about 16GB
total parameters
25.2B
active parameter
3.8B
native context
262,144 tokens

This is an MoE model in which 3.8B of the total 25.2B parameters are activated per token. Supports 256K (262,144) native contexts and provides fast decoding per load.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

Gemma 4 26B-A4B: a serving recipe for 24–32GB hardware

Verify prefill and token generation speed of this model