Q4 · 4K context standard
Gemma 4 26B-A4B (MoE): MoE model with 3.8B activated per token out of total 25.2B · Q4 execution memory approximately 16GB
Compare memory required for prefill, token generation speed, quantization, and 262,144 token context length for Gemma 4 26B-A4B (MoE) Q4 execution across different hardware.
Model specifications
- Q4 Memory Required
- about 16GB
- total parameters
- 25.2B
- active parameter
- 3.8B
- native context
- 262,144 tokens
This is an MoE model in which 3.8B of the total 25.2B parameters are activated per token. Supports 256K (262,144) native contexts and provides fast decoding per load.
Main equipment
- RTX 5060 Ti 16GB — available 15GB, 448GB/s Feel the speed · Price, power, and sales composition
- 2× RTX 5090 64GB (PCIe) — available 61GB, 3584GB/s Feel the speed · Price, power, and sales composition
- MBA M4 16GB — available 13GB, 120GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.