Q4 · 4K context standard
Gemma 4 12B: 11.95B density model · Q4 execution memory approximately 8GB
Compare memory requirements of approximately 8GB for Gemma 4 12B Q4 execution, device-specific prefill, token generation speed, quantization, and 262,144 token context length.
Model specifications
- Q4 Memory Required
- about 8GB
- total parameters
- 11.95B
- active parameter
- 11.95B
- native context
- 262,144 tokens
A single 11.95B parameter dense model supports 256K (262,144) native contexts. High-speed standalone operation is possible in 16GB~24GB GPU and Mac environments.
Main equipment
- RTX 5060 Ti 16GB — available 15GB, 448GB/s Feel the speed · Price, power, and sales composition
- 2× RTX 5090 64GB (PCIe) — available 61GB, 3584GB/s Feel the speed · Price, power, and sales composition
- MBA M4 16GB — available 13GB, 120GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.