Q4 · 4K context standard

Gemma 4 12B: 11.95B density model · Q4 execution memory approximately 8GB

Compare memory requirements of approximately 8GB for Gemma 4 12B Q4 execution, device-specific prefill, token generation speed, quantization, and 262,144 token context length.

Model specifications

Q4 Memory Required
about 8GB
total parameters
11.95B
active parameter
11.95B
native context
262,144 tokens

A single 11.95B parameter dense model supports 256K (262,144) native contexts. High-speed standalone operation is possible in 16GB~24GB GPU and Mac environments.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

Gemma 4 12B: LM Studio and vLLM on 16GB hardware

Verify prefill and token generation speed of this model