Q4 · 4K context standard
Qwen3.6 35B-A3B (MoE): MoE model with 3B weights activated per token out of total 35B · Q4 execution memory approximately 22GB
Qwen3.6 35B-A3B (MoE) Q4 execution requires approximately 22GB of memory and compares device-specific prefill, token generation speed, quantization, and 262,144 token context length.
Model specifications
- Q4 Memory Required
- about 22GB
- total parameters
- 35B
- active parameter
- 3B
- native context
- 262,144 tokens
This is an MoE model in which approximately 3B of the total 35B parameters are activated per token. Memory loading requires 35B capacity, and decoding bandwidth is consumed based on 3B active parameters.
Main equipment
- RTX 5060 Ti 16GB — available 15GB, 448GB/s Feel the speed · Price, power, and sales composition
- 2× RTX 5090 64GB (PCIe) — available 61GB, 3584GB/s Feel the speed · Price, power, and sales composition
- MBA M4 24GB — available 20GB, 120GB/s · Feel the speed · Price, power, and sales composition
Speed tuning for each device
When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.