Q4 · 4K context standard

Qwen3.6 35B-A3B (MoE): MoE model with 3B weights activated per token out of total 35B · Q4 execution memory approximately 22GB

Qwen3.6 35B-A3B (MoE) Q4 execution requires approximately 22GB of memory and compares device-specific prefill, token generation speed, quantization, and 262,144 token context length.

Model specifications

Q4 Memory Required
about 22GB
total parameters
35B
active parameter
3B
native context
262,144 tokens

This is an MoE model in which approximately 3B of the total 35B parameters are activated per token. Memory loading requires 35B capacity, and decoding bandwidth is consumed based on 3B active parameters.

Main equipment

Speed tuning for each device

When selecting equipment, you can view quantization, safe context, GPU offloading, Flash Attention, KV cache, and batch start values together.

Qwen3.6 35B-A3B: vLLM setup and the qwen3 parser

Verify prefill and token generation speed of this model