Q4_K_M · 4K context

K2 Horizon MoVA 36B-A4B MoE model with 36B total parameters and 4B active per token · estimated Q4_K_M memory about 22GB

This is a MoE model with a MoVA structure that activates approximately 4B per token out of a total of 36B. The decode computation is lightweight, but the full weights must be accessible in memory, and the official server example is the multi-GPU route.

Estimated Q4_K_M memory
about 22GB
total parameters
36B
active parameter
4B
native context
512K

Start with a model

Find hardware for this model

Also running document search?

Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.

Matching devices 40· 5 shown

Mac mini M4 32GB

Comfortable fit · Q4_K_M · Headroom 5.1GB

₩2,100,000

34.6 ~ 38.6 tok/s

Mac mini M6 32GB

Comfortable fit · Q4_K_M · Headroom 5.1GB

₩2,774,000

50.4 ~ 56.2 tok/s

MBP M3 Pro 36GB

Comfortable fit · Q4_K_M · Headroom 8.1GB

₩2,850,000

44.5 ~ 49.6 tok/s

Mac mini M4 Pro 48GB

Comfortable fit · Q4_K_M · Headroom 19.1GB

₩3,035,000

83.3 ~ 92.5 tok/s

MBA M5 32GB

Comfortable fit · Q4_K_M · Headroom 5.1GB

₩3,095,000

44.1 ~ 49.3 tok/s

Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.