Q4_K_M · 4K context
K2 Horizon MoVA 36B-A4B MoE model with 36B total parameters and 4B active per token · estimated Q4_K_M memory about 22GB
This is a MoE model with a MoVA structure that activates approximately 4B per token out of a total of 36B. The decode computation is lightweight, but the full weights must be accessible in memory, and the official server example is the multi-GPU route.
- Estimated Q4_K_M memory
- about 22GB
- total parameters
- 36B
- active parameter
- 4B
- native context
- 512K
Start with a model
Find hardware for this model
Also running document search?
Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.
Matching devices 40· 5 shown
Comfortable fit · Q4_K_M · Headroom 5.1GB
₩2,100,000
34.6 ~ 38.6 tok/s
Comfortable fit · Q4_K_M · Headroom 5.1GB
₩2,774,000
50.4 ~ 56.2 tok/s
Comfortable fit · Q4_K_M · Headroom 8.1GB
₩2,850,000
44.5 ~ 49.6 tok/s
Comfortable fit · Q4_K_M · Headroom 19.1GB
₩3,035,000
83.3 ~ 92.5 tok/s
Comfortable fit · Q4_K_M · Headroom 5.1GB
₩3,095,000
44.1 ~ 49.3 tok/s
Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.