Q4_K_M · 4K context

K2 Horizon 32B 32B dense model · estimated Q4_K_M memory about 21GB

This is a 32B dense model of K2 Horizon. On 24GB devices, check Q4 and short context first, and if you want to use long inputs and simultaneous requests, you should also check the available memory of 32GB or more.

Estimated Q4_K_M memory
About 21GB
total parameters
32B
active parameter
32B
native context
512K

Start with a model

Find hardware for this model

Also running document search?

Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.

Matching devices 40· 5 shown

2× RTX 3090 48GB (NVLink)

Comfortable fit · Q4_K_M · Headroom 24.3GB

₩5,900,000

48.5 ~ 67.4 tok/s

Studio M4 Max 64GB

Comfortable fit · Q4_K_M · Headroom 35.3GB

₩6,050,000

21.4 ~ 23.7 tok/s

Studio M5 Max 64GB

Comfortable fit · Q4_K_M · Headroom 35.3GB

₩6,925,000

24 ~ 26.6 tok/s

MBP M5 Max 64GB

Comfortable fit · Q4_K_M · Headroom 35.3GB

₩7,300,000

24 ~ 26.6 tok/s

NVIDIA A40 48GB

Comfortable fit · Q4_K_M · Headroom 25.3GB

₩7,500,000

26.5 ~ 29.5 tok/s

Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.