Q4_K_M · 4K context

Holo4 35B-A3B MoE model with 35B total parameters and 3B active per token · estimated Q4_K_M memory about 22GB

A computer-enabled MoE model that activates approximately 3B per token out of a total of 35B. The activation parameter describes the amount of computation, and local loading requires full weights and vision/KV cache headroom.

Estimated Q4_K_M memory
about 22GB
total parameters
35B
active parameter
3B
native context
256K

Start with a model

Find hardware for this model

Also running document search?

Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.

Matching devices 40· 5 shown

Mac mini M4 32GB

Comfortable fit · Q4_K_M · Headroom 5.7GB

₩2,100,000

46.1 ~ 51.5 tok/s

Mac mini M6 32GB

Comfortable fit · Q4_K_M · Headroom 5.7GB

₩2,774,000

67.2 ~ 74.9 tok/s

MBP M3 Pro 36GB

Comfortable fit · Q4_K_M · Headroom 8.7GB

₩2,850,000

59.3 ~ 66.1 tok/s

Mac mini M4 Pro 48GB

Comfortable fit · Q4_K_M · Headroom 19.7GB

₩3,035,000

111.1 ~ 123.4 tok/s

MBA M5 32GB

Comfortable fit · Q4_K_M · Headroom 5.7GB

₩3,095,000

58.8 ~ 65.7 tok/s

Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.