Q4_K_M · 4K context
Holo4 35B-A3B MoE model with 35B total parameters and 3B active per token · estimated Q4_K_M memory about 22GB
A computer-enabled MoE model that activates approximately 3B per token out of a total of 35B. The activation parameter describes the amount of computation, and local loading requires full weights and vision/KV cache headroom.
- Estimated Q4_K_M memory
- about 22GB
- total parameters
- 35B
- active parameter
- 3B
- native context
- 256K
Start with a model
Find hardware for this model
Also running document search?
Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.
Matching devices 40· 5 shown
Comfortable fit · Q4_K_M · Headroom 5.7GB
₩2,100,000
46.1 ~ 51.5 tok/s
Comfortable fit · Q4_K_M · Headroom 5.7GB
₩2,774,000
67.2 ~ 74.9 tok/s
Comfortable fit · Q4_K_M · Headroom 8.7GB
₩2,850,000
59.3 ~ 66.1 tok/s
Comfortable fit · Q4_K_M · Headroom 19.7GB
₩3,035,000
111.1 ~ 123.4 tok/s
Comfortable fit · Q4_K_M · Headroom 5.7GB
₩3,095,000
58.8 ~ 65.7 tok/s
Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.