TERNARY_PQ2 · 4K 문맥 기준
Bonsai 2 27B 27.36B dense model · estimated TERNARY_PQ2 memory about 9GB
A ternary-weight model derived from Qwen3.8-27B. The PQ2_0 language weights are 7.21 GB and require a Hadamard-aware runtime rather than stock llama.cpp. The speed experience prioritizes PrismML's published PQ2_0 measurements.
- Estimated TERNARY_PQ2 memory
- About 9GB
- total parameters
- 27.36B
- active parameter
- 27.36B
- native context
- 256K
Start with a model
Find hardware for this model
Also running document search?
Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.
Matching devices 35· 5 shown
Comfortable fit · TERNARY_PQ2 · Headroom 6.2GB
₩2,100,000
46.6 ~ 51.6 tok/s
Comfortable fit · TERNARY_PQ2 · Headroom 11.2GB
₩2,650,000
27.3 ~ 30.3 tok/s
Comfortable fit · TERNARY_PQ2 · Headroom 32.2GB
₩3,035,000
27.3 ~ 30.3 tok/s
Comfortable fit · TERNARY_PQ2 · Headroom 13.7GB
₩3,125,000
93.5 ~ 103.9 tok/s
Comfortable fit · TERNARY_PQ2 · Headroom 32.2GB
₩3,450,000
27.3 ~ 30.3 tok/s
Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.
Bonsai 2 27B Local Serving Run Recipe
Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.
If you want to choose equipment again
You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.