TERNARY_PQ2 · 4K 문맥 기준

Bonsai 2 27B 27.36B dense model · estimated TERNARY_PQ2 memory about 9GB

A ternary-weight model derived from Qwen3.8-27B. The PQ2_0 language weights are 7.21 GB and require a Hadamard-aware runtime rather than stock llama.cpp. The speed experience prioritizes PrismML's published PQ2_0 measurements.

Estimated TERNARY_PQ2 memory
About 9GB
total parameters
27.36B
active parameter
27.36B
native context
256K

Start with a model

Find hardware for this model

Also running document search?

Reserve this for embedding and reranking models sharing the same GPU or unified memory. Leave it at none if they run separately on the CPU.

Matching devices 35· 5 shown

RTX 5060 Ti 16GB

Comfortable fit · TERNARY_PQ2 · Headroom 6.2GB

₩2,100,000

46.6 ~ 51.6 tok/s

Mac mini M4 Pro 24GB

Comfortable fit · TERNARY_PQ2 · Headroom 11.2GB

₩2,650,000

27.3 ~ 30.3 tok/s

Mac mini M4 Pro 48GB

Comfortable fit · TERNARY_PQ2 · Headroom 32.2GB

₩3,035,000

27.3 ~ 30.3 tok/s

RTX 3090 24GB

Comfortable fit · TERNARY_PQ2 · Headroom 13.7GB

₩3,125,000

93.5 ~ 103.9 tok/s

MBP M4 Pro 48GB

Comfortable fit · TERNARY_PQ2 · Headroom 32.2GB

₩3,450,000

27.3 ~ 30.3 tok/s

Price is the range midpoint, with a basic host added for GPU-only prices. Speed is an estimate for the same model and 4K input.

Running Guide

Bonsai 2 27B Local Serving Run Recipe

Selecting a device changes the quantization, safety context, GPU offload, Flash Attention, KV cache, and batch start values.

Speed tuning for each device

If you want to choose equipment again

You can first narrow down the model configurations based on budget, noise level, operating system and whether it's acceptable used.