multi equipment · Officially released product · Updated 2026-09-06
Configuration to run a model across two devices · Total memory 48GB
2× NVIDIA GeForce RTX 3090 NVLink 48GB local LLM speed·memory·electricity cost
Dual GPU configuration with 2× 24GB VRAM (48GB combined, 1,872 GB/s combined local bandwidth) and 3rd generation NVLink (112.5 GB/s bi-directional). When supporting tensor/pipeline parallelism, the 70B class model is fully loaded.
Core specifications
- available memory
- 45GB
- Memory bandwidth
- 1872GB/s
- local LLM sustained load
- about 780W
- Korean price range (KRW)
- 330–440 × KRW 10,000
a person who fits well
- Utilize 45GB of available memory
- Two-Device Distributed Inference
- Comparison of Qwen3.8-Flash-Next-grade models
Check before purchasing
Even if you connect two devices, the memories are not automatically merged into one. Distributed runtime support and interconnect bottlenecks must be checked together.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces | 8GB | 129.8 ~ 180.3 tok/s | 1.6 ~ 2.7 s | 1,510 ~ 2,562 tok/s |
| Qwen3.8-Flash-Next llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces | 77.01GB | 9.5 ~ 10.7 tok/s | 2.1 ~ 9.6 s | 433 ~ 2,003 tok/s (NVMe cold~warm) |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
What local LLM can be run on 2× RTX 3090 48GB (NVLink)?
Gemma 4 12B, Qwen3.8-Flash-Next, and others can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.
What is the power consumption of a local LLM using 2× RTX 3090 48GB (NVLink)?
Expected to consume approximately 780W under sustained load, with an actual range estimated between 650–920W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.