multi equipment · Officially released product · Updated 2026-09-06

Configuration to run a model across two devices · Total memory 64GB

2× NVIDIA GeForce RTX 5090 PCIe 64GB local LLM speed·memory·electricity cost

2× 32GB GDDR7 VRAM (64GB combined, 3,584 GB/s combined local bandwidth). The RTX 5090 does not support NVLink and communicates only through the PCIe Gen5 interface, resulting in a decrease in parallel efficiency due to an all-reduce communication bottleneck.

Core specifications

available memory
61GB
Memory bandwidth
3584GB/s
local LLM sustained load
about 1100W
Korean price range (KRW)
1,240–1,560 × KRW 10,000

a person who fits well

  • Utilize 61GB of available memory
  • Two-Device Distributed Inference
  • GLM-5.3-Flash (MoE) model comparison

Check before purchasing

Even if you connect two devices, the memories are not automatically merged into one. Distributed runtime support and interconnect bottlenecks must be checked together.

4K context · Q4 model-specific expected performance

Modelrequired memoryDecodeFirst tokenPrefill
Gemma 4 12B
llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces
8GB198 ~ 305.6 tok/s0.89 ~ 1.4 s2,893 ~ 4,622 tok/s
GLM-5.3-Flash (MoE)
llama.cpp CUDA · Q4_K_M · GPU split into 2 pieces
190.21GB5.1 ~ 5.2 tok/s4.2 ~ 6.6 s641 ~ 1,025 tok/s

Single-user expected range and may vary depending on runtime, cooling, and quantization files.

Frequently Asked Questions

What local LLM can be run on 2× RTX 5090 64GB (PCIe)?

Gemma 4 12B, GLM-5.3-Flash (MoE), etc. can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.

What is the power consumption of a local LLM using 2× RTX 5090 64GB (PCIe)?

Expected to consume approximately 1100W under sustained load, with an actual range estimated between 900–1350W.

Are the displayed token speeds ground truth?

This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.