multi equipment · Official product specifications · Updated 2026-09-06
Configuration to run a model across two devices · Total memory 256GB
2× NVIDIA DGX Spark cluster 256GB local LLM speed·memory·electricity cost
Two nodes connected by a ConnectX-7 200 Gbps (25 GB/s) RoCE link (256 GB total distributed memory, 546 GB/s combined local bandwidth). When optimizing TP=2, even a single request can be faster, and the measured configuration for each model is applied first.
Core specifications
- available memory
- 232GB
- Memory bandwidth
- 546GB/s
- local LLM sustained load
- about 380W
- Korean price range (KRW)
- 1,420–1,740 × KRW 10,000
a person who fits well
- Utilize 232GB of available memory
- Two-Device Distributed Inference
- GLM-5.3-Flash (MoE) model comparison
Check before purchasing
Even if you connect two devices, the memories are not automatically merged into one. Distributed runtime support and interconnect bottlenecks must be checked together.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B 4-bit · 2-node inference | 8GB | 27.8 ~ 44.5 tok/s | 2.3 ~ 4.4 s | 930 ~ 1,779 tok/s |
| GLM-5.3-Flash (MoE) 4-bit · 2-node inference | 190.21GB | 18.5 ~ 29.6 tok/s | 5.5 ~ 10.5 s | 393 ~ 752 tok/s |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
What LLM can be run on 2× DGX Spark 256GB?
Gemma 4 12B, GLM-5.3-Flash (MoE), etc. can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.
How much local LLM computing power does 2× DGX Spark with 256GB provide?
Expected to consume approximately 380W under sustained load, with an actual range estimated between 320–466W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.