CUDA large capacity · Official product specifications · Updated 2026-09-06
This is a device that puts a large model into one node. This is a different choice compared to the speed of a gaming GPU.
NVIDIA DGX Spark 128GB Local LLM Selection Criteria
This is ideal when you need large integrated memory and the NVIDIA development environment and prioritize model loading over single GPU maximum speed.
Core specifications
- available memory
- 116GB
- Memory bandwidth
- 273GB/s
- local LLM sustained load
- about 190W
- Korean price range (KRW)
- 710–870 × KRW 10,000
a person who fits well
- Loading large models
- CUDA server development
- Review of 2 Spark expansions
Check before purchasing
You shouldn't just look at the number 128GB and judge it to be faster than the RTX 5090; you should look at the single request speed and memory capacity separately.
What tasks is it suitable for?
| task | determination | reason |
|---|---|---|
| Large LLMs | comfortable | Its strength is the 128GB memory capacity of a single node. |
| serving development | fit | Great for using NVIDIA runtime and server tools. |
| Create image | conditional | It's possible, but you'll need to compare it to GeForce speeds for the same budget. |
Configuration Selection Criteria
The first bottleneck encountered
Although large capacity loads are strong, the single request rate of smaller models may be lower than that of the high-performance GeForce due to memory bandwidth.
Recommended purchase setup
DGX Spark 128GB single node · NVMe free space · Wired network
Criteria for spending more money
There is a reason to pay for large models and CUDA when you need them at the same time. If you only want creation speed within 32GB, RTX 5090 is right for you.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B llama.cpp · Q4_K_M | 8GB | 27.8 ~ 30.9 tok/s | 3.8 ~ 5.1 s | 802 ~ 1,084 tok/s |
| Qwen3.8-Flash-Next llama.cpp · Q4_K_M | 107.13GB | 16.7 ~ 23.4 tok/s | 3.1 ~ 4.2 s | 987 ~ 1,335 tok/s |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
Which local LLM can be run on DGX Spark 128GB?
Gemma 4 12B, Qwen3.8-Flash-Next, and others can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.
How much local LLM computing power does DGX Spark 128GB have?
Expected to be around 190W under sustained load, with an actual range of 160–233W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.