large model · Officially released product · Updated 2026-09-06

96GB GPU memory, body cost separate

NVIDIA RTX PRO 6000 Blackwell 96GB local LLM speed/memory/electricity rate

1,792 GB/s 96GB ECC VRAM workstation GPU. Fully load 70B to 120B 4-bit models on a single card, all within the available 93GB VRAM budget.

Core specifications

available memory
93GB
Memory bandwidth
1792GB/s
local LLM sustained load
about 650W
Korean price range (KRW)
2,450–2,650 × KRW 10,000

a person who fits well

  • Utilizes 93GB of available memory
  • CUDA-based local inference
  • GLM-5.3-Flash (MoE) level high-speed generation

Check before purchasing

Price and power shown are GPU-centric comparisons. In a real system, a host PC, power supply, cooling and slot configuration are added.

4K context · Q4 model-specific expected performance

Modelrequired memoryDecodeFirst tokenPrefill
Gemma 4 12B
llama.cpp CUDA · Q4_K_M
8GB198 ~ 218.3 tok/s1.1 ~ 1.4 s3,025 ~ 3,698 tok/s
GLM-5.3-Flash (MoE)
llama.cpp CUDA · Q4_K_M
190.21GB6.6 ~ 6.6 tok/s4.2 ~ 5.1 s823 ~ 1,006 tok/s

Single-user expected range and may vary depending on runtime, cooling, and quantization files.

Frequently Asked Questions

Which local LLM can be run on RTX PRO 6000 96GB?

Gemma 4 12B, GLM-5.3-Flash (MoE), etc. can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.

How much power does a local LLM on an RTX PRO 6000 96GB consume?

Expected to consume approximately 650W under sustained load, with an actual range estimated between 520–780W.

Are the displayed token speeds ground truth?

This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.