speed priority · Officially released product · Updated 2026-09-06

Local creation operations that fit 32GB are the fastest.

RTX 5090 32GB local LLM/image/video generation performance

This makes sense when you prioritize both single GPU generation speed and 32GB VRAM and are willing to accept high power and deployment costs.

Core specifications

available memory
30.5GB
Memory bandwidth
1792GB/s
local LLM sustained load
about 620W
Korean price range (KRW)
620–780 × KRW 10,000

a person who fits well

  • Fast 27B to 35B inference
  • Image / video
  • Modern CUDA workflow

Check before purchasing

You shouldn't just look at the price of the graphics card alone; you must also include high-capacity power, cooling, and body costs.

What tasks is it suitable for?

taskdeterminationreason
27B·35B conversationcomfortableYou can expect high token generation rates on a single GPU.
Image/videocomfortable32GB VRAM and the latest CUDA support are key.
24 hour serverconditionalThe trade-off for high performance is high power and cooling costs.

Configuration Selection Criteria

The first bottleneck encountered

Larger models exceeding 32GB will not fully utilize peak computing performance and require offloading.

Recommended purchase setup

RTX 5090 32GB · System RAM 96GB or more · 1200W power · Large case

Criteria for spending more money

Images/videos and 35B speed are suitable for reducing profits or working time. If you only keep the conversation light, you are overinvesting.

4K context · Q4 model-specific expected performance

Modelrequired memoryDecodeFirst tokenPrefill
Gemma 4 12B
llama.cpp CUDA · Q4_K_M
8GB198 ~ 218.3 tok/s1.2 ~ 1.4 s3,014 ~ 3,399 tok/s
Qwen3.6 35B-A3B (MoE)
llama.cpp CUDA · Q4_K_M
21.28GB111.5 ~ 280 tok/s0.55 ~ 0.62 s6,712 ~ 7,569 tok/s
Qwen3.8-Flash-Next
llama.cpp CUDA · Q4_K_M
77.01GB12.7 ~ 14 tok/s1.9 ~ 5.8 s710 ~ 2,183 tok/s (NVMe cold~warm)

Single-user expected range and may vary depending on runtime, cooling, and quantization files.

Frequently Asked Questions

What LLM can be run on an RTX 5090 with 32GB?

Gemma 4 12B, Qwen3.6 35B-A3B (MoE), Qwen3.8-Flash-Next, and others can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.

What is the power consumption of a local LLM on an RTX 5090 with 32GB?

Expected to be around 620W under sustained load, with an actual range of 480–760W.

Are the displayed token speeds ground truth?

This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.