35B An Jeong-kwon · Officially released product · Updated 2026-09-06
48GB GPU memory, main body cost separate
NVIDIA A40 48GB local LLM speed, memory, electricity rate
696 GB/s 384-bit 48GB ECC VRAM. This is a high-capacity GPU that can load 70B class 4-bit or 27B~35B high-precision models on a single card in used workstation and server environments.
Core specifications
- available memory
- 46GB
- Memory bandwidth
- 696GB/s
- local LLM sustained load
- about 400W
- Korean price range (KRW)
- 680–780 × KRW 10,000
a person who fits well
- Utilizes 46GB of available memory
- CUDA-based local inference
- Qwen3.8-Flash-Next level high-speed generation
Check before purchasing
Price and power shown are GPU-centric comparisons. In a real system, a host PC, power supply, cooling and slot configuration are added.
4K context · Q4 model-specific expected performance
| Model | required memory | Decode | First token | Prefill |
|---|---|---|---|---|
| Gemma 4 12B llama.cpp CUDA · Q4_K_M | 8GB | 71 ~ 78.9 tok/s | 3.0 ~ 3.8 s | 1,091 ~ 1,388 tok/s |
| Qwen3.8-Flash-Next llama.cpp CUDA · Q4_K_M | 77.01GB | 9.1 ~ 10.3 tok/s | 3.8 ~ 12.9 s | 319 ~ 1,107 tok/s (NVMe cold~warm) |
Single-user expected range and may vary depending on runtime, cooling, and quantization files.
Frequently Asked Questions
Which local LLM can be run on an NVIDIA A40 48GB?
Gemma 4 12B, Qwen3.8-Flash-Next, and others can be compared with 4K input and device-specific recommended formats. The required memory varies depending on the model and context length.
How much power does a local LLM on an NVIDIA A40 48GB consume?
Expected to be around 400W under sustained load, with an actual range of 330–500W.
Are the displayed token speeds ground truth?
This is an estimated range calibrated to the calculator against official hardware specifications and verified benchmark ranges. Runtime, cooling, quantization may vary depending on file and context length.