Executable programs and extensions

When PyTorch seems to use the CPU despite an NVIDIA GPU

If the GPU appears but PyTorch runs on the CPU, check one step at a time in the same environment.

If the NVIDIA GPU1 appears but answers are slow, first run nvidia-smi in the same terminal or notebook where you launched the model to check whether the driver sees the GPU. Then run the read-only Python diagnostic below in that same environment to see whether PyTorch2 can use it. The diagnostic does not run the model or change files. Even if CUDA3 is available, that does not prove the model is computing on the GPU, so check the model settings and app logs last. nvidia-smi, torch.version.cuda and nvcc --version each report different information.

Requirements and key details
  • nvidia-smi reports driver and GPU visibility; nvcc --version reports the installed development Toolkit compiler version.
  • torch.version.cuda identifies the CUDA version of the loaded PyTorch build; torch.cuda.is_available() checks availability in the current process.
  • CUDA availability does not guarantee that the whole model runs on the GPU. Check the model device, app logs and VRAM headroom.

PyTorch may use the CPU even when the GPU is visible

Suppose a document-summary model responds slowly in an IDE notebook while CPU4 use is high. In the same notebook, nvidia-smi lists the NVIDIA GPU. Do not diagnose from the system monitor alone: check in order whether the driver sees the GPU, which PyTorch the notebook loaded, and which device the model uses for computation.

The terminal and IDE may use different Python installations. For example, the terminal can see the GPU while the notebook loads a CPU-only PyTorch build. Having nvcc installed does not show that the IDE uses the same environment. Run the checks below in the notebook that launched the model.

First check whether the driver can see the GPU

Run nvidia-smi from a terminal. It shows whether the NVIDIA driver detects the GPU, the driver version and currently visible GPU and memory information. If the command is missing or cannot find a device, check the operating system and driver before changing PyTorch packages. Under WSL2 on Windows, the Windows host provides the CUDA GPU driver. NVIDIA instructs users not to install a Linux NVIDIA driver inside WSL.

Do not read the CUDA Version in the output as the version of the installed Toolkit. It indicates the CUDA level supported by the driver. Check the development Toolkit compiler with nvcc --version. The absence of nvcc does not mean GPU execution is impossible. A prebuilt PyTorch package may include CUDA runtime5 components, but a compatible driver is still required.

A GPU workstation with separate objects representing a driver, software package and compiler toolkit.
The driver, framework runtime and development Toolkit describe different parts of a CUDA setup.

Diagnose from the same Python process that failed

Now run this code in the same terminal or notebook where the model failed. It reports the Python executable, PyTorch build and CUDA availability in the current process. This read-only diagnostic neither creates tensors nor performs model computation, and it does not change files.

First check that the Python path in the output is the environment that launched the model. If it points to a different virtual environment6, correct the IDE’s interpreter and run the diagnostic again. If PyTorch CUDA build is None, the imported PyTorch is likely not a CUDA build. torch.version.cuda reports the CUDA version for this PyTorch build; torch.cuda.is_available() reports whether this Python process can use CUDA.

The CUDA support level shown by nvidia-smi can differ from the PyTorch build version. Since CUDA 11, minor-version compatibility is available within a major release when minimum driver requirements are met, with some feature limits. New features, GPU architectures, PTX compilation and library combinations may have additional requirements. If the diagnostic returns False or an error points to the driver, check that combination in PyTorch’s install guide and NVIDIA’s compatibility table.

Read-only CUDA availability diagnostic
import sys
import torch

print("Python:", sys.executable)
print("PyTorch:", torch.__version__)
print("PyTorch CUDA build:", torch.version.cuda)
print("CUDA available:", torch.cuda.is_available())

if torch.cuda.is_available():
    print("GPU:", torch.cuda.get_device_name(0))
    print("Device count:", torch.cuda.device_count())
This does not start model computation. Run it in the same Python environment that failed.
A blank device card, graphics card and developer Toolkit case arranged on a workstation.
Check that the driver sees the GPU before troubleshooting the framework or compiler.

Check which device holds the model parameters

CUDA available: True and a GPU name mean this Python process found a CUDA device. Next check where the model parameters live. For a single-device model, inspect a parameter with next(model.parameters()).device. The code may select CPU, or the model and tensors may be on different devices. A standard PyTorch nn.Module has no universal .device attribute, so for a distributed model check the app logs and distributed settings instead of relying on the first parameter.

Run the same document summary and check whether the app log selects the GPU or reports an error or CPU fallback. A short request may register as low GPU utilization in the system monitor. VRAM7 use only suggests that the model was loaded; it does not show that computation is underway, so check the parameter device and app log together.

PyTorch’s ROCm8 build also uses the torch.cuda.* namespace APIs as compatibility interfaces. To identify AMD GPU use, check torch.version.hip9 and the current PyTorch/ROCm installation guidance. CUDA-only libraries or custom operations may not work in another build. Framework device detection and support for every project feature are separate matters.

A local AI workstation with a text-free monitor and an unmarked graphics card.
A visible GPU and a successful framework check still need confirmation in the application’s own logs.

Change one thing based on the result

If nvidia-smi cannot see the device, check that the card is connected and that the operating system driver supports the GPU and OS. With WSL2 on a Windows host, check both the host driver and guest environment. Changing PyTorch repeatedly before resolving this can add new errors and obscure the original cause.

If the GPU and driver are visible but torch.cuda.is_available() is False, check the failed process’s Python path and PyTorch build. Review the project’s required version and CUDA platform, then select your OS, package manager, Python and compute platform in PyTorch’s official install selector to get a command for your environment. Apply the instructions inside the project’s virtual environment and rerun the same code. Avoid relying on an old CUDA wheel URL or deleting global packages in bulk.

If CUDA is True but the model remains on CPU, stop reinstalling the framework and inspect the app and model device settings. Check whether the app includes a CUDA backend, whether GPU offloading10 is enabled, and whether the model and context fit in VRAM with execution headroom. If VRAM is insufficient, some work may move to the CPU or execution may fail. Follow the current instructions for that app or model repository for exact actions and option names.

Rerun the same summary; record these values if it still fails

After changing one driver or PyTorch setting, summarize the same document again in the same IDE notebook. Check the app log for the selected GPU and backend and for errors during processing. If the problem is resolved, keep that setup. If the same failure remains, gather the diagnostic details below before asking for help.

Share the GPU model and operating system, driver version and CUDA display from nvidia-smi, Python path from the failing notebook, PyTorch version, torch.version.cuda, is_available() result, selected app device and error message. Redact personal directory names from the path before sharing. Full environment variables or a successful result from another Python session are not needed to explain this failure.

Terminology notes

  1. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  2. PyTorch — A software framework for building and running AI models. Check the compatible PyTorch version and hardware support along with the model.

    Back to the text
  3. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text
  4. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  5. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  6. Python virtual environment — An isolated space for installing Python packages per project. It helps reduce version conflicts and is not a virtual machine.

    Back to the text
  7. VRAM — Memory used by a graphics card’s GPU for model weights and intermediate values. It is distinct from system RAM.

    Back to the text
  8. ROCm — AMD’s software platform for AI and high-performance computing on GPUs. Support depends on the combination of GPU, operating system, driver, and framework versions.

    Back to the text
  9. HIP — An API and runtime for writing portable GPU C++ code. It helps with CUDA porting but does not guarantee compatibility of every library or operation, or equivalent speed.

    Back to the text
  10. Offloading — Moving some model data from GPU memory to system RAM or storage when capacity is limited. This adds data transfer.

    Back to the text