Independent tools and practical guides. · Free app downloads with an account.
NeuroDynamic.TechAccount

Run Whisper on an NVIDIA GPU

The recorded CPU dictation setup took around twelve seconds to transcribe a short clip. Using the server’s NVIDIA GPU brought a comparable request below half a second. These are results from that deployment; another model or machine may behave differently.

The GPU commands below come from the original server setup. They have not been retested in the guide VM because it has no NVIDIA card.

Why use a GPU?

GPUs can perform many of the calculations used by speech models in parallel. Performance also depends on the model, precision, available graphics memory and the rest of the pipeline.

Before changing the server

Get Run your own dictation server working first. For the GPU version you also need a supported NVIDIA card, a compatible driver and the NVIDIA Container Toolkit configured for Docker.

The command used in this build

This example makes the GPU available to the container and selects the GPU configuration. Check the selected model’s memory requirements before using it on a different card.

sudo docker run -d --name whisper \
  --gpus all \
  -p 9000:9000 \
  -v /srv/whisper/data:/var/lib/whisper \
  -e WHISPER_MODEL=large-v3-turbo \
  -e WHISPER_DEVICE=cuda \
  -e WHISPER_COMPUTE=float16 \
  -e WHISPER_API_KEY=$(cat ~/whisper-api-key.txt) \
  --restart unless-stopped \
  hwdsl2/whisper-server:cuda

A compatibility error encountered during setup

One deployment restarted instead of loading the model, with this message:

CUDA forward compatibility was attempted on non supported HW

A CUDA error needs checking against the exact driver, libraries, container image and hardware. Error 803 indicates a driver-version mismatch; error 804 indicates an unsupported forward-compatibility configuration. They are not interchangeable diagnoses.

In this deployment, rebuilding the image on a compatible CUDA base resolved the issue. That does not establish that driver updates cannot help other machines. Check the supported combinations before changing either the driver or the image. NVIDIA’s compatibility documentation explains the error codes and hardware restrictions.

Recorded timings

The original comparison is retained below. It shows the results reported for that setup, rather than a speed guarantee.

Setup How long the answer took
Processor only about twelve to fourteen seconds
Graphics card under half a second

Check the result

Confirm that the server starts, accepts a recording and returns usable text. Measure the complete request, including model loading if it starts cold. Keep the working CPU configuration available while investigating the GPU setup.

Display settings

Changes apply immediately and stay in this browser.

Colour theme