Run your own dictation server
Original tested setup: I ran every command in this guide top-to-bottom on a fresh Ubuntu 24.04 virtual machine (4 CPUs, 8 GB RAM) on 13 August 2026 — and originally on two clean Ubuntu 22.04 machines. Versions at test time: Docker 29.1.3 (Ubuntu
docker.iopackage);hwdsl2/whisper-server:latestimage dated 2026-08-06; Whisper modelsmall. If something's broken since, tell me.
Setup notes reviewed on 18 September 2026 against the image maintainer’s documentation. The safer key-file and network settings below, and the Murmur connection steps, have not had a fresh end-to-end installation test. The dated result above is the original test, not a new acceptance result.
Dictation can be useful for getting a draft onto the page. Running the transcription service yourself gives you control over where it processes the recording, but you also take on its setup and maintenance.
This guide builds a small server and tests it with an audio clip. The basics page explains unfamiliar terms.
What you are building
A service on your own computer receives an audio clip and returns text. The model runs locally. Keep the service on an appropriate private network and check the settings of any app you connect to it.
You can connect Murmur to this service. Transcription then runs on your server, with no Murmur word limit or transcription fee.
Why these choices
- Whisper provides the speech-recognition model. Check its output: it can mishear speech or add words, especially with unclear audio.
- Docker packages the server and its dependencies in a container, which makes this setup easier to repeat.
- A CPU setup lets you begin without a dedicated GPU. Processing may take longer than the recording itself.
What you need
- A computer running Ubuntu (I used 22.04, and re-tested on 24.04 — every command worked unchanged). Other systems work too; the commands differ slightly.
- About 1.5 GB of free disk space for the program and the speech model. Mine used 1.3 GB.
- Time for the installation and model download; connection speed will affect the wait.
The testing note above records the machines and versions used for this walkthrough.
Step 1: install Docker
Install Docker on the Ubuntu machine:
sudo apt-get update
sudo apt-get install -y docker.io
Check the installation:
docker --version
sudo systemctl is-active docker
Look for a Docker version and an active service. These commands use sudo to access Docker. A permission error means the current user does not have access to the Docker socket; it should not be ignored.
Step 2: make somewhere to keep the model
sudo mkdir -p /srv/whisper/data
The speech model is a biggish download. This folder is where it lives so it only downloads once.
Step 3: make a key so only you can use it
An API key restricts requests to clients that have it. Generate a key and keep its saved file private:
(umask 077; set -C; openssl rand -hex 24 > ~/whisper-api-key.txt)
The key is saved in ~/whisper-api-key.txt, readable only by your user. This command refuses to overwrite an existing file. If you already have a key, keep it and check its permissions with chmod 600 ~/whisper-api-key.txt. Treat the key like a password; do not post it in screenshots or support messages.
Step 4: start the server
One command does it:
sudo docker run -d --name whisper \
-p 127.0.0.1:9000:9000 \
-v /srv/whisper/data:/var/lib/whisper \
-e WHISPER_MODEL=small \
-e WHISPER_DEVICE=cpu \
-e WHISPER_API_KEY=$(cat ~/whisper-api-key.txt) \
--restart unless-stopped \
hwdsl2/whisper-server:latest
You do not need to memorise that. In plain terms it says: run the Whisper
program in the background, let it be reached on port 9000 from this server only, keep the model in
the folder from Step 2, use the model size named small on the processor
(Whisper comes in sizes from tiny to large; small balances speed
against accuracy), protect it with the key from Step 3, and start it again automatically if the machine
reboots.
The first time it starts it downloads the model, so give it a minute. Watch it get ready:
sudo docker logs whisper
Wait for the line Whisper speech-to-text server is ready. When I first
built this I rebooted the test machine on purpose to check the "start again
automatically" part, and it came back on its own. My latest re-test did not
repeat the reboot, so treat that as a 2025 result, not a 2026 one.
One thing to ignore: the startup log prints a suggested address that happens
to be your public internet address. Do not use it and do not open this up to
the internet. On your own network, localhost or your machine's local
address is all you need. The log also warns that you are "sending
unauthenticated requests to the HF Hub" and suggests an HF_TOKEN. That is
a nag about download speed, not a problem. Ignore it too.
Step 5: test it with your voice
Put a short audio clip in your home folder. A phone voice memo is fine. Then:
curl -s http://localhost:9000/v1/audio/transcriptions \
-H "Authorization: Bearer $(cat ~/whisper-api-key.txt)" \
-F [email protected] \
-F model=whisper-1
The request sends the recording and prints the transcription. First check that it works on the server itself.
Connect Murmur to your server
The command above binds to 127.0.0.1: only this server can connect. To use a different Windows PC, arrange a private connection first. For a trusted LAN, recreate the container with -p YOUR_PRIVATE_SERVER_IP:9000:9000, replacing the placeholder with the server’s private interface address. Limit access to the PCs that need it. HTTP does not encrypt the key or audio; use a secured private tunnel or HTTPS where the network is not trusted. Do not forward this port from your router.
- In Murmur, choose an own-server profile and open its connection settings.
- Enter the reachable base URL, including
/v1: for example,http://192.168.1.50:9000/v1if that is your server’s private address. On a different Windows PC,localhostmeans that PC, not your server. - Enter the private API key you made above. Keep Murmur hosted-service keys out of this profile.
- Use a model name accepted by your server. This image documents
whisper-1as the API name; it uses the server’s configured model,smallin this example. - Record a short test in Murmur and check the returned text. If it fails, check the address, key and private network connection before trying a longer recording.
Murmur does not bundle Whisper. The service must be running and reachable whenever you dictate. Other compatible server implementations may need different model names or settings.
Recorded processing time
On the recorded four-core CPU test, a three-second clip took about eleven seconds to return. This setup processes a submitted recording; it does not stream words onto the page as you speak.
The Run Whisper on an NVIDIA GPU guide describes a faster setup and a compatibility problem encountered along the way.
When something goes wrong
permission denied while trying to connect to the docker APImeans you left offsudo.- An error about a missing authorisation header means the request did not include the expected key. Check the client configuration.
- Lost your key? Read it back with
cat ~/whisper-api-key.txt.
Remove the practice setup
sudo docker rm -f whisper
sudo docker rmi hwdsl2/whisper-server:latest
sudo rm -rf /srv/whisper
These commands remove the container and the named model directory. Docker remains installed, and the API-key file created earlier remains unless you remove it separately. Check the paths before removing any files.