Back to basics
The ideas behind the guides, in plain words
A few terms used in the server and local-AI guides. Follow the links from an article, or browse the explanations below.
Running AI on your own computer
Running a model locally means your own computer processes the request, rather than sending it to a hosted model service. Other features in the app may still use a network connection, so check their settings separately.
A local setup gives you more control over the processing. You also provide the hardware, electricity, updates and maintenance.
A model
A model is a trained set of parameters used by software to produce an output, such as transcribed speech, text or an image. You download the model files and run them with compatible software.
Model size affects memory requirements. A larger model is not automatically better for every task; the training, software and way you run it matter too.
Tokens, and “tokens per second”
A token is a small chunk of text, roughly a short word or part of a word. AI programs read and write in tokens rather than whole sentences. “Tokens per second” is simply how fast the program produces its answer.
Tokens per second measures output generation. It does not include every delay, such as loading the model or processing a long input. Try a typical request to see what the complete wait feels like.
GPUs and VRAM (why the graphics card matters)
A GPU is a processor suited to doing many calculations in parallel, including graphics and much of the work used by AI models. Its dedicated memory is usually called VRAM.
The model and the working data need memory. Software can sometimes split work between the GPU and system memory, but this may reduce speed. Check the requirements of the model and its settings.
Self-hosting
“Self-hosting” means running a service yourself instead of paying someone else to run it for you. The dictation server in these guides is self-hosted: it is a program on your own computer that you start, rather than a website you sign into.
Self-hosting gives you control over a service’s configuration and data. You become responsible for access, updates, backups and keeping it available.
Whisper and speech-to-text
Speech-to-text turns spoken words into written ones. Whisper is a well-known, free speech-to-text model. When a guide here talks about “a Whisper server”, it means a small program running that model, waiting for audio and sending back the text.
Dictation can be useful for drafts or notes. Check the result for transcription errors. Recording retention depends on the server and client settings; running Whisper locally does not by itself guarantee deletion.
Ready to try one?
The build-log guides put these ideas to work, one command at a time, with the reasons for each step written out.
Read the how-to guides