Independent tools and practical guides. · Free app downloads with an account.
NeuroDynamic.TechAccount

Building a local voice assistant

This project connects a microphone to Home Assistant so spoken requests can control lights and other devices. Speech recognition and spoken replies run locally. The question was whether a language model added anything useful to everyday commands.

The three parts

  • Speech recognition turns the recording into text.
  • A command handler decides what action to take.
  • Text-to-speech reads the reply aloud.

The setup used a small board for speech input and output, with the main server available for a language model. Home Assistant connected the parts through its voice pipeline.

Trying a model on the small board

The first model ran, but generated about 0.4 tokens per second in the recorded test. A token is a short piece of text. At that rate, even a brief reply took too long for a light-switch command. Memory was limited too.

Moving the model to the server

The GPU server recorded about 165 tokens per second once the model was loaded. Shorter instructions and replies helped, and limiting the devices included in the request reduced unnecessary context.

Those figures describe different test setups, not a controlled hardware comparison. They also leave out the delay before generation begins: the server model took around twelve seconds to load from cold.

Using Home Assistant’s command handler

For the household commands tried here, Home Assistant’s built-in handling was enough. Sending those requests directly to it removed the model’s loading delay and let voice control continue while the main server was being updated.

Clear device names also helped. A name people naturally use is easier to ask for than an internal label chosen during installation.

What to take from the experiment

Start with the commands you want to use and measure the whole wait, from speaking to hearing a reply. Add a model if you need the conversation it provides and can accommodate its memory use and startup time.

This account is based on the original deployment notes. The small-board setup has not been freshly rebuilt for this preview. For the speech-recognition part, see Run your own dictation server.

Display settings

Changes apply immediately and stay in this browser.

Colour theme