Local speech recognition & synthesis

Creating a local voice assistant: ASR·LLM·MCP·TTS connection guide

Understanding speech and operating a computer are two different steps.

The local voice assistant connects ASR1, LLM2, tool execution, and TTS3 in sequence. You can make it quickly and safely by looking at the delays and permissions of each step separately.

Requirements and key details
  • Measure ASR, judgment, tools, and TTS delays separately.
  • The MCP server must be trusted with the same level of trust as any installed program.
  • Connect to read-only tools first.
  • Leave a confirmation question and cancellation path.

Divide a sentence into four steps

ASR converts the microphone input into text, and LLM determines the intent of the request. If you need to search for a file or schedule, call the MCP4 tool, and TTS will read the completed answer aloud. If you are late at any stage, the entire conversation will seem broken.

Initially, streaming ASR, short system prompts, small decision models, and short voice replies reduce round-trip time. A hierarchical configuration using larger models or longer TTS output is practical only when long descriptions are needed.

Voice assistant with microphone, local AI equipment, tools, and speaker
You can find slow sections by measuring the time for voice recognition, judgment, tool execution, and voice synthesis separately.

Speed is considered the time for the first partial response.

We record the time ASR waits for the end of the sentence, the first token5 for LLM, the tool execution, and the first audio for TTS, respectively. If you only look at the overall completion time, it is difficult to know which parts need to be changed.

You can reduce perceived waiting by turning over partial transcriptions before the end of speech and starting TTS once the first sentence of your LLM answer is complete. However, because intermediate results may change, command execution is performed after final transcription or user confirmation.

MCP server reviews like installing a local app

The local MCP server can use the file and program permissions that the Run As account has. The source and code are verified, only necessary directories are accessed, and servers exposed to the public network are authenticated.

Initially, only link tools that are easy to revert, such as search and lookup. Deleting files, making payments, and sending external messages are confirmed before execution, and requests and results are left in a user-readable format.

Scene where local AI is connected to tool modules with different permissions
Open tools that are easy to undo, such as searches, and restrict transfer and change permissions separately.

The first version connects one safe thing all the way to the end

Decide on a flow with clear inputs and outputs, such as ‘transcribing the meeting recording and reading today’s tasks.’ ASR accuracy, tool success rate, and first voice response time are measured repeatedly using the same scenario.

After the flow is stable, add tools with different permissions one by one, such as calendar, document search, and home automation. It's more about being able to stop when something fails and tell the user what didn't work, rather than the number of features.

A physical study with a microphone, local AI computer, and speakers
The first version focuses on stabilizing a single flow to the end, such as meeting records or calendar views.

Terminology notes

  1. Automatic Speech Recognition — Technology that recognizes spoken language in audio and converts it to text. The recognized text represents the utterance, not a reconstruction of the audio.

    Back to the text
  2. Large language model — A language model trained on large text datasets to process and generate text. Capabilities and supported inputs vary by model.

    Back to the text
  3. Text-to-Speech — Technology that converts text into spoken audio with pronunciation and prosody. Output characteristics depend on the model and settings.

    Back to the text
  4. MCP — A protocol for connecting clients to tools and data sources a model may use. Connectivity and permission to execute a particular tool must still be configured separately.

    Back to the text
  5. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text