Local speech recognition & synthesis
Edit Audio Locally with AuK: Speech Revisions, Cleanup, and TTS
AuK is a local model that turns an instruction and optional source audio into generated or edited speech.
AuK is a 1.5B model that generates speech or edits recordings from instructions. Provide a source file and replacement words to revise speech, or request noise and reverberation reduction. This guide starts with Python setup on CUDA1 and short-file editing, and also points to MLX2 on Mac and audio.cpp.
What does AuK take in and produce?
AuK is a 1.5B local model controlled by an instruction. Give it text to synthesize speech, or include an audio file to request a content edit or changes to delivery such as speed, emotion, or voice quality. Noise and reverberation reduction or source isolation use the same instruction-based interface.
Start by deciding whether the task needs source audio. For new narration, provide the text and a voice description; for revising or cleaning a recording, provide the audio and the requested change. AuK is not an ASR3 system that returns only a transcript: its output is an audio file.
Choose between revising words, reducing noise, or generating new speech
AuK released its code and weights on September 9, 2026, with Base and four-step distilled Flash checkpoints4. In editing, the instruction describes what to change or preserve while reference audio supplies the speaker and delivery context. For cleanup, the instruction names the source to keep and noise to remove.
Official examples focus on English and Chinese. If Korean pronunciation or a particular delivery style matters, generate a short segment first and check pronunciation, edit boundaries, and remaining background sound before processing a longer recording.

Prerequisites and installation
The official README specifies a Python 3.10 environment. Clone the repository, install the core inference5 package, and download AuK-Base and Qwen2.5-Omni-3B separately under `ckpts` with the Hugging Face CLI6. AuK-Flash is optional for a first Base inference check.
First install a PyTorch7 build that matches your GPU8. The next table shows the AuK README’s measurement conditions alongside its memory results.
For Apple Silicon, follow the MLX instructions on the official `feat/mlx-apple-silicon` branch. audio.cpp also supports AuK and Flash as experimental GGUF9 models. This guide's `--cpu_offload` command and A800 memory table concern CUDA; on Mac, use the installation documentation and model components for the chosen runtime10.
git clone https://github.com/Tencent-Hunyuan/AuK
cd AuK
conda create -n auk python=3.10 -y
conda activate auk
pip install -e .
pip install -U "huggingface_hub[cli]"
hf download tencent/AuK --local-dir ./ckpts/AuK
hf download Qwen/Qwen2.5-Omni-3B --local-dir ./ckpts/Qwen2.5-Omni-3BStart by replacing words in a short recording.
Use a short audio file that you have permission to process. The official example input is `assets/demo-input-audio/content-edit/content.wav`; for your own file, change the path and specify the phrase to replace. `--gen_seconds` sets the target output duration. If the result is clipped, increase that value first.
A successful run creates `out_content_edit.wav`. Compare the same segment with the source to hear whether the words changed and surrounding speech connects naturally. The second command adapts the official cleanup request into English. Check the output for reduced noise and reverberation, and ensure the voices you wanted to keep remain.
auk-infer \
--audio assets/demo-input-audio/content-edit/content.wav \
--instruction "Replace 'but accepting what we cannot have' with 'and living well with dreams unmet'." \
--output out_content_edit.wav \
--gen_seconds 7.0auk-infer \
--audio assets/demo-input-audio/vocal-extraction/vocal-1-input.wav \
--instruction "Please restore this audio to clean vocals, preserve the original speakers, and remove noise and reverberation." \
--output out_denoise.wav
Under what conditions did CPU offload reduce memory?
The official README table records `torch.cuda.max_memory_allocated` during BF1611 inference on one NVIDIA A800-SXM4-80GB and compares CPU12 offload13 disabled and enabled. Its two input conditions are text-only generation of 1.5 seconds and generation with five seconds of reference audio.
The table lets you compare each model under the same A800 conditions. Do not treat these figures as minimum memory requirements for another GPU. Check input duration, precision, and memory used by other applications before choosing an offload setting.
| Model and input | Offload off | Offload on |
|---|---|---|
| AuK-Base, text-only, 1.5 s output | 24.78 GiB | 16.75 GiB |
| AuK-Base, 5 s reference audio | 25.00 GiB | 16.98 GiB |
| AuK-Flash, text-only, 1.5 s output | 24.77 GiB | 16.75 GiB |
| AuK-Flash, 5 s reference audio | 24.97 GiB | 16.98 GiB |
Peak allocated memory by text/reference input on NVIDIA A800-SXM4-80GB with BF16
AuK-Base, text-only, 1.5 s output
- Offload off
- 24.78 GiB
- Offload on
- 16.75 GiB
AuK-Base, 5 s reference audio
- Offload off
- 25.00 GiB
- Offload on
- 16.98 GiB
AuK-Flash, text-only, 1.5 s output
- Offload off
- 24.77 GiB
- Offload on
- 16.75 GiB
AuK-Flash, 5 s reference audio
- Offload off
- 24.97 GiB
- Offload on
- 16.98 GiB
auk-infer --cpu_offload \
--audio assets/demo-input-audio/content-edit/content.wav \
--instruction "Replace 'but accepting what we cannot have' with 'and living well with dreams unmet'." \
--output out_content_edit.wav \
--gen_seconds 7.0
Check the encoder license before commercial use.
AuK code is MIT-licensed. Its required `Qwen/Qwen2.5-Omni-3B` encoder uses `qwen-research`, with non-commercial research and evaluation terms. Before integration into a commercial service, check the conditions for the code and every model component.
AuK fits short spoken edits, voice-over drafts, and recording cleanup. Choose a dedicated ASR system when you need document transcription or speaker timestamps across long recordings; when Korean output matters, start with a tool you can evaluate on Korean speech samples.
Terminology notes
CUDA — A software platform for general-purpose computing on NVIDIA GPUs.
Back to the textMLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.
Back to the textAutomatic Speech Recognition — Technology that recognizes spoken language in audio and converts it to text.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textInference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.
Back to the textCLI — Short for Command-Line Interface: operating a program by entering commands in a terminal.
Back to the textPyTorch — A software framework for building and running AI models. Check the compatible PyTorch version and hardware support along with the model.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textGGUF — A file format for model data, widely used by llama.cpp-based tools.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.
Back to the textBF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textCPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.
Back to the textOffloading — Moving some model data from GPU memory to system RAM or storage when capacity is limited. This adds data transfer.
Back to the text
Read next
Local speech recognition & synthesis
What is Qwen3-TTS? Turn text into speech on your computer
Local music generation
Create and revise ACE-Step music in a DAW: a local VST3 guide
Local speech recognition & synthesis
Orukeet local transcription: transcribe English and European-language audio
Local speech recognition & synthesis
Creating a local voice assistant: ASR·LLM·MCP·TTS connection guide
