Local speech recognition & synthesis

Edit Audio Locally with AuK: Speech Revisions, Cleanup, and TTS

AuK is a local model that turns an instruction and optional source audio into generated or edited speech.

AuK is a 1.5B model that generates speech or edits recordings from instructions. Provide a source file and replacement words to revise speech, or request noise and reverberation reduction. This guide starts with Python setup on CUDA1 and short-file editing, and also points to MLX2 on Mac and audio.cpp.

Requirements and key details
  • AuK-Base and four-step distilled AuK-Flash accept an instruction with text and optional source audio.
  • The official CLI requires Python 3.10 and separately downloaded AuK checkpoints and the Qwen2.5-Omni-3B encoder.
  • The official BF16 table reports 24.78 GiB without CPU offload and 16.75 GiB with it for 1.5-second text-only output on an A800-SXM4-80GB.

What does AuK take in and produce?

AuK is a 1.5B local model controlled by an instruction. Give it text to synthesize speech, or include an audio file to request a content edit or changes to delivery such as speed, emotion, or voice quality. Noise and reverberation reduction or source isolation use the same instruction-based interface.

Start by deciding whether the task needs source audio. For new narration, provide the text and a voice description; for revising or cleaning a recording, provide the audio and the requested change. AuK is not an ASR3 system that returns only a transcript: its output is an audio file.

Choose between revising words, reducing noise, or generating new speech

AuK released its code and weights on September 9, 2026, with Base and four-step distilled Flash checkpoints4. In editing, the instruction describes what to change or preserve while reference audio supplies the speaker and delivery context. For cleanup, the instruction names the source to keep and noise to remove.

Official examples focus on English and Chinese. If Korean pronunciation or a particular delivery style matters, generate a short segment first and check pronunciation, edit boundaries, and remaining background sound before processing a longer recording.

A short audio waveform and an editing sentence appear on a laptop beside a microphone
Pair the source clip with the requested replacement and test a short segment first.

Prerequisites and installation

The official README specifies a Python 3.10 environment. Clone the repository, install the core inference5 package, and download AuK-Base and Qwen2.5-Omni-3B separately under `ckpts` with the Hugging Face CLI6. AuK-Flash is optional for a first Base inference check.

First install a PyTorch7 build that matches your GPU8. The next table shows the AuK README’s measurement conditions alongside its memory results.

For Apple Silicon, follow the MLX instructions on the official `feat/mlx-apple-silicon` branch. audio.cpp also supports AuK and Flash as experimental GGUF9 models. This guide's `--cpu_offload` command and A800 memory table concern CUDA; on Mac, use the installation documentation and model components for the chosen runtime10.

Install AuK and download the model components
git clone https://github.com/Tencent-Hunyuan/AuK
cd AuK
conda create -n auk python=3.10 -y
conda activate auk
pip install -e .
pip install -U "huggingface_hub[cli]"
hf download tencent/AuK --local-dir ./ckpts/AuK
hf download Qwen/Qwen2.5-Omni-3B --local-dir ./ckpts/Qwen2.5-Omni-3B

Start by replacing words in a short recording.

Use a short audio file that you have permission to process. The official example input is `assets/demo-input-audio/content-edit/content.wav`; for your own file, change the path and specify the phrase to replace. `--gen_seconds` sets the target output duration. If the result is clipped, increase that value first.

A successful run creates `out_content_edit.wav`. Compare the same segment with the source to hear whether the words changed and surrounding speech connects naturally. The second command adapts the official cleanup request into English. Check the output for reduced noise and reverberation, and ensure the voices you wanted to keep remain.

Official content-editing CLI
auk-infer \
  --audio assets/demo-input-audio/content-edit/content.wav \
  --instruction "Replace 'but accepting what we cannot have' with 'and living well with dreams unmet'." \
  --output out_content_edit.wav \
  --gen_seconds 7.0
Reduce speech noise and reverberation
auk-infer \
  --audio assets/demo-input-audio/vocal-extraction/vocal-1-input.wav \
  --instruction "Please restore this audio to clean vocals, preserve the original speakers, and remove noise and reverberation." \
  --output out_denoise.wav
Two audio waveforms, representing the source and edited versions, sit side by side on a laptop screen
Listen for the changed words, the transition around them, and the voice you intended to keep.

Under what conditions did CPU offload reduce memory?

The official README table records `torch.cuda.max_memory_allocated` during BF1611 inference on one NVIDIA A800-SXM4-80GB and compares CPU12 offload13 disabled and enabled. Its two input conditions are text-only generation of 1.5 seconds and generation with five seconds of reference audio.

The table lets you compare each model under the same A800 conditions. Do not treat these figures as minimum memory requirements for another GPU. Check input duration, precision, and memory used by other applications before choosing an offload setting.

Peak allocated memory by text/reference input on NVIDIA A800-SXM4-80GB with BF16
Model and inputOffload offOffload on
AuK-Base, text-only, 1.5 s output24.78 GiB16.75 GiB
AuK-Base, 5 s reference audio25.00 GiB16.98 GiB
AuK-Flash, text-only, 1.5 s output24.77 GiB16.75 GiB
AuK-Flash, 5 s reference audio24.97 GiB16.98 GiB

Peak allocated memory by text/reference input on NVIDIA A800-SXM4-80GB with BF16

AuK-Base, text-only, 1.5 s output

Offload off
24.78 GiB
Offload on
16.75 GiB

AuK-Base, 5 s reference audio

Offload off
25.00 GiB
Offload on
16.98 GiB

AuK-Flash, text-only, 1.5 s output

Offload off
24.77 GiB
Offload on
16.75 GiB

AuK-Flash, 5 s reference audio

Offload off
24.97 GiB
Offload on
16.98 GiB
Enable CPU offload for CUDA
auk-infer --cpu_offload \
  --audio assets/demo-input-audio/content-edit/content.wav \
  --instruction "Replace 'but accepting what we cannot have' with 'and living well with dreams unmet'." \
  --output out_content_edit.wav \
  --gen_seconds 7.0
An audio waveform is visible on a laptop beside a small speaker
For cleanup, name the sound to keep and the noise or reverberation to reduce.

Check the encoder license before commercial use.

AuK code is MIT-licensed. Its required `Qwen/Qwen2.5-Omni-3B` encoder uses `qwen-research`, with non-commercial research and evaluation terms. Before integration into a commercial service, check the conditions for the code and every model component.

AuK fits short spoken edits, voice-over drafts, and recording cleanup. Choose a dedicated ASR system when you need document transcription or speaker timestamps across long recordings; when Korean output matters, start with a tool you can evaluate on Korean speech samples.

Terminology notes

  1. CUDA — A software platform for general-purpose computing on NVIDIA GPUs.

    Back to the text
  2. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  3. Automatic Speech Recognition — Technology that recognizes spoken language in audio and converts it to text.

    Back to the text
  4. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  5. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  6. CLI — Short for Command-Line Interface: operating a program by entering commands in a terminal.

    Back to the text
  7. PyTorch — A software framework for building and running AI models. Check the compatible PyTorch version and hardware support along with the model.

    Back to the text
  8. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  9. GGUF — A file format for model data, widely used by llama.cpp-based tools.

    Back to the text
  10. Runtime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.

    Back to the text
  11. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  12. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  13. Offloading — Moving some model data from GPU memory to system RAM or storage when capacity is limited. This adds data transfer.

    Back to the text