From recording to text. From text to voice.

Compare the wait across devices and runtimes.

qwen3-asr-rs · Metal · BF16

Model languages: KO · EN · JA +

Input recording · shared English sample

10.2 s · English · Qwen3-TTS 1.7B CustomVoice · Aiden

Estimated processing time

2.4 s

Processing speed vs playback

4.3×

Estimated-timing demonstration · no live model inference.

Start to see the processing sequence.0.0 s / 2.4 s

0%

A shared, pre-generated example—not a comparison of the selected models’ voices or recognition accuracy.

Timing method and runtime conditions

Published processing ratios are applied to the sample duration. Download and model loading are excluded; short-clip overhead, language and text can change actual time.

M4 · Metal · BF16 · EN/ZH · 3–36s · 3 runs/sample · model already loaded

RTF = 0.230 · 10.2 s × 0.2302.4 s

Transcript reveal is illustrative, not measured word-level timing.

Sample audio credits

For longer audio

1 min audio

13.8 s

10 min audio

138.0 s

60 min audio

828.0 s

A simple projection using the same processing ratio.

Which speech model fits your task?

The guides also compare Whisper turbo, Parakeet, Kokoro and Chatterbox by language and runtime. The timing demo includes only configurations with published timing evidence.