From recording to text. From text to voice.
Compare the wait across devices and runtimes.
qwen3-asr-rs · Metal · BF16
Model languages: KO · EN · JA +
Input recording · shared English sample
10.2 s · English · Qwen3-TTS 1.7B CustomVoice · Aiden
Estimated processing time
≈ 2.4 s
Processing speed vs playback
4.3×
Estimated-timing demonstration · no live model inference.
0%
A shared, pre-generated example—not a comparison of the selected models’ voices or recognition accuracy.
Timing method and runtime conditions
Published processing ratios are applied to the sample duration. Download and model loading are excluded; short-clip overhead, language and text can change actual time.
M4 · Metal · BF16 · EN/ZH · 3–36s · 3 runs/sample · model already loaded
RTF = 0.230 · 10.2 s × 0.230 ≈ 2.4 s
Transcript reveal is illustrative, not measured word-level timing.
Sample audio creditsFor longer audio
1 min audio
≈ 13.8 s
10 min audio
≈ 138.0 s
60 min audio
≈ 828.0 s
A simple projection using the same processing ratio.
Which speech model fits your task?
The guides also compare Whisper turbo, Parakeet, Kokoro and Chatterbox by language and runtime. The timing demo includes only configurations with published timing evidence.