Local speech recognition & synthesis

Orukeet local transcription: transcribe English and European-language audio

Orukeet is a local ASR model for transcribing English and European-language audio files.

Orukeet is a 0.6B-class ASR1 model that converts recordings to text. It supports 25 languages centered on English and European languages, and its roughly 714 MB Q8 GGUF2 can run with Metal3 on Mac. Use the installer to prepare the model and runtime4, then transcribe a short WAV file.

Requirements and key details
  • Supports 25 languages including English, French, and German. Korean and Japanese are not in the supported list.
  • The official native installer verifies Q8 weights and selects Apple Silicon Metal, detected CUDA, or CPU paths.
  • The 9.85% pooled WER on 25-language FLEURS used 20,146 recordings, greedy-batch TDT, FP32 weights, and CUDA BF16 autocast.

What languages does Orukeet transcribe, and what does it return?

Orukeet is an automatic speech recognition (ASR) model that takes a recording and returns recognized words as text. Released on September 11, 2026, it is a 0.6B-class model based on NVIDIA Parakeet TDT 0.6B v3, according to its model card. Its main purpose is transcription, not speaker diarization5 or emotion scoring.

The model card lists 25 languages: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, and Ukrainian. Korean and Japanese are not listed. If Korean recordings are your main workload, choose an ASR system that supports Korean.

A portable voice recorder sits beside a laptop showing an audio waveform
First check that the recording language appears in Orukeet’s supported list.

Choose an inference format before judging by hardware

The NeMo source is closest to the research and official evaluation path; native GGUF is a simpler starting point for local transcription apps or Mac use. The native Q8 file is 714,456,704 bytes (about 714 MB), and F16 is about 1.30 GB. The ONNX INT8 archive is about 487 MB and expands to about 672 MB. Download size is not runtime RAM6 or VRAM7 usage, so it should not be presented as a memory requirement.

The native package can select Metal on Apple Silicon, CUDA8 when it detects an NVIDIA device, and CPU9 otherwise. The choice still depends on the runtimes available on the system, so keep the installation JSON’s `device`, `model`, and `runtime` values to identify the path used on later runs.

Check model files and usage rights separately from code
PathApproximate artifactKey note
NeMo2,509,342,720 bytesOfficial evaluation path; requires NeMo
native Q8 GGUF714,456,704 bytesMetal, CUDA, or CPU runtime
ONNX INT8486,807,585 bytes compressed; about 672 MB extractedsherpa-onnx path

Check model files and usage rights separately from code

NeMo

Approximate artifact
2,509,342,720 bytes
Key note
Official evaluation path; requires NeMo

native Q8 GGUF

Approximate artifact
714,456,704 bytes
Key note
Metal, CUDA, or CPU runtime

ONNX INT8

Approximate artifact
486,807,585 bytes compressed; about 672 MB extracted
Key note
sherpa-onnx path

Start native transcription with Python 3.12 or later

The official model card specifies Python 3.12 or later and the v0.1.1 wheel. Install the wheel in an activated virtual environment10 and use `orukeet install` to fetch the Q8 weights and native runtime. In an application, keep the `Orukeet` object alive and call `transcribe()` for each file instead of reloading the model every time.

Begin with a short WAV file and inspect both the returned `text` field and the installation JSON. If words are missing, first check that the language is supported, the file can be opened, and the selected runtime is installed. The Python snippet follows the model card’s interface; it was not executed as part of this writing task.

Install the official v0.1.1 wheel
python -m pip install --upgrade \
  https://github.com/Oruk-AI/orukeet/releases/download/v0.1.1/orukeet-0.1.1-py3-none-any.whl
orukeet install --device auto --cache ./orukeet-cache --output installation.json
Transcribe one audio file
import json
from pathlib import Path
from orukeet import Orukeet

config = json.loads(Path("installation.json").read_text(encoding="utf-8-sig"))
with Orukeet(config["model"], config["runtime"], device=config["device"]) as asr:
    result = asr.transcribe("recording.wav")
    print(result["text"])
An open handwritten notebook sits beside a laptop displaying an audio waveform and transcript workspace
Check omissions and punctuation on a short segment before processing longer files.
Transcription desk with a compact Mac and monitor
Keep the Metal runtime and device values from the installation JSON to identify the path used.

Read the FLEURS score with its evaluation conditions

Oruk AI reports that pooled WER on 25-language FLEURS changed from a baseline of 11.01% to 9.85% with Orukeet. The evaluation contains 20,146 audio recordings and uses the same NeMo greedy-batch TDT path, FP3211 weights, and CUDA BF1612 autocast. The 9.85% figure belongs to that dataset and inference13 setup; it is not an error-rate guarantee for every recording.

The paper states that LibriSpeech test-other was used for further adaptation and checkpoint14 selection. Since this evaluation material also informed model development, take that into account when interpreting results. For your recordings, compare missing or incorrect words using the same file and runtime settings.

FLEURS aggregate across 25 languages with the same NeMo evaluation path.
ModelWord error rate (WER)Evaluation conditions
Parakeet TDT 0.6B v311.01%20,146 recordings, greedy-batch TDT, FP32 weights and CUDA BF16 autocast
Orukeet9.85%Same conditions

FLEURS aggregate across 25 languages with the same NeMo evaluation path.

Parakeet TDT 0.6B v3

Word error rate (WER)
11.01%
Evaluation conditions
20,146 recordings, greedy-batch TDT, FP32 weights and CUDA BF16 autocast

Orukeet

Word error rate (WER)
9.85%
Evaluation conditions
Same conditions

Check the code and weight licenses separately

The repository code is MIT-licensed, while the Hugging Face model card lists the weights and fitted kernels under CC BY-SA 4.0. If you modify or redistribute the model or weights, checking only the code license is insufficient. Review the weight terms separately, including attribution and share-alike requirements.

If your goal is local transcription of English or European-language recordings and you need transcript text with segment timing, inspect the native output and test a short sample. If speaker identity, emotion analysis, or Korean support is central, those needs fall outside the published scope and another tool is a better fit.

If the first run fails, check language, file, and runtime

If download verification or runtime selection fails, check whether `installation.json` was created and whether its `device` and `runtime` paths match the current OS. If the model file is missing, inspect write permission on the `--cache` directory and confirm the Q8 file finished downloading. Because the native installer checks both model and runtime, do not start by changing the Python code before this step is resolved.

If the transcript is empty, first verify that the recording is in one of the model’s 25 listed languages and that its format and audio track can be read. If a transcript appears but omits words, test the same short segment in a supported language and separate input audibility from transcription behavior. Check performance in another language with recordings in that language.

Terminology notes

  1. Automatic Speech Recognition — Technology that recognizes spoken language in audio and converts it to text.

    Back to the text
  2. GGUF — A file format for model data, widely used by llama.cpp-based tools.

    Back to the text
  3. Metal — Apple’s low-level technology for graphics and parallel GPU computation.

    Back to the text
  4. Runtime — The software environment that provides facilities needed while a program runs. In local AI, it can also refer to a model execution engine.

    Back to the text
  5. Speaker diarization — The task of dividing a recording into who-spoke-when segments and labeling them by speaker.

    Back to the text
  6. System RAM — System memory that temporarily holds data while programs run.

    Back to the text
  7. VRAM — Memory used by a graphics card’s GPU to store model weights and intermediate values.

    Back to the text
  8. CUDA — A software platform for general-purpose computing on NVIDIA GPUs.

    Back to the text
  9. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  10. Python virtual environment — An isolated space for installing Python packages per project, helping reduce package-version conflicts.

    Back to the text
  11. FP32 — A 32-bit floating-point format. It uses more memory per value than BF16 and can represent values more precisely.

    Back to the text
  12. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  13. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  14. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text