Model setup recipes
Qwen3-Omni Local Setup Conditions: Testing Audio and Video Paths
Multimodal capability does not mean your local runtime supports every path.
Qwen3-Omni Instruct includes thinker and talker components and supports text, audio, and video input with text or audio output. Model-card capability is not the same as local backend support. The guide uses the official Transformers path on Linux CUDA1 for one local WAV file; test audio, text, and video paths separately and check memory against the exact model and input.
Choose the desired input/output combination first
If you want to use Qwen3-Omni as a local voice assistant or for video summaries, first decide what combination of inputs and outputs you need. The required model path differs depending on whether you want to take microphone speech and answer in text, reply with generated speech, or send a short video with a question to summarize a scene. Do not assume every runtime2 supports every combination simply because the model is multimodal.
The official Instruct configuration includes thinker and talker components. The thinker interprets text, audio, and video signals and forms the response; the talker path is involved in speech generation. Start by confirming text input/output, then add audio or video token3 paths. This makes it easier to identify which stage fails.
You need network access to the model card, an installation environment for the chosen backend, storage and sufficient memory, plus a test audio or video file. Begin with a short, non-sensitive sample, not a real work file or recording containing personal information. After downloading the weights and input file, separately verify whether the test can run with the network disconnected; this helps confirm that data is not sent elsewhere.
| Task | Path to check first | Validation question |
|---|---|---|
| Text dialogue | Thinker input/output | Do the model and template load? |
| Audio understanding | Audio-input preprocessing and recognition | Is the spoken question transcribed or interpreted correctly? |
| Speech response | Talker audio output | Can the generated audio be saved and played? |
| Video question | Video-frame sampling and thinker | Are timing and key scenes reflected? |
Separate model capabilities, execution paths, and task types when planning tests.
Text dialogue
- Path to check first
- Thinker input/output
- Validation question
- Do the model and template load?
Audio understanding
- Path to check first
- Audio-input preprocessing and recognition
- Validation question
- Is the spoken question transcribed or interpreted correctly?
Speech response
- Path to check first
- Talker audio output
- Validation question
- Can the generated audio be saved and played?
Video question
- Path to check first
- Video-frame sampling and thinker
- Validation question
- Are timing and key scenes reflected?

Establish backend support before proceeding
The verifiable reference path is to load the official `Qwen3OmniMoeForConditionalGeneration` and processor with Hugging Face Transformers on a Linux NVIDIA GPU4, then submit one audio file. The Qwen team specifies Transformers 5.2.0 or later, `accelerate`, `qwen-omni-utils`, and ffmpeg; FlashAttention5 2 is an optional choice for compatible CUDA GPUs using FP16/BF16. Qwen also lists vLLM-Omni as an option for larger serving workloads.
The example below is not an Apple Silicon support path. Official support for Qwen3-Omni on macOS/MLX and compatibility across all audio/video features have not been confirmed, so do not assume this combination will run. If you are targeting an Apple device, wait until an official converted checkpoint6 and processor plus audio/video backend support are confirmed.
This is a 30B-A3B-class checkpoint and requires substantial GPU memory and local storage. The official memory table is for BF167 with Transformers and FlashAttention 2. Before installing, check the GPU, CUDA/PyTorch combination, and checkpoint storage needs. Changing precision or backend can change both requirements and supported functionality.
python -m venv .venv
source .venv/bin/activate
python -m pip install -U 'transformers>=5.2.0' accelerate qwen-omni-utils soundfile
# Install a PyTorch build matched to your CUDA and GPU from the official PyTorch selector.
python -m pip install -U flash-attn --no-build-isolation
ffmpeg -versionComplete one short audio round trip first
Save a short, non-sensitive WAV file as `sample.wav`, then run the code below. The conversation uses a local file path; `process_mm_info` prepares the audio input and the thinker returns a text answer. The first run can take a long time because it downloads and loads a large checkpoint, and it can use substantial GPU memory. If the file cannot be opened, check its codec8 and sample rate9, ffmpeg, and `qwen-omni-utils` first.
Next, test speech output through the path that includes the talker. Check whether the same answer can be returned as both text and audio, and whether the generated file plays. If you receive only text, compare the request format, backend limitations, and model configuration with the official example. Enabling every audio input/output path at once makes it difficult to tell whether recognition or speech generation failed.
Record model-load time, input processing, time to the first text response, and total response time separately. Distinguish cold and warm runs10 and repeat with the same file. The code below requests text output only and does not include speech output from the talker.
from transformers import Qwen3OmniMoeForConditionalGeneration, Qwen3OmniMoeProcessor
from qwen_omni_utils import process_mm_info
model_id = 'Qwen/Qwen3-Omni-30B-A3B-Instruct'
model = Qwen3OmniMoeForConditionalGeneration.from_pretrained(
model_id, dtype='auto', device_map='auto', attn_implementation='flash_attention_2'
)
processor = Qwen3OmniMoeProcessor.from_pretrained(model_id)
messages = [{'role': 'user', 'content': [
{'type': 'audio', 'audio': './sample.wav'},
{'type': 'text', 'text': 'Summarize the spoken message in one sentence.'},
]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
audios, images, videos = process_mm_info(messages, use_audio_in_video=False)
inputs = processor(text=text, audio=audios, images=images, videos=videos,
return_tensors='pt', padding=True, use_audio_in_video=False)
inputs = inputs.to(model.device).to(model.dtype)
text_ids, _ = model.generate(**inputs, return_audio=False, thinker_return_dict_in_generate=True)
answer = processor.batch_decode(
text_ids.sequences[:, inputs['input_ids'].shape[1]:],
skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(answer)
Increase video length while tracking frame count
After the audio test succeeds, ask a question about a short video clip of about five seconds. The path that decodes and passes frames11 to the model matters, as does temporal information. Even videos with the same file size can produce different input volumes if frame-sampling rate or resolution differs. Start with a short clip using the preprocessing and frame-sampling settings in an official example.
The official model-card memory figures are theoretical and assume BF16 Transformers with FlashAttention 2. The card gives examples of about 78.85 GB for 15 seconds of video and about 144.81 GB for 120 seconds. These are not actual minimum-memory requirements for every backend, nor results for a small quantized model. Treat them as a warning that longer videos can require substantially more memory.
Instead of increasing only clip length, change frame rate and resolution one at a time, and record input-frame count and peak memory. If temporal order matters, ask about the sequence of events rather than only “What is in the last frame?” A change in frame sampling changes what the model sees, so speed and accuracy results from different settings are not directly comparable.
| Task | Path to check first | Validation question |
|---|---|---|
| Text dialogue | Thinker input/output | Do the model and template load? |
| Audio understanding | Audio-input preprocessing and recognition | Is the spoken question transcribed or interpreted correctly? |
| Frame sampling | Talker audio output | Can the generated audio be saved and played? |
| Video question | Video-frame sampling and thinker | Are timing and key scenes reflected? |
Separate model capabilities, execution paths, and task types when planning tests.
Text dialogue
- Path to check first
- Thinker input/output
- Validation question
- Do the model and template load?
Audio understanding
- Path to check first
- Audio-input preprocessing and recognition
- Validation question
- Is the spoken question transcribed or interpreted correctly?
Frame sampling
- Path to check first
- Talker audio output
- Validation question
- Can the generated audio be saved and played?
Video question
- Path to check first
- Video-frame sampling and thinker
- Validation question
- Are timing and key scenes reflected?

Locate bottlenecks in latency and memory
Multimodal dialogue may involve file preprocessing, vision/audio encoders, thinker prefill12, generation decode13, and talker output. If you record only total time, you may conclude that “the model is slow” without knowing which stage caused the delay. Record only detailed timings available in the logs; do not invent unsupported monitoring figures.
If memory runs short, separate model weights, input length, video-frame count, concurrent conversations, and other-app usage. First test whether shortening the video stabilizes the run. If audio and video do not need to be processed together, compare separate paths. Do not treat the model card’s theoretical memory condition as the same value as runtime memory use.
For real-time dialogue, prioritize time to first response and when speech output begins. For batch work such as summarizing long recordings, focus more on total processing time and memory stability. If the task fails on a particular hardware/backend combination, label it unsupported and consider a smaller model or an officially supported serving environment.
Official model, runtime, and license references
Modality and backend support can change with model and runtime releases. Before deployment, recheck required versions, usage terms, and limitations in the official repository and model card. This guide does not measure local-device performance.
Terminology notes
CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textFlashAttention — An implementation that improves memory access in attention computation. Availability and effects depend on hardware, model, and runtime.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textBF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textAudio codec — A method or software for encoding, compressing, and decoding digital audio. Size, compatibility, and lossiness vary by codec.
Back to the textSampling rate — The number of times per second an audio signal is measured when digitized. It is distinct from bit depth.
Back to the textWarm run — A measurement made after model loading and initialization. It may exclude the loading wait from the first run.
Back to the textFrame — A single image that makes up part of a video. Frame rate and frame resolution are separate properties.
Back to the textPrefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.
Back to the textDecode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.
Back to the text