Model setup recipes

Qwen3-Omni Local Setup Conditions: Testing Audio and Video Paths

Multimodal capability does not mean your local runtime supports every path.

Qwen3-Omni Instruct includes thinker and talker components and supports text, audio, and video input with text or audio output. Model-card capability is not the same as local backend support. The guide uses the official Transformers path on Linux CUDA1 for one local WAV file; test audio, text, and video paths separately and check memory against the exact model and input.

Requirements and key details
  • Thinker handles perception and reasoning; the talker is part of the audio-output path.
  • Separate model-card modality support from each inference backend’s support.
  • Test audio input, video input, text output, and audio output in separate small runs.
  • Record video duration and frame count because they affect memory and latency.

Choose the desired input/output combination first

If you want to use Qwen3-Omni as a local voice assistant or for video summaries, first decide what combination of inputs and outputs you need. The required model path differs depending on whether you want to take microphone speech and answer in text, reply with generated speech, or send a short video with a question to summarize a scene. Do not assume every runtime2 supports every combination simply because the model is multimodal.

The official Instruct configuration includes thinker and talker components. The thinker interprets text, audio, and video signals and forms the response; the talker path is involved in speech generation. Start by confirming text input/output, then add audio or video token3 paths. This makes it easier to identify which stage fails.

You need network access to the model card, an installation environment for the chosen backend, storage and sufficient memory, plus a test audio or video file. Begin with a short, non-sensitive sample, not a real work file or recording containing personal information. After downloading the weights and input file, separately verify whether the test can run with the network disconnected; this helps confirm that data is not sent elsewhere.

Separate model capabilities, execution paths, and task types when planning tests.
TaskPath to check firstValidation question
Text dialogueThinker input/outputDo the model and template load?
Audio understandingAudio-input preprocessing and recognitionIs the spoken question transcribed or interpreted correctly?
Speech responseTalker audio outputCan the generated audio be saved and played?
Video questionVideo-frame sampling and thinkerAre timing and key scenes reflected?

Separate model capabilities, execution paths, and task types when planning tests.

Text dialogue

Path to check first
Thinker input/output
Validation question
Do the model and template load?

Audio understanding

Path to check first
Audio-input preprocessing and recognition
Validation question
Is the spoken question transcribed or interpreted correctly?

Speech response

Path to check first
Talker audio output
Validation question
Can the generated audio be saved and played?

Video question

Path to check first
Video-frame sampling and thinker
Validation question
Are timing and key scenes reflected?
A local computer with a microphone, webcam, and speakers
Check peripheral and runtime compatibility separately from model capability.

Establish backend support before proceeding

The verifiable reference path is to load the official `Qwen3OmniMoeForConditionalGeneration` and processor with Hugging Face Transformers on a Linux NVIDIA GPU4, then submit one audio file. The Qwen team specifies Transformers 5.2.0 or later, `accelerate`, `qwen-omni-utils`, and ffmpeg; FlashAttention5 2 is an optional choice for compatible CUDA GPUs using FP16/BF16. Qwen also lists vLLM-Omni as an option for larger serving workloads.

The example below is not an Apple Silicon support path. Official support for Qwen3-Omni on macOS/MLX and compatibility across all audio/video features have not been confirmed, so do not assume this combination will run. If you are targeting an Apple device, wait until an official converted checkpoint6 and processor plus audio/video backend support are confirmed.

This is a 30B-A3B-class checkpoint and requires substantial GPU memory and local storage. The official memory table is for BF167 with Transformers and FlashAttention 2. Before installing, check the GPU, CUDA/PyTorch combination, and checkpoint storage needs. Changing precision or backend can change both requirements and supported functionality.

Install the official Transformers environment
python -m venv .venv
source .venv/bin/activate
python -m pip install -U 'transformers>=5.2.0' accelerate qwen-omni-utils soundfile
# Install a PyTorch build matched to your CUDA and GPU from the official PyTorch selector.
python -m pip install -U flash-attn --no-build-isolation
ffmpeg -version
Meet the official example's package requirements, then prepare a short WAV file. This is the Linux CUDA GPU path, not an MLX command for Mac.

Complete one short audio round trip first

Save a short, non-sensitive WAV file as `sample.wav`, then run the code below. The conversation uses a local file path; `process_mm_info` prepares the audio input and the thinker returns a text answer. The first run can take a long time because it downloads and loads a large checkpoint, and it can use substantial GPU memory. If the file cannot be opened, check its codec8 and sample rate9, ffmpeg, and `qwen-omni-utils` first.

Next, test speech output through the path that includes the talker. Check whether the same answer can be returned as both text and audio, and whether the generated file plays. If you receive only text, compare the request format, backend limitations, and model configuration with the official example. Enabling every audio input/output path at once makes it difficult to tell whether recognition or speech generation failed.

Record model-load time, input processing, time to the first text response, and total response time separately. Distinguish cold and warm runs10 and repeat with the same file. The code below requests text output only and does not include speech output from the talker.

Send one local WAV file and receive a text response
from transformers import Qwen3OmniMoeForConditionalGeneration, Qwen3OmniMoeProcessor
from qwen_omni_utils import process_mm_info

model_id = 'Qwen/Qwen3-Omni-30B-A3B-Instruct'
model = Qwen3OmniMoeForConditionalGeneration.from_pretrained(
    model_id, dtype='auto', device_map='auto', attn_implementation='flash_attention_2'
)
processor = Qwen3OmniMoeProcessor.from_pretrained(model_id)
messages = [{'role': 'user', 'content': [
    {'type': 'audio', 'audio': './sample.wav'},
    {'type': 'text', 'text': 'Summarize the spoken message in one sentence.'},
]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
audios, images, videos = process_mm_info(messages, use_audio_in_video=False)
inputs = processor(text=text, audio=audios, images=images, videos=videos,
                    return_tensors='pt', padding=True, use_audio_in_video=False)
inputs = inputs.to(model.device).to(model.dtype)
text_ids, _ = model.generate(**inputs, return_audio=False, thinker_return_dict_in_generate=True)
answer = processor.batch_decode(
    text_ids.sequences[:, inputs['input_ids'].shape[1]:],
    skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(answer)
`sample.wav` is a local file path. This Transformers example targets Linux CUDA GPUs and does not guarantee success on an untested system.
Video, an audio waveform, and a document enter a local AI interface
Test one input modality at a time before combining them.

Increase video length while tracking frame count

After the audio test succeeds, ask a question about a short video clip of about five seconds. The path that decodes and passes frames11 to the model matters, as does temporal information. Even videos with the same file size can produce different input volumes if frame-sampling rate or resolution differs. Start with a short clip using the preprocessing and frame-sampling settings in an official example.

The official model-card memory figures are theoretical and assume BF16 Transformers with FlashAttention 2. The card gives examples of about 78.85 GB for 15 seconds of video and about 144.81 GB for 120 seconds. These are not actual minimum-memory requirements for every backend, nor results for a small quantized model. Treat them as a warning that longer videos can require substantially more memory.

Instead of increasing only clip length, change frame rate and resolution one at a time, and record input-frame count and peak memory. If temporal order matters, ask about the sequence of events rather than only “What is in the last frame?” A change in frame sampling changes what the model sees, so speed and accuracy results from different settings are not directly comparable.

Separate model capabilities, execution paths, and task types when planning tests.
TaskPath to check firstValidation question
Text dialogueThinker input/outputDo the model and template load?
Audio understandingAudio-input preprocessing and recognitionIs the spoken question transcribed or interpreted correctly?
Frame samplingTalker audio outputCan the generated audio be saved and played?
Video questionVideo-frame sampling and thinkerAre timing and key scenes reflected?

Separate model capabilities, execution paths, and task types when planning tests.

Text dialogue

Path to check first
Thinker input/output
Validation question
Do the model and template load?

Audio understanding

Path to check first
Audio-input preprocessing and recognition
Validation question
Is the spoken question transcribed or interpreted correctly?

Frame sampling

Path to check first
Talker audio output
Validation question
Can the generated audio be saved and played?

Video question

Path to check first
Video-frame sampling and thinker
Validation question
Are timing and key scenes reflected?
A local PC with a GPU and visible memory headroom
Track memory as video duration and frame count increase.

Locate bottlenecks in latency and memory

Multimodal dialogue may involve file preprocessing, vision/audio encoders, thinker prefill12, generation decode13, and talker output. If you record only total time, you may conclude that “the model is slow” without knowing which stage caused the delay. Record only detailed timings available in the logs; do not invent unsupported monitoring figures.

If memory runs short, separate model weights, input length, video-frame count, concurrent conversations, and other-app usage. First test whether shortening the video stabilizes the run. If audio and video do not need to be processed together, compare separate paths. Do not treat the model card’s theoretical memory condition as the same value as runtime memory use.

For real-time dialogue, prioritize time to first response and when speech output begins. For batch work such as summarizing long recordings, focus more on total processing time and memory stability. If the task fails on a particular hardware/backend combination, label it unsupported and consider a smaller model or an officially supported serving environment.

Official model, runtime, and license references

Modality and backend support can change with model and runtime releases. Before deployment, recheck required versions, usage terms, and limitations in the official repository and model card. This guide does not measure local-device performance.

Terminology notes

  1. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text
  2. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  3. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text
  4. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  5. FlashAttention — An implementation that improves memory access in attention computation. Availability and effects depend on hardware, model, and runtime.

    Back to the text
  6. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  7. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  8. Audio codec — A method or software for encoding, compressing, and decoding digital audio. Size, compatibility, and lossiness vary by codec.

    Back to the text
  9. Sampling rate — The number of times per second an audio signal is measured when digitized. It is distinct from bit depth.

    Back to the text
  10. Warm run — A measurement made after model loading and initialization. It may exclude the loading wait from the first run.

    Back to the text
  11. Frame — A single image that makes up part of a video. Frame rate and frame resolution are separate properties.

    Back to the text
  12. Prefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.

    Back to the text
  13. Decode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.

    Back to the text