Model setup recipes

Qwen3.8-27B settings by hardware: oMLX on Mac and llama.cpp on RTX

Choose your hardware to see balanced, generation-speed and long-context profiles, with checks to work through.

Suppose Qwen3.8-27B seemed fine on the first day, but a long coding question now takes a while. The model name alone is not enough to compare your results with someone else's speed figures. First match the file and context used by oMLX on Mac or llama.cpp on RTX, then identify whether loading, the first answer, or generation is slow and change one setting at a time.

Requirements and key details
  • oMLX 0.7.0 supports Lightning MTP and DFlash2 for compatible Qwen3.8-27B checkpoints. Compare the two generation-acceleration paths one at a time.
  • Treat weight quantization and KV cache precision as separate settings. For long contexts, consider Q8 KV cache and a smaller prefill batch.
  • Save a baseline with the balanced profile, change one setting at a time, and use the median of three runs to decide which changes to keep.
Jump to hardware-specific settings

Qwen3.8 27B · 4K

Qwen3.8 27B · 4,096 input tokens1 · Estimated for one user. These are the same calculated results as the speed experience, not a measured run.

Time to first token2 is the wait before the answer starts. Token generation speed describes how quickly the rest follows. For long documents, consider both.

Qwen3.8 27B · 4K
Qwen3.8 27B · 4K

Mac mini M4 32GB · Qwen3.8 27B

Quantization3: Q4_K_M · Acceleration: none · Token generation: 5.1 ~ 5.7 tok/s · First token: 16.4 to 35.5 seconds

Prompt processing: 116 ~ 252 tok/s · Memory required: 17.6GB · Available memory: 27GB

Mac mini M4 Pro 48GB · Qwen3.8 27B

Quantization: Q4_K_M · Acceleration: none · Token generation: 12.3 ~ 13.7 tok/s · First token: 7.8 to 14.3 seconds

Prompt processing: 289 ~ 530 tok/s · Memory required: 17.6GB · Available memory: 41GB

RTX 4090 24GB · Qwen3.8 27B

Quantization: Q4_K_M · Acceleration: none · Token generation: 38.3 ~ 51.9 tok/s · First token: 2.6 to 3.6 seconds

Prompt processing: 1,141 ~ 1,591 tok/s · Memory required: 17.6GB · Available memory: 22.5GB

Match the model file to your runtime

Start by selecting the computer you actually plan to use. This guide separates the Mac path using oMLX from the RTX path using llama.cpp; the desktop starting point is a compatible Q4 GGUF4 or MLX5 conversion. Even with the same model name, different file formats and apps lead the selector to provide different commands. Match your device and format before copying someone else's settings.

The roughly 17 GB figure is an approximate guide when looking for a Q4 file, not the total memory needed at runtime6. After downloading it, check its exact size and whether enough room remains for the operating system, app, and cache during execution. A model file being present on disk is not the same as the model being loaded into memory and ready to answer.

The path that passes the official BF167 model ID to a server is separate as well. A simple calculation for the 27B weights alone gives roughly 54 GB, with runtime workspace needed on top. The vLLM example below therefore assumes an NVIDIA environment capable of running the official server weights; it cannot be reused as a Q4 command for a 24 GB GPU8. Even under the same model name, the required environment changes with the distribution file and precision.

Establish a text-question baseline first; it will be easier to compare later when you add images or MTP9. If you change the Q4 file, vision input, and MTP all at once, it is hard to tell which of the three caused a failure. In this example, first reach a state where a text answer completes, then add only the extra features supported by the selected runtime.

Match the model file to your runtime
Match the model file to your runtime

Keep a working baseline

Apply the command produced by the selector for Mac oMLX or RTX llama.cpp, and send one request at a time. First get a short coding question to complete at the indicated context. The goal is not a high score but to establish that the basic path works from model loading through the answer. Starting with a much larger context changes cache use and makes it harder to establish a baseline.

Once a normal answer appears, record that configuration. After a preparation run, note prefill10, time to first token, decode11, and peak memory across three runs so you are less likely to judge from one unusually fast result. Record a repeat of the same input separately from the first request. An input cache may remain for the second run, so do not interpret the difference as a fixed speed difference between models or devices.

For example, if you include a function that fails on empty input along with its caller, note which files were included in the request and which conditions the answer identified. Before producing performance figures, this record helps confirm that both settings were given the same task. If someone else used a shorter question or files that were already open, you cannot directly apply their tok/s to the completion time for your work.

The vLLM command below is a different configuration from the desktop selector. It is an example for an NVIDIA server that can hold the official BF16 weights, so do not mix it with the Q4 commands for Mac or a 24 GB RTX GPU. A client sending requests to 127.0.0.1 does not prove that the server is inaccessible from external networks. Check the actual binding and firewall; prepare authentication and access controls before allowing other devices to connect.

Official vLLM server setup (BF16 weights)
vllm serve Qwen/Qwen3.8-27B   --host 127.0.0.1   --port 8000   --max-model-len 8192   --gpu-memory-utilization 0.90
Run on 127.0.0.1 on an NVIDIA server with room for the official BF16 weights and runtime overhead.

A running server still needs a reply test

For an API12 test, use the model identifier actually shown in the model list. The port must also match the runtime you launched. This guide's commands use 8000 for oMLX and 8080 for llama.cpp; repeating the example curl against port 8000 will not help if no server is listening there. First check that the address responds, then verify that a short chat request completes.

Check whether the response arrives in streaming chunks, and if the app separates reasoning from the final answer, inspect how both stages appear. When the runtime is set to emit a long reasoning trace, the final sentence may appear late even if input processing is not the cause. Looking at the display from the first token through the final answer alongside server logs helps distinguish where the time was spent.

A name appearing in the model list confirms that the server knows an identifier; it does not verify that your app requested the intended model or that the context setting took effect. An incorrect name may cause a rejected request or route the answer elsewhere. Verify the actual model identifier and server logs in a short test response before moving to document input; this reduces confusion later between a slow answer caused by loading and one caused by targeting a different model.

Once the text question completes, send code or a document you actually plan to use. For example, include the function to change and nearby callers, then check whether the answer addresses the requested file scope. A high tok/s on a short greeting cannot stand in for the first-answer and completion times after a long code input, so use your own task as the comparison baseline.

Check the API endpoint and test a chat request
# 1. List loaded models
curl -N http://127.0.0.1:8000/v1/models

# 2. Call the OpenAI-compatible chat completions API
curl -N http://127.0.0.1:8000/v1/chat/completions   -H "Content-Type: application/json"   -d '{
    "model": "<local-model-identifier>",
    "messages": [
      {"role": "user", "content": "Summarize three advantages of local LLM serving."}
    ],
    "temperature": 0.6,
    "max_tokens": 1024,
    "stream": true
  }'
These examples use port 8000 for oMLX and 8080 for llama.cpp. Match the address to the runtime you started.

Find where the wait occurs

If the first answer is late after sending code or a document, check the input-token count, prefill chunk, and cache state. If generation is slow after the answer has started, inspect memory use, swapping, and whether some weights have moved to the CPU13. The former may be an input-processing bottleneck and the latter a generation bottleneck; increasing MTP steps to fix both can miss the cause.

The oMLX 0.7.0 release includes Lightning MTP and DFlash2 paths for Qwen3.8-27B. They are different speculative-decoding14 methods, so select one at a time in the UI rather than enabling both, and check that the checkpoint15 and draft model match the chosen path. Comparing the baseline decode with one acceleration path on the same prompt shows what actually changed.

The Qwen3.8-27B oQ4e comparison in PR #3958 generated 512 tokens from a long code_python prompt on an M3 Ultra with an 80-core GPU and 512 GB of unified memory16. At temperature 1, top_p 0.95, and top_k 20, Lightning MTP reached 74.1, 64.9, and 49.6 tok/s at 8K, 16K, and 64K context; DFlash2 reached 79.4, 84.7, and 59.9 tok/s. ANE prefill was off, and each row compares the PR branch with main. These are PR validation results, not a fresh repeated test of stable oMLX 0.7.0 under the same conditions, so repeat the comparison on your Mac with the same model, input, and context before applying them.

PR #3958 does not list a memory-guard tier or KV-cache17 precision in its results table. In oMLX 0.7.0, balanced leaves about 8% of memory (3–8 GB) free for other apps, while aggressive leaves about 2% (1.5–4 GB). Start with balanced for everyday use; consider aggressive only after checking stability with other apps closed. Record KV-cache precision and MTP as separate conditions so you can tell which setting changed the result.

Also validate the code result, not just speed. For example, send the same function-change request with the default and MTP settings, then check whether the function was renamed as requested, the empty-input branch remains, and the existing return format is preserved. If a test fails, record the difference to see whether the faster answer actually completed the task. This checks the code requirements; it is not the same test as whether candidate-token verification works.

If memory runs short only after increasing the context, first reduce the request count and input length to find a range that answers reliably. Check whether the limit recurs with documents you use every day, and consider a larger-memory configuration only if needed. Rather than declaring the device inadequate after one failed long prompt, compare the text baseline with the actual workload to find what changed.

Keep settings you can use every day, not just a fast result

If you plan to use the model for both code changes and document summaries, keep one representative prompt for each. A prefill batch or cache setting that helps one task may change memory pressure or waiting time at another input length. Record more than answer speed: compare whether the code runs and whether important document conditions remain in the summary against a reference answer.

In the final record, include the model filename, app version, context length18, request count, and acceleration options. If an update changes execution, you can return to this state and identify which setting changed. First establish a setup that completes your usual task and returns the necessary information accurately. Once that baseline is stable, adjust one item and compare whether waiting actually decreased.

For example, if your usual questions complete reliably with the Q4 file but only long inputs fail, first check how the selector configured that context and prefill before buying new hardware. Conversely, if the model will not load even with a short input, inspect the file/runtime combination and available memory before cache precision or MTP steps. Locating the problem makes it easier to choose what to change.

Keep settings you can use every day, not just a fast result
Keep settings you can use every day, not just a fast result

Running an NVFP4 checkpoint

Requires a compatible NVFP419 kernel. Spark measurements and RTX estimates are kept separate.

vLLM / SGLang · Inferact/Qwen3.8-27B-NVFP4

Terminology notes

  1. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text
  2. Time to First Token — The time from sending a request until its first output token arrives. It is separate from the generation rate of later tokens.

    Back to the text
  3. Quantization — Representing model values with fewer bits. Memory use, accuracy, or execution speed may change; the effects depend on the format and implementation.

    Back to the text
  4. GGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.

    Back to the text
  5. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  6. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  7. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  8. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  9. MTP — A training method or model component for predicting multiple future tokens. Whether it improves generation speed depends on implementation and runtime conditions.

    Back to the text
  10. Prefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.

    Back to the text
  11. Decode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.

    Back to the text
  12. API — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.

    Back to the text
  13. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  14. Speculative decoding — A generation method that proposes output candidates for the main model to verify. Candidates may come from a separate draft model, an MTP head, or lookup of repeated context; speed effects depend on implementation and conditions.

    Back to the text
  15. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  16. Unified memory — An architecture where the CPU and GPU share one physical memory pool. It does not increase total memory capacity; available capacity depends on the system.

    Back to the text
  17. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  18. Context window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.

    Back to the text
  19. NVFP4 — A 4-bit floating-point data format defined by NVIDIA. Support depends on GPU generation, model, and software implementation.

    Back to the text

Serving recipe

Settings for your hardware

For one user

Choose your hardware to see the launch command and settings you can adjust.

Available memory: 56GB / Bandwidth: 307GB/s

Profile

Apple MacBook Pro M5 Pro (64GB)

Memory fits

Recommended runtime: oMLX (4-bit MLX (a different format from GGUF Q4_K_M)) · Documented runtime settings

For your first run or reliable everyday use

Profile context length

16,384 tokens

Based on runtime documentationMBP M5 Pro 64GB · Q4 · 1 request

A starting setup based on the runtime documentation. Follow the steps below and compare speed as you make changes.

Estimated decode speed

13.9 ~ 15.4 tok/s

Estimate for one user

Estimated TTFT

36.0 ~ 41.9 s

Prefill: 392 ~ 456 tok/s

Estimated memory use

About 21.21GB / 56GB

Remaining margin: approx. 34.8GB

Server address

http://127.0.0.1:8000

Local only · 127.0.0.1

Estimates use the calculator’s recommended quantization and acceleration with 16,384tokens. They are not measurements of the configuration above.

Settings in this profile

ValueWhy this setting
Model weights4-bit MLX (a different format from GGUF Q4_K_M)Use an MLX model directory. This is a different file format from Q4_K_M GGUF.
Check input length16,384Check the input limit in model settings before sending a request. This launch command does not set the context length.
KV cacheCheck model settingsThis launch command does not set the KV cache precision.
Prefill batch sizeCheck runtime settingsDo not use llama.cpp's batch-size and ubatch-size values as substitutes for MLX settings.
Concurrent requests1Measure the speed of one request, separately from the combined throughput of multiple requests.
Memory headroom2 GB or moreAfter loading the model, check actual memory use and whether the system is swapping.

MBP M5 Pro 64GB Run command (oMLX)

Apple Silicon · MLX Model · Memory Guard
MODEL_DIR="$HOME/.omlx/models"
omlx serve \
  --model-dir "$MODEL_DIR" \
  --host 127.0.0.1 \
  --port 8000 \
  --max-concurrent-requests 1 \
  --memory-guard balanced
Record a baseline with this setup, then follow the tuning steps below, changing one setting at a time.

Checks after startup

  1. 01Check the startup log for completed model loading and the actual context length.
  2. 02After warming up, send one request at a time. Use three different inputs of the same length to record fresh-input processing time.
  3. 03Record repeated-input results separately as cache-reuse runs. Do not combine them with fresh-input results.
  4. 04Record prefill tok/s, decode tok/s and peak memory together, then change one setting at a time.

Trade-offs

  • Conservative context and concurrency settings may produce less than the hardware's maximum throughput.

Adjust one item at a time

Speed tuning, step by step

If several values are changed at once, it is difficult to find the cause.

  1. STEP 1

    Save a baseline

    Record first-input and cache-reuse runs separately, three times each. Compare prefill, decode and peak memory together.

    When to stop: Do not move on if there are errors or less than 2GB of memory headroom.

  2. STEP 2

    Tune prefill chunk size

    Increase in steps of 1,024 → 2,048 → 4,096, comparing TTFT and peak memory on long inputs.

    When to stop: Revert to the previous value if TTFT does not improve or peak memory rises sharply.

  3. STEP 3

    KV Cache Tuning

    Only when you need more context, compare Q8 cache against the F16/BF16 baseline using the same question.

    When to stop: Keep the default precision if the output changes or you already have enough context.

  4. STEP 4

    MTP/speculative decoding

    Enable it only for supported models. Test code and prose separately, measuring acceptance rates and actual decode tok/s.

    When to stop: Turn it off if the median of three runs does not improve on the baseline.

  5. STEP 5

    Context expansion

    Double context length at each step until you reach what you need. Check retrieval from the middle of the input and whether swapping occurs.

    When to stop: Reduce by one step if retrieval accuracy drops or swapping or memory compression starts.

Submit a measurement from my hardware

Import a JSON file with at least three runs under the same conditions. The file stays in this browser until you submit it.

Include only hardware, model, runtime settings and measurements. Do not include prompts, responses, raw logs or file paths.

If you do not have a measurement file, use the tool with your running local server. It requires Node.js 20 or later and does not upload results automatically.

These are reviewed community submissions, not measurements made by this site.

Report a setup issue

Tell us where this setup failed. Only the site administrator can read your report.

Setup being reported · MBP M5 Pro 64GB · Qwen3.8 27B · Balanced

Change log

These entries record changes to the site's guidance. They do not automatically check your installed engine or model version.

  1. Corrected model formats and runtime settings for Mac

    Separated the Mac MLX path from GGUF Q4_K_M and aligned the engine name, model format and command. KV-cache precision and prefill batching that the command does not specify are now marked for checking in the model or runtime settings.

    Also removed MLX speed profiles that implied a larger batch despite leaving the command unchanged.

  2. Removed another engine's options from Qwen 27B Spark settings

    Prevented llama.cpp-specific speculative-decoding flags from appearing in single- and dual-Spark SGLang settings. The DFlash2 and DSpark configurations retain their respective models and runtime paths.

  3. Separated uncached-input and cache-reuse records

    Updated the checks to record first-time inputs separately from repeated inputs. Cache effects are not folded into uncached prefill throughput or hardware differences.

  4. Added saving and reopening hardware profiles

    Runnable profiles can now be saved with their hardware selection in this browser, up to five entries, and reopened from the guide directory. A notice prompts you to recheck before running if the published settings have changed or the profile is no longer available.

  5. Added links to runtime explanations

    Engine names in the settings now link to the corresponding runtime guide. Paths without an identified engine are not linked to a guessed program.