Model setup recipes

Run a llama.cpp MTP Server on DGX Spark: Qwen3.6 Agent Requests

Go beyond starting the model: verify that prior thinking carries into the next request.

To serve a local agent with Qwen3.6-35B-A3B MTP1 on DGX Spark, build llama.cpp for the GB10 CUDA2 architecture and choose a Q4_K_XL GGUF3 that includes MTP. `preserve_thinking` keeps Qwen's interleaved thinking in later turns. Check, in order, that the model responds, the MTP draft initializes, and the multi-turn API4 works.

Requirements and key details
  • NVIDIA's playbook builds llama.cpp for Spark with `CMAKE_CUDA_ARCHITECTURES=121a-real`.
  • MTP uses a compatible `unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL` GGUF with `draft-mtp`.
  • `preserve_thinking` tells Qwen to retain earlier thinking blocks in multi-turn history.

Start with a Spark-verified model file, not just the 35B label

Suppose a coding agent proposes a change in a repository, then receives test logs and revises it. Entering a chat URL may look sufficient, but the server model must include an MTP head and the request format must carry thinking history forward. This recipe serves Qwen3.6-35B-A3B from one DGX Spark through an OpenAI-compatible endpoint.

NVIDIA's llama.cpp playbook specifies DGX Spark, DGX OS, 128 GB unified memory5, and CUDA architecture `121a-real`. The example model download is about 35 GB, and the llama.cpp build also needs disk space, so check available storage. The playbook lists about 30 GB of available memory for model weights and KV cache6, but actual needs vary with other apps and context length7.

Use `unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL`. The MTP label does not mean other quantized repositories or ordinary Q4 files include MTP. Because this is a large file, confirm Hugging Face access, network reliability, and resumable downloads.

A small AI computer beside a code-editing screen on a desk
For a server used on the same Spark, bind to a local address and verify the connection first.

Prepare a CUDA build specifically for Spark

Open a Linux terminal on DGX Spark and confirm Git, CMake, and the CUDA Toolkit are available. The commands below follow NVIDIA's documented dependency-install and source-build sequence. `-DGGML_CUDA=ON` builds the CUDA backend; `121a-real` specifies Spark's GB10 GPU8 architecture. Record the executable version and source revision if you rebuild from the same clone.

After the build, run `build/bin/llama-server --version` to confirm the executable was created. If CUDA is reported missing, check `nvcc --version` and your PATH. A build for another GPU architecture may complete but fail to recognize Spark's GPU, so do not copy settings from an RTX build.

Install dependencies, clone the repository, and build with CUDA
sudo apt update
sudo apt install -y git clang cmake libcurl4-openssl-dev libssl-dev
git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp
cd ~/llama.cpp
cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CURL=ON -DGGML_RPC=ON -DCMAKE_CUDA_ARCHITECTURES=121a-real
cmake --build build --config Release --target llama-server -j
./build/bin/llama-server --version
These CUDA build flags come from NVIDIA's Spark playbook. After building, verify the server version produced on Spark's DGX OS.

Start with a local-only endpoint

Only an agent app on this Spark needs to connect, so bind the server to loopback with `--host 127.0.0.1`. NVIDIA's `0.0.0.0` example allows connections from other devices; do not use it unless you intend to expose the server to your network. `-hf` downloads the Hugging Face GGUF into cache and can also load a vision projector automatically when supported by the model.

The first startup includes downloading a model tens of gigabytes in size and loading it into CUDA. The API is not ready until `server is listening` appears in the server terminal. CPU9 and GPU share the 128 GB unified memory, so subtracting model size alone cannot establish that enough memory remains. Account for context, the operating system, and other apps as well.

The MTP command includes `--spec-type draft-mtp` and `--spec-draft-n-max 3`, values from NVIDIA's compatible MTP example. Do not transfer them to similarly named checkpoints10 or other models without checking support. `preserve_thinking` retains template-generated thinking blocks in later turn history so the agent does not lose prior reasoning from the conversation.

Spark server with MTP and thinking history enabled
cd ~/llama.cpp/build
./bin/llama-server -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL --host 127.0.0.1 --port 30000 --chat-template-kwargs '{"preserve_thinking":true}' --spec-type draft-mtp --spec-draft-n-max 3 --alias qwen3.6-35b-a3b --ctx-size 8192 -ngl 99
Use the recommended Q4_K_XL GGUF for DGX Spark. The first run downloads and loads the model, so wait for the server-ready log.
A small draft-candidate block and a larger verification block beside an AI computer
MTP candidates can be accepted together only after the target model verifies them.

Check health, then send a short API request

Watch the server log for the model load and `speculative decoding11 context initialized`, then check the health endpoint from another Spark terminal. This log means the MTP context initialized; it is not a measurement that the model will always run faster for your requests. If startup fails, first check the model download, CUDA build, and memory errors.

Next, send a one-sentence request to the OpenAI-compatible `/v1/chat/completions` endpoint using the model ID served by the server. A returned answer confirms the endpoint connection. This is a smoke test, not a speed measurement. The coding agent should continue with a revision request using the same conversation ID and history. If the app removes or rewrites thinking blocks, `preserve_thinking` cannot retain content the app has discarded.

For a multi-turn check, send the function and requirements in the first turn, then send the test log in the same conversation and ask to fix only the cause of failure. Confirm that the answer connects the earlier proposed change with the test output. Whether an API wrapper resends message history and whether a raw response exposes thinking content can depend on app settings and the model's response format.

Check server readiness and the first JSON response
curl --fail-with-body http://127.0.0.1:30000/health
curl --fail-with-body http://127.0.0.1:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"qwen3.6-35b-a3b","messages":[{"role":"user","content":"Review this Python function for empty-list input and suggest a minimal fix: def first_item(items): return items[0]"}],"max_tokens":128,"temperature":0}'
After the health check succeeds, send a request with the exact model ID. Confirm that the JSON contains an assistant message under `choices`.
A local AI workspace connecting code and conversation screens
A follow-up request should include earlier messages and the test results.

Evaluate agent conversations separately from speed tests

After checking the first response, send a second request with the same conversation history to see whether thinking carries forward. A coding agent with a long context uses it for both history and code files. NVIDIA's playbook recommends at least 32K, preferably 100K or more, for agentic and coding work, but does not guarantee that these lengths fit comfortably in every Spark setup. Set the context based on the repository input and memory used by other apps.

For a speed comparison, keep the model revision, input, and generation limit fixed, warm up the server, and then measure. Check initialization logs and server timings or speculative statistics to confirm MTP is active. To compare response speed, remove only the MTP options from the baseline command and send the same prompt several times. Exclude the initial download and model-loading time from generation results. NVIDIA's playbook gives no direct comparative tokens12-per-second figure for this combination, so do not estimate or invent one.

Check the model ID, backend, and memory first

If you see `curl: (7) Failed to connect`, this is not a model-speed problem. Check that the server is listening and that client and server use `127.0.0.1` and port `30000` on the same Spark. If startup exits, inspect standard error for an incorrect Hugging Face model name, download failure, CUDA initialization issue, or OOM message.

If the server starts but returns no answer or malformed text, first check that you selected an MTP-compatible GGUF, that the executable initialized `draft-mtp`, and that the agent uses the same model ID and chat template. Setting `preserve_thinking` to true does not make the app retain history automatically. Confirm that the second request's messages include the required earlier turns.

If memory runs short, close other apps and reduce unnecessary context first; set an explicit limit with `--ctx-size` if needed. Reducing context in a long agent session can hide part of the conversation history from the model, so this is not merely a speed setting. If you switch quantized files, recheck that the new file includes an MTP head and is supported.

Official implementation references

Installation and support details were checked against the official sources below. Record the model and runtime13 versions used when reproducing the setup.

Terminology notes

  1. MTP — A training method or model component for predicting multiple future tokens. Whether it improves generation speed depends on implementation and runtime conditions.

    Back to the text
  2. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text
  3. GGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.

    Back to the text
  4. API — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.

    Back to the text
  5. Unified memory — An architecture where the CPU and GPU share one physical memory pool. It does not increase total memory capacity; available capacity depends on the system.

    Back to the text
  6. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  7. Context window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.

    Back to the text
  8. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  9. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  10. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  11. Speculative decoding — A generation method that proposes output candidates for the main model to verify. Candidates may come from a separate draft model, an MTP head, or lookup of repeated context; speed effects depend on implementation and conditions.

    Back to the text
  12. Token — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.

    Back to the text
  13. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text