Model setup recipes
Run a llama.cpp MTP Server on DGX Spark: Qwen3.6 Agent Requests
Go beyond starting the model: verify that prior thinking carries into the next request.
To serve a local agent with Qwen3.6-35B-A3B MTP1 on DGX Spark, build llama.cpp for the GB10 CUDA2 architecture and choose a Q4_K_XL GGUF3 that includes MTP. `preserve_thinking` keeps Qwen's interleaved thinking in later turns. Check, in order, that the model responds, the MTP draft initializes, and the multi-turn API4 works.
Start with a Spark-verified model file, not just the 35B label
Suppose a coding agent proposes a change in a repository, then receives test logs and revises it. Entering a chat URL may look sufficient, but the server model must include an MTP head and the request format must carry thinking history forward. This recipe serves Qwen3.6-35B-A3B from one DGX Spark through an OpenAI-compatible endpoint.
NVIDIA's llama.cpp playbook specifies DGX Spark, DGX OS, 128 GB unified memory5, and CUDA architecture `121a-real`. The example model download is about 35 GB, and the llama.cpp build also needs disk space, so check available storage. The playbook lists about 30 GB of available memory for model weights and KV cache6, but actual needs vary with other apps and context length7.
Use `unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL`. The MTP label does not mean other quantized repositories or ordinary Q4 files include MTP. Because this is a large file, confirm Hugging Face access, network reliability, and resumable downloads.

Prepare a CUDA build specifically for Spark
Open a Linux terminal on DGX Spark and confirm Git, CMake, and the CUDA Toolkit are available. The commands below follow NVIDIA's documented dependency-install and source-build sequence. `-DGGML_CUDA=ON` builds the CUDA backend; `121a-real` specifies Spark's GB10 GPU8 architecture. Record the executable version and source revision if you rebuild from the same clone.
After the build, run `build/bin/llama-server --version` to confirm the executable was created. If CUDA is reported missing, check `nvcc --version` and your PATH. A build for another GPU architecture may complete but fail to recognize Spark's GPU, so do not copy settings from an RTX build.
sudo apt update
sudo apt install -y git clang cmake libcurl4-openssl-dev libssl-dev
git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp
cd ~/llama.cpp
cmake -B build -DGGML_NATIVE=ON -DGGML_CUDA=ON -DGGML_CURL=ON -DGGML_RPC=ON -DCMAKE_CUDA_ARCHITECTURES=121a-real
cmake --build build --config Release --target llama-server -j
./build/bin/llama-server --versionStart with a local-only endpoint
Only an agent app on this Spark needs to connect, so bind the server to loopback with `--host 127.0.0.1`. NVIDIA's `0.0.0.0` example allows connections from other devices; do not use it unless you intend to expose the server to your network. `-hf` downloads the Hugging Face GGUF into cache and can also load a vision projector automatically when supported by the model.
The first startup includes downloading a model tens of gigabytes in size and loading it into CUDA. The API is not ready until `server is listening` appears in the server terminal. CPU9 and GPU share the 128 GB unified memory, so subtracting model size alone cannot establish that enough memory remains. Account for context, the operating system, and other apps as well.
The MTP command includes `--spec-type draft-mtp` and `--spec-draft-n-max 3`, values from NVIDIA's compatible MTP example. Do not transfer them to similarly named checkpoints10 or other models without checking support. `preserve_thinking` retains template-generated thinking blocks in later turn history so the agent does not lose prior reasoning from the conversation.
cd ~/llama.cpp/build
./bin/llama-server -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL --host 127.0.0.1 --port 30000 --chat-template-kwargs '{"preserve_thinking":true}' --spec-type draft-mtp --spec-draft-n-max 3 --alias qwen3.6-35b-a3b --ctx-size 8192 -ngl 99
Check health, then send a short API request
Watch the server log for the model load and `speculative decoding11 context initialized`, then check the health endpoint from another Spark terminal. This log means the MTP context initialized; it is not a measurement that the model will always run faster for your requests. If startup fails, first check the model download, CUDA build, and memory errors.
Next, send a one-sentence request to the OpenAI-compatible `/v1/chat/completions` endpoint using the model ID served by the server. A returned answer confirms the endpoint connection. This is a smoke test, not a speed measurement. The coding agent should continue with a revision request using the same conversation ID and history. If the app removes or rewrites thinking blocks, `preserve_thinking` cannot retain content the app has discarded.
For a multi-turn check, send the function and requirements in the first turn, then send the test log in the same conversation and ask to fix only the cause of failure. Confirm that the answer connects the earlier proposed change with the test output. Whether an API wrapper resends message history and whether a raw response exposes thinking content can depend on app settings and the model's response format.
curl --fail-with-body http://127.0.0.1:30000/health
curl --fail-with-body http://127.0.0.1:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"qwen3.6-35b-a3b","messages":[{"role":"user","content":"Review this Python function for empty-list input and suggest a minimal fix: def first_item(items): return items[0]"}],"max_tokens":128,"temperature":0}'
Evaluate agent conversations separately from speed tests
After checking the first response, send a second request with the same conversation history to see whether thinking carries forward. A coding agent with a long context uses it for both history and code files. NVIDIA's playbook recommends at least 32K, preferably 100K or more, for agentic and coding work, but does not guarantee that these lengths fit comfortably in every Spark setup. Set the context based on the repository input and memory used by other apps.
For a speed comparison, keep the model revision, input, and generation limit fixed, warm up the server, and then measure. Check initialization logs and server timings or speculative statistics to confirm MTP is active. To compare response speed, remove only the MTP options from the baseline command and send the same prompt several times. Exclude the initial download and model-loading time from generation results. NVIDIA's playbook gives no direct comparative tokens12-per-second figure for this combination, so do not estimate or invent one.
Check the model ID, backend, and memory first
If you see `curl: (7) Failed to connect`, this is not a model-speed problem. Check that the server is listening and that client and server use `127.0.0.1` and port `30000` on the same Spark. If startup exits, inspect standard error for an incorrect Hugging Face model name, download failure, CUDA initialization issue, or OOM message.
If the server starts but returns no answer or malformed text, first check that you selected an MTP-compatible GGUF, that the executable initialized `draft-mtp`, and that the agent uses the same model ID and chat template. Setting `preserve_thinking` to true does not make the app retain history automatically. Confirm that the second request's messages include the required earlier turns.
If memory runs short, close other apps and reduce unnecessary context first; set an explicit limit with `--ctx-size` if needed. Reducing context in a long agent session can hide part of the conversation history from the model, so this is not merely a speed setting. If you switch quantized files, recheck that the new file includes an MTP head and is supported.
Official implementation references
Installation and support details were checked against the official sources below. Record the model and runtime13 versions used when reproducing the setup.
Terminology notes
MTP — A training method or model component for predicting multiple future tokens. Whether it improves generation speed depends on implementation and runtime conditions.
Back to the textCUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.
Back to the textGGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.
Back to the textAPI — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.
Back to the textUnified memory — An architecture where the CPU and GPU share one physical memory pool. It does not increase total memory capacity; available capacity depends on the system.
Back to the textKV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.
Back to the textContext window — The token span of input and generated content a model can handle in one request. The supported limit and memory use depend on the model and runtime settings.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textCPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textSpeculative decoding — A generation method that proposes output candidates for the main model to verify. Candidates may come from a separate draft model, an MTP head, or lookup of repeated context; speed effects depend on implementation and conditions.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the text