Model setup recipes
LFM2.5-VL-3B DSpark: more than 2× faster visual responses on Mac
DSpark does not reduce the time to read the photo, but accelerates the time to write the answer after reading.
You don't necessarily need a huge vision model to do things like extract dates and amounts from receipt photos or short descriptions of product photos locally. LFM2.5-VL-3B is geared towards fast processing of these single requests with approximately 3B size and 32K contexts. Using a separate DSpark draft model together speeds up answer generation, but image encoding and prefilling are not accelerated by the same factor.
First, divide the tasks into reading the picture and writing the answer.
A request in a vision language model can be broken down into steps that turn a picture into features, a prefill1 that reads the input context, and a decode2 that concatenates the answer token3. DSpark reduces decode iterations by having a small draft model propose the next batch of tokens and have it verified by this model. The first part itself, which encodes the camera picture, is not omitted.
Therefore, receipt OCR, which requires only one short answer line, may not see the overall completion time reduced by the same factor even with faster decode. Conversely, when generating tens to hundreds of tokens, such as image descriptions or translations, acceleration is easily felt. The site speed screen applies a token generation multiplier, and the overall task judgment is based on a combination of image size and answer length.

Official figures are bundled with equipment and engine
In six vision challenges released by Liquid AI, M5 Max's MLX4-VLM decode was 2.30 to 3.13 times faster and overall completion time was 1.56 to 2.62 times faster. M3 Ultra's llama.cpp decode was 1.57 to 2.14 times faster, and overall 1.30 to 1.77 times faster. Even if the model is the same, the magnification will vary if the engine, hardware, and output length are different.
The model card introduces the base LFM2.5-VL-3B at 228 tok/s on the Apple M5 Max and 116 tok/s on the AMD Ryzen AI Max+ 395. This value is a reference point for a specific published path and is not a guaranteed speed for all quantizations and images. The site will only use the DSpark range if the same equipment has been identified and will not arbitrarily move it to other equipment.

For Mac, choose between MLX-VLM and llama.cpp according to your needs.
The simplest formal path in M5 Max is to specify both the target model and the DSpark draft model on the MLX-VLM server. If you need app integration, check for an OpenAI compatible endpoint and first save a basic run with one fixed image and fixed output length. After you turn on acceleration, make sure your answers are the same and you have a record of draft approvals.
When using the llama.cpp path like M3 Ultra, the target F16 GGUF5, vision projector and DSpark draft files must match respectively. You can start small with context 8192 and use the formal example with a draft maximum length of 8 as a guideline. Because even a single file change breaks the comparison, we keep the command and checkpoint6 hashes together.
mlx_vlm.server --model LiquidAI/LFM2.5-VL-3B --draft-model LiquidAI/LFM2.5-VL-3B-DSparkOn NVIDIA server, check DSpark path for SGLang
The SGLang formal example specifies the DSpark algorithm and draft model, FlashInfer draft attention7, and block size of 9. We start with one request and a short context to make sure the server uploads the model and receives the image. It then records the number of draft proposals, number of approvals, and completion time compared to the same request with acceleration turned off.
A draft model is not a replacement model that will output a final answer that is different from the original model. Since this model verifies the proposal every time, the results of the same target model can be maintained under greedy generation conditions. However, if the runtime8 version or sampling settings are different, the comparison conditions will be different, so you must use the exact same request and creation options.
python -m sglang.launch_server \
--model-path LiquidAI/LFM2.5-VL-3B \
--speculative-algorithm DSPARK \
--speculative-draft-model-path LiquidAI/LFM2.5-VL-3B-DSpark \
--speculative-draft-attention-backend flashinfer \
--speculative-dspark-block-size 9 \
--disable-radix-cache --mem-fraction-static 0.8 --host 127.0.0.1 --port 30000
In my work, I only need to time three things:
We measure the base run and the DSpark run three times each, with the same pictures and the same questions. We leave aside the overall completion time, from image delivery to the first token, and from the first token to the last token. If photo preprocessing time is long, reducing image size may have a greater impact on the overall experience than accelerating decode.
If draft acceptance rates are low or responses are very short, the costs of an additional draft model may offset the gains. Pick three actual tasks to repeat, such as OCR, captioning, or translation, and compare the medians. The overall latency reduced for my workload, not the public multiplier, is the criterion for maintaining acceleration.
Terminology notes
Prefill — The stage where an LLM reads the input prompt and computes representations for its tokens. Longer prompts contain more tokens to process.
Back to the textDecode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.
Back to the textToken — A unit into which a model divides input or output for processing. One token does not equal one character or a fixed duration.
Back to the textMLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.
Back to the textGGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.
Back to the textCheckpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.
Back to the textAttention — A computation that compares positions in an input so a model can select information relevant to its current step. Details and cost depend on the architecture.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the text