Image and video generation

Run Pruna-Qwen-Image-2.1 locally: choosing its 5- or 8-step adapter

Use 5 steps to sort through mug-scene candidates quickly; start with 8 when you need to judge the rim and handle in more detail.

When choosing a ceramic mug scene, use 5 steps to scan lighting and composition, then recheck the final candidate with 8 steps and the base model. Pruna-Qwen-Image-2.1 is not a standalone small model; it is a LoRA1 adapter applied to the existing Qwen-Image-2.1, and v0.1 quality is below the base model. This comparison uses a CUDA2 setup.

Requirements and key details
  • Five steps are faster but visibly lower in quality than eight; the project recommends 8 steps as the quality-first default.
  • Each adapter file is about 335.6 MB; that is not the full model memory footprint or GPU memory requirement.
  • The published up-to-6.3× result is tied to a specific H100 80 GB benchmark; it does not establish equal image quality or the same speedup on a personal GPU.

First choose the lighting and composition, not the final image

In short, use 5 steps to scan lighting and composition candidates, then check the rim and handle again at 8 steps. Consider a ceramic mug on a small table, comparing soft window light with a darker studio setup. If you also need to vary the background and camera angle, it is more useful to generate candidates for comparing light direction, framing and overall color before spending time polishing one image.

Pruna-Qwen-Image-2.1 is a LoRA adapter that adds weights on top of the base model. It was trained to generate images in fewer steps; this is not the same as simply lowering the base model from 40 to 8 steps. A step repeatedly updates the image representation that starts from noise, and this adapter uses either 5 or 8 steps. Its developer says the v0.1 quality is still below the base model’s, so candidate-generation speed and final-image quality need to be checked separately.

By step count alone, 8 steps use one-fifth as many denoising iterations as the 40-step base model, and 5 steps use one-eighth. This division compares iteration counts; it does not mean total generation time falls by 5× or 8×. Per-step compute, prompt encoding and image decoding3 also contribute to the job, and actual processing differs by GPU4. Treat faster candidate generation as a possible benefit, then check the time and image quality on the same computer.

A monitor on a warm studio desk shows several ceramic mug still-life candidates with different lighting and compositions.
Choose the light direction and mug placement first; inspect fine detail after selecting a candidate.

It is an adapter on the existing model, not a small checkpoint

Despite the Pruna name, you are not downloading a smaller standalone version of Qwen-Image-2.1. Load the original Qwen/Qwen-Image-2.1 checkpoint5 and apply the LoRA weights on top. The pipeline6 is the execution setup that connects the image-generation stages; the text encoder converts the prompt into a representation the model can process; and the VAE7 converts between the model’s internal representation and the final pixels. These components remain those of the base model.

Each adapter file is about 335.6 MB. That helps estimate the extra download, but it is not the total memory the GPU needs during image generation. The base weights, text encoder, VAE and intermediate working space—which varies with resolution—also use memory. So the file size alone cannot justify the conclusion that an 8 GB graphics card is enough. The model card does not establish a VRAM8 reduction or prove that the adapter will run on a smaller GPU.

The recipe assumes a CUDA GPU and non-commercial research or evaluation. Commercial use requires a separate license. First prepare enough memory and a CUDA setup to run the base model, then add the adapter. The illustrations in this guide explain the workflow; they are not Pruna-generated samples.

A workbench with books and a small expansion card beside a desktop fitted with a graphics card
A small adapter download does not remove the memory needed to run the base model.

Five and eight steps serve different parts of the same job

The roles become clearer if you compare both versions side by side with the same mug prompt and seed. The 8-step adapter is recommended as the quality-first default; the 5-step version is faster but visibly lower in quality. You can try 5 steps to scan background colors, light direction and mug placement. For a candidate where the rim’s ellipse or the handle’s shape matters, inspect it again with 8 steps or the base model. ‘Useful for quick selection’ does not mean ‘good enough for the final image.’

Each adapter was trained for its own sigma schedule9. Sigma is the noise level at a step; the schedule is the list of how that value changes from step to step. Match the 5-step list to the 5-step file and the 8-step list to the 8-step file. Load only one adapter at a time and keep the LoRA strength at 1.0. Turn dynamic shifting off and set shift=1 so the sigmas are not transformed twice. The scheduler appends the final 0 automatically, so do not add it yourself. CFG10 adjusts how strongly generation follows the prompt condition; this example sets true_cfg_scale=1.0 and uses no negative prompt.

To run it, save the Python block below as pruna_mug.py and execute python pruna_mug.py in a CUDA-enabled Python environment after installing the packages. Begin with 8 steps to confirm the setup produces an image. If you later choose 5, the adapter filename and sigma list change together. Using the same seed, 42, matches the starting noise condition across runs, but does not guarantee identical compositions from both versions.

Editing the background around a reference mug uses the same pipeline. Training covered 1K resolution and up to three reference images. Start with one image you have prepared and a 1024×1024 output, then check whether the instruction to preserve the mug’s shape is followed. More reference images or a higher resolution are separate experiments after this basic edit works.

Install the CUDA packages and pinned Diffusers version
pip install 'torch>=2.4.0' 'transformers>=5.17' accelerate peft pillow
pip install git+https://github.com/huggingface/diffusers@6256aa7666cedd47443adc8f82da9a10e110b09c
The official example assumes a CUDA GPU setup.
Generate a 1024×1024 image: pruna_mug.py
import torch
from diffusers import FlowMatchEulerDiscreteScheduler, QwenImage21Pipeline

STEPS = 8  # 8: quality-first default; 5: faster candidate selection
SIGMAS = {
    5: [1.0, 0.94, 6 / 7, 2 / 3, 0.4],
    8: [1.0, 14 / 15, 6 / 7, 10 / 13, 2 / 3, 6 / 11, 0.4, 2 / 9],
}[STEPS]

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
pipe.load_lora_weights(
    "PrunaAI/Pruna-Qwen-Image-2.1",
    weight_name=f"p_qwen_image_2.1_{STEPS}step_v0.1.safetensors",
)
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(
    pipe.scheduler.config, use_dynamic_shifting=False, shift=1.0, shift_terminal=None
)

image = pipe(
    prompt=(
        "A ceramic mug on a small table, soft window light from the left, "
        "warm neutral background, three-quarter view, realistic still life"
    ),
    width=1024, height=1024,
    generator=torch.Generator("cuda").manual_seed(42),
    num_inference_steps=STEPS, sigmas=SIGMAS,
    true_cfg_scale=1.0, use_kv_cache=True,
).images[0]
image.save(f"mug-{STEPS}-step.png")
Start with 8 steps. Changing to 5 also selects its matching file and sigmas.
Optional image-editing example using the same pipeline
from PIL import Image

source = Image.open("mug-reference.png").convert("RGB")
edited = pipe(
    prompt="Keep the ceramic mug shape; change the background to a warm gray studio.",
    image=source,
    width=1024, height=1024,
    generator=torch.Generator("cuda").manual_seed(42),
    num_inference_steps=STEPS, sigmas=SIGMAS,
    true_cfg_scale=1.0, use_kv_cache=True,
).images[0]
edited.save("mug-edited.png")
Optional: append this below the preceding code and place mug-reference.png in the same folder. It reuses pipe and STEPS from above.

Read the published 6.3× result with its hardware and workload

The model card reports up to 6.3× faster, the maximum multiplier reported for its test on one H100 80 GB GPU with BF1611 and batch size12 1. It uses the Qwen model card’s example prompt and reports the median of three generations after one warmup. Prompt encoding, denoising and decoding are included; PNG saving, model loading and warmup are excluded. This is therefore not the full wait from starting a run on a computer to saving the file. The figure cannot be used to calculate seconds on a personal device.

The chart runs base Qwen-Image-2.1 at 40 steps with KV cache13 on, while the 5- and 8-step Pruna runs have KV cache off. A KV cache stores some intermediate results for reuse. The chart leaves the LoRA unmerged and uses no CFG, compilation or CPU14 offload15. By contrast, the official quickstart sets use_kv_cache=True for Pruna inference16, so running the preceding example does not reproduce the published chart settings.

For a local comparison, keep the prompt, resolution, seed, runtime17 and save method the same, then inspect both adapters’ candidates. If you time completion, separate model loading and the first warmup from repeated generation, run the same setup several times and compare medians. A CPU timer can stop before CUDA work is complete and report too little time, so do not use an unsynchronized timer. Even a careful result describes that GPU and software setup; it is not a guarantee for other devices.

Inspect the final mug candidate for small defects

After using 5 steps to narrow the lighting and background choices, generate the selected scene again with 8 steps. Zoom in on the ellipse of the cup opening, the join between handle and body, and the highlights on the glaze. A defect can be hard to notice at a small preview but more visible at the size you will use or after cropping. Be specific in the prompt—for example, ‘soft window light from the left.’ The model card also says that detailed prompts generally give better results.

Eight steps do not mean that the adapter has reached base-model quality. The v0.1 notes say its overall quality is still below the base model, with a more noticeable drop at 5 steps. If fine detail matters in the final image, compare the same scene with the original 40-step base model as well. That may take longer, but it gives you a direct basis for deciding which version suits the job. The published 2K chart shows speed only; it does not validate 2K image quality. Training covered 1K resolution, so 1024 square is the appropriate place to start a comparison.

A round magnifying glass compares the rim and handle across three printed mug still-life images on a tabletop.
After choosing a candidate, zoom in on the rim ellipse and the handle joint.

Keep the CUDA example separate from Mac runtime paths

The official quickstart requires a CUDA GPU, and the code above uses .to("cuda") and a CUDA random-number generator. This recipe is therefore for a CUDA setup. The adapter card does not confirm Mac Metal18, MLX19 or GGUF20 integrations, but that does not prove other implementations or conversion paths impossible. If you are considering another runtime, check whether that project supports this Pruna adapter and version.

The site’s /generate page shows estimated generation times for the base Qwen-Image by hardware, not for the Pruna adapter. Fewer steps may help when evaluating candidates across several lighting and background choices. If you have only a few candidates and can wait, using the base model directly is also reasonable. Whichever route you choose, inspect the selected one or two images from the 8-step adapter and 40-step base model to decide whether their detail is adequate for the final use.

Check the weight terms separately from the evaluation workflow

The Pruna adapter derives from Qwen-Image-2.1, and its model card and included license point to the Qwen Research License. Section 1 defines non-commercial as research or evaluation, and Section 2 limits the granted use to those non-commercial purposes. It says a separate commercial license is required for commercial use. The article’s copyright and the license for the model weights are separate matters; the fact that an article is publicly available does not expand the allowed use of the weights.

A non-commercial evaluation comparing mug scenes is therefore not the same stated purpose as creating images for products to sell. This guide reports the distinction written in the license; it is not legal advice. Before downloading model files or distributing outputs, read the full license, the current official card and any separate permission that may be needed. Whether the image-generation workflow works and what purpose the result may be used for are different questions.

Terminology notes

  1. LoRA — A method that learns and applies small additional weights to adapt a base model. The adapter is not a standalone model; uses include style adaptation and few-step generation.

    Back to the text
  2. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text
  3. Decode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.

    Back to the text
  4. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  5. Checkpoint — A file containing saved model weights and related state. Versions or tasks in one model family may use different checkpoints.

    Back to the text
  6. Pipeline — A sequence of processing stages from input to output. Different models or tools may be used at each stage.

    Back to the text
  7. VAE — Short for Variational Autoencoder. It can encode input into a compact latent representation or decode that representation into an output.

    Back to the text
  8. VRAM — Memory used by a graphics card’s GPU for model weights and intermediate values. It is distinct from system RAM.

    Back to the text
  9. Sigma schedule — The sequence of noise levels used across generation steps. Values can differ even at the same step count; an adapter trained for a particular schedule should use that schedule.

    Back to the text
  10. CFG — A technique that combines conditioned and unconditioned predictions to adjust how generation follows its input. Distilled models may require it to be disabled; recommended values vary by model.

    Back to the text
  11. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  12. Batch size — The number of inputs or requests processed together in one batch. Increasing it can affect both throughput and memory requirements.

    Back to the text
  13. KV cache — Memory that stores attention keys and values from earlier tokens for reuse during later token generation. Its size depends on context length and batch size.

    Back to the text
  14. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  15. Offloading — Moving some model data from GPU memory to system RAM or storage when capacity is limited. This adds data transfer.

    Back to the text
  16. Inference — The process of using a trained model to compute an output for an input. Here, local inference means running the model on the user’s device.

    Back to the text
  17. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  18. Metal — Apple’s low-level technology for graphics and parallel GPU computation. It is not itself a model-selection or chat app.

    Back to the text
  19. MLX — A machine-learning framework developed by Apple. On Apple silicon it uses unified memory and Metal; separate Linux backends are also available. Model and feature support depends on the MLX-based tool.

    Back to the text
  20. GGUF — A file format for model data, widely used by llama.cpp-based tools. The format alone does not guarantee compatibility or speed on particular hardware.

    Back to the text