See the speed before choosing hardware

Image and video generation

Make and edit images locally with LLaDA-Image: choosing Base or Turbo

Few image tasks end with the first output. An illustration may need a new composition; a product photo may need one small change while everything else stays put. LLaDA-Image offers generation and reference-image editing in the same model family. Try two revisions of one image before deciding whether Base or Turbo fits your routine.

Begin with one image you actually need

The easiest mistake is collecting attractive examples without deciding what you need to produce. Is it an illustration for an article, or an existing photo whose background needs changing? LLaDA-Image has a text-to-image mode and an instruction-guided editing mode that accepts a reference image. Support for both does not mean both will satisfy every task equally. Prepare one image you can inspect and one simple instruction. For a new image, look at composition and subject; for an edit, look just as closely at what should have remained untouched.

Begin with one image you actually need
Begin with one image you actually need

Use Turbo to explore, then check Base

The official examples use 50 sampling steps for Base and four for the distilled Turbo checkpoint. Turbo is a reasonable way to explore several directions before trying Base on a promising composition. Base is not a magical final-save button, though. Repeat the same request with several seeds and inspect the details your work depends on, whether those are lettering, hands, color or edges. Turbo may already be sufficient, while Base may still need a revised prompt. The 4-to-50 step ratio cannot be treated as a hardware speed multiplier: loading, text processing, editing input and saving all contribute to a real session.

For editing, inspect what stays the same

In editing mode, pair the reference image with an instruction that names both the change and what must survive: for example, soften the background but keep the cup's logo. Put the result beside the original. Look for altered object shapes, letters, fingers and expressions that you did not request. On such work, the number of corrections may matter more than the time for one pass. The official pipeline accepts an image and output dimensions for editing, with dimensions divisible by 32. Start from its example; if a run fails on size or input, check those conditions before assuming the hardware is inadequate.

For editing, inspect what stays the same
For editing, inspect what stays the same

Record the path and version you ran

The official repository provides Diffusers-based inference and Base and Turbo checkpoints. Its example environment uses Python 3.11 and CUDA-based PyTorch, with text and editing modes selected explicitly. There are also separate ComfyUI workflows, but check their required files and software versions before downloading a single checkpoint from a screenshot. A working small-scale trial is worth recording: checkpoint name, precision, mode, dimensions and step count. That note makes the next failure much easier to diagnose.

A checkpoint size is not a memory requirement

Neither 6B parameters nor the existence of an FP8 checkpoint proves that a particular GPU will handle your work comfortably. The text encoder, image backbone, intermediate data, resolution and runtime placement all consume memory. Try Base and Turbo, and supported lighter precision variants where appropriate, without assuming identical output quality. Before buying, complete both generation and editing at the size you will use, with your everyday applications open. If your present machine completes a single run, measure repeatability and the actual wait before deciding it needs replacement.

A checkpoint size is not a memory requirement
A checkpoint size is not a memory requirement

Measure time to a usable image

For real work, the relevant clock stops when you have an image you can use, not when the first file appears. Run the same reference and instruction several times, noting initial load, generation and any corrective edit separately. Frequent composition searches may favor Turbo's quicker iterations. If you mostly refine one image, preservation and edit success may dominate. The site's image-generation experience simulates a hardware wait; it is not a measured LLaDA-Image result or a quality test. Make the final call on a workflow you have saved and can rerun.