Local music generation

Create and revise ACE-Step music in a DAW: a local VST3 guide

Before editing in a DAW, separate the roles of the generation server and the VST3 plugin.

If you want to bring ACE-Step generation into a DAW, acestep.vst3 combines an independent C++17/GGML backend with a JUCE 8 plugin. The README documents Metal1 on Apple Silicon and CPU/CUDA build paths on Linux. Download four Q8_0 Turbo model files totaling about 7.7 GB, verify a short render through the local HTTP server, then connect the DAW plugin. The reported 18.3 seconds on an M2 Pro is VAE2 decoding3 for 86.8 seconds of audio—not end-to-end song generation.

Requirements and key details
  • The repository is an independent C++17/GGML implementation with CPU, CUDA, Metal, and Vulkan paths.
  • The Turbo Q8_0 setup uses four files—text encoder, LM, DiT, and VAE—about 7.7 GB total.
  • The M2 Pro’s 18.3-second result is VAE decoding for an 86.8-second audio clip, not an end-to-end generation benchmark.

Start with the path from a song idea to a DAW track

Suppose you want to make a 20-second background cue for a short video, then trim and refine it in a DAW. The acestep.vst3 repository bundles a local model server, web UI, standalone tools, and a JUCE VST3 plugin. Confirm that generation works through the server first, then route the plugin in your DAW; this keeps model errors separate from host-integration issues.

Despite the ACE-Step name, do not describe this as ACE Studio’s official app. Its README calls the project an independent C++17/GGML implementation and a native backend using ACE-Step 1.5 weights. Code and model licenses may differ, so check each license before building or redistributing.

A Mac workstation with a DAW timeline and plugin panel, MIDI keyboard, and audio interface
After confirming server output, check the track and plugin routing in the DAW.

Choose the backend for your OS and prepare the four model files

On Apple Silicon macOS, the documented CMake path enables Metal automatically. For Linux NVIDIA, the example uses `-DGGML_CUDA=ON`; for Linux Vulkan, `-DGGML_VULKAN=ON`. The build example uses `nproc`, which may not exist in the default macOS shell; if so, set a modest job count or use a command that reports available CPUs. For Windows, check release binaries and current support rather than pasting macOS/Linux commands unchanged.

The documented model set is Qwen3-Embedding 0.6B Q8_0 (~748 MB), ACE-Step 5Hz LM 4B Q8_0 (~4.2 GB), Turbo DiT4 Q8_0 (~2.4 GB), and BF165 VAE (~322 MB). Their roughly 7.7 GB total is download/storage size—not a guarantee that 7.7 GB of RAM/VRAM and scratch space will be enough at runtime6. Verify every model path and keep the first generation short with defaults.

Verify a short render through the local server first

Follow the README to clone the repository, initialize submodules, and build with the option for your OS. It documents `models.sh` for downloading the Q8_0 Turbo set. Start the server with all four model paths and open `http://localhost:8080`. A page loading is not proof generation works: enter a short description and lyrics, then confirm a WAV file is actually written.

If the first render fails, match the model name and path in the error to the actual filenames under `models/`. Check capitalization, Turbo versus SFT, and Q8_0 versus BF16 mismatches first. If loading succeeds but the process exits during generation, distinguish GPU7-memory from system-memory pressure; close other heavy apps and retry with a shorter output. Server logs help identify whether model initialization, `/lm`, or `/synth` failed.

Build the server and provide model paths
git clone https://github.com/ace-step/acestep.vst3.git
cd acestep.vst3
git submodule update --init
mkdir build && cd build
cmake ..
cmake --build . --config Release -j4
cd ..
pip install hf
./models.sh
./build/ace-server \
  --lm models/acestep-5Hz-lm-4B-Q8_0.gguf \
  --embedding models/Qwen3-Embedding-0.6B-Q8_0.gguf \
  --dit models/acestep-v15-turbo-Q8_0.gguf \
  --vae models/vae-BF16.gguf \
  --port 8080
A DAW arrangement with a short generated clip and waveform card beside an existing track
Place the generated WAV beside an existing track and listen through its beginning and end.

After the server works, connect VST3 to your DAW

Once the server has produced a WAV, build the plugin. Follow the repository instructions to configure and build under `plugins/acestep_vst3`, then copy the `.vst3` bundle to a user plugin directory scanned by your DAW. Fully restart the host and rescan plugins. Check permissions if using a system-wide folder, and do not bypass security prompts for an unverified binary simply because the plugin is missing.

When the plugin appears, load it on a new audio track and check its server connection. In this setup the VST3 UI communicates with the ace-server API8, so the server must remain running. Confirm the health indicator, enter a description, lyrics, and duration, and generate once. Drag the resulting WAV onto a DAW track and listen through the beginning and end; for the next pass, change only the prompt or duration. This connects generation and editing in one project—it does not mean every DAW host plays the model as a real-time instrument.

A studio comparing generated music takes on a Mac beside headphones, speakers, and an SSD
Changing one condition at a time makes it easier to understand differences between generated takes.

Do not read 18.3 seconds as song-generation time

The repository’s M2 Pro 16 GB table measures only VAE decoding of 86.8 seconds of 48 kHz stereo audio. It reports 18.3 seconds with chunk 1024, overlap 16, and `im2col_1d`. It does not include language-model processing, 8-step DiT generation, prompt enrichment, server startup, or DAW editing. Do not rewrite it as “an M2 Pro generates a 20-second song in 18 seconds.”

To understand your own timing, separate first model load from the second render and record lyrics, duration, seed, backend, and output length. Do not directly compare runs on different CPU9, Metal, CUDA10, or Vulkan paths. Before speed, verify that the requested output completes and produces a WAV you can edit.

Check maintenance and licenses before using it in production

The README and release page can change. Check the current release assets and known limitations on the day you install rather than relying on an older binary name or UI workflow. Review host-specific issues such as VST3 WebView input behavior, and test in a duplicated project first. Model-use terms, rights to generated output, and licenses for the plugin and dependencies are separate questions.

Do not choose a DAW or GPU based on one VAE timing. Judge the setup only after a short server render succeeds, the plugin appears reliably in your host, and the music you need can be saved and edited. Then compare backends or precisions against your actual number of revisions and project duration.

Original model and execution documentation

Consult the original sources below for installation and model usage requirements.

Terminology notes

  1. Metal — Apple’s low-level technology for graphics and parallel GPU computation. It is not itself a model-selection or chat app.

    Back to the text
  2. VAE — Short for Variational Autoencoder. It can encode input into a compact latent representation or decode that representation into an output.

    Back to the text
  3. Decode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.

    Back to the text
  4. DiT — Short for Diffusion Transformer: a model architecture that uses a transformer to refine noisy representations across diffusion steps.

    Back to the text
  5. BF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.

    Back to the text
  6. Runtime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.

    Back to the text
  7. GPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.

    Back to the text
  8. API — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.

    Back to the text
  9. CPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.

    Back to the text
  10. CUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.

    Back to the text