Local music generation
Create and revise ACE-Step music in a DAW: a local VST3 guide
Before editing in a DAW, separate the roles of the generation server and the VST3 plugin.
If you want to bring ACE-Step generation into a DAW, acestep.vst3 combines an independent C++17/GGML backend with a JUCE 8 plugin. The README documents Metal1 on Apple Silicon and CPU/CUDA build paths on Linux. Download four Q8_0 Turbo model files totaling about 7.7 GB, verify a short render through the local HTTP server, then connect the DAW plugin. The reported 18.3 seconds on an M2 Pro is VAE2 decoding3 for 86.8 seconds of audio—not end-to-end song generation.
Start with the path from a song idea to a DAW track
Suppose you want to make a 20-second background cue for a short video, then trim and refine it in a DAW. The acestep.vst3 repository bundles a local model server, web UI, standalone tools, and a JUCE VST3 plugin. Confirm that generation works through the server first, then route the plugin in your DAW; this keeps model errors separate from host-integration issues.
Despite the ACE-Step name, do not describe this as ACE Studio’s official app. Its README calls the project an independent C++17/GGML implementation and a native backend using ACE-Step 1.5 weights. Code and model licenses may differ, so check each license before building or redistributing.

Choose the backend for your OS and prepare the four model files
On Apple Silicon macOS, the documented CMake path enables Metal automatically. For Linux NVIDIA, the example uses `-DGGML_CUDA=ON`; for Linux Vulkan, `-DGGML_VULKAN=ON`. The build example uses `nproc`, which may not exist in the default macOS shell; if so, set a modest job count or use a command that reports available CPUs. For Windows, check release binaries and current support rather than pasting macOS/Linux commands unchanged.
The documented model set is Qwen3-Embedding 0.6B Q8_0 (~748 MB), ACE-Step 5Hz LM 4B Q8_0 (~4.2 GB), Turbo DiT4 Q8_0 (~2.4 GB), and BF165 VAE (~322 MB). Their roughly 7.7 GB total is download/storage size—not a guarantee that 7.7 GB of RAM/VRAM and scratch space will be enough at runtime6. Verify every model path and keep the first generation short with defaults.
Verify a short render through the local server first
Follow the README to clone the repository, initialize submodules, and build with the option for your OS. It documents `models.sh` for downloading the Q8_0 Turbo set. Start the server with all four model paths and open `http://localhost:8080`. A page loading is not proof generation works: enter a short description and lyrics, then confirm a WAV file is actually written.
If the first render fails, match the model name and path in the error to the actual filenames under `models/`. Check capitalization, Turbo versus SFT, and Q8_0 versus BF16 mismatches first. If loading succeeds but the process exits during generation, distinguish GPU7-memory from system-memory pressure; close other heavy apps and retry with a shorter output. Server logs help identify whether model initialization, `/lm`, or `/synth` failed.
git clone https://github.com/ace-step/acestep.vst3.git
cd acestep.vst3
git submodule update --init
mkdir build && cd build
cmake ..
cmake --build . --config Release -j4
cd ..
pip install hf
./models.sh
./build/ace-server \
--lm models/acestep-5Hz-lm-4B-Q8_0.gguf \
--embedding models/Qwen3-Embedding-0.6B-Q8_0.gguf \
--dit models/acestep-v15-turbo-Q8_0.gguf \
--vae models/vae-BF16.gguf \
--port 8080
After the server works, connect VST3 to your DAW
Once the server has produced a WAV, build the plugin. Follow the repository instructions to configure and build under `plugins/acestep_vst3`, then copy the `.vst3` bundle to a user plugin directory scanned by your DAW. Fully restart the host and rescan plugins. Check permissions if using a system-wide folder, and do not bypass security prompts for an unverified binary simply because the plugin is missing.
When the plugin appears, load it on a new audio track and check its server connection. In this setup the VST3 UI communicates with the ace-server API8, so the server must remain running. Confirm the health indicator, enter a description, lyrics, and duration, and generate once. Drag the resulting WAV onto a DAW track and listen through the beginning and end; for the next pass, change only the prompt or duration. This connects generation and editing in one project—it does not mean every DAW host plays the model as a real-time instrument.

Do not read 18.3 seconds as song-generation time
The repository’s M2 Pro 16 GB table measures only VAE decoding of 86.8 seconds of 48 kHz stereo audio. It reports 18.3 seconds with chunk 1024, overlap 16, and `im2col_1d`. It does not include language-model processing, 8-step DiT generation, prompt enrichment, server startup, or DAW editing. Do not rewrite it as “an M2 Pro generates a 20-second song in 18 seconds.”
To understand your own timing, separate first model load from the second render and record lyrics, duration, seed, backend, and output length. Do not directly compare runs on different CPU9, Metal, CUDA10, or Vulkan paths. Before speed, verify that the requested output completes and produces a WAV you can edit.
Check maintenance and licenses before using it in production
The README and release page can change. Check the current release assets and known limitations on the day you install rather than relying on an older binary name or UI workflow. Review host-specific issues such as VST3 WebView input behavior, and test in a duplicated project first. Model-use terms, rights to generated output, and licenses for the plugin and dependencies are separate questions.
Do not choose a DAW or GPU based on one VAE timing. Judge the setup only after a short server render succeeds, the plugin appears reliably in your host, and the music you need can be saved and edited. Then compare backends or precisions against your actual number of revisions and project duration.
Original model and execution documentation
Consult the original sources below for installation and model usage requirements.
Terminology notes
Metal — Apple’s low-level technology for graphics and parallel GPU computation. It is not itself a model-selection or chat app.
Back to the textVAE — Short for Variational Autoencoder. It can encode input into a compact latent representation or decode that representation into an output.
Back to the textDecode — For an LLM, this is the stage that generates output tokens after input processing. For a VAE or audio codec, decoding can mean reconstructing the original form from a compressed representation or encoded data.
Back to the textDiT — Short for Diffusion Transformer: a model architecture that uses a transformer to refine noisy representations across diffusion steps.
Back to the textBF16 — A 16-bit floating-point format for storing and computing model values. Support depends on the hardware and runtime.
Back to the textRuntime — The software environment that provides facilities needed while a program runs. In local AI it can also refer to a model execution engine; a GPU runtime library and a complete serving app are different components.
Back to the textGPU — A processor designed to handle many calculations in parallel. It performs model computations during AI inference.
Back to the textAPI — A defined interface that lets other code call a program’s functions. The term API alone does not imply sending data to an external server.
Back to the textCPU — The central processor that executes general-purpose program instructions. AI workloads may divide work between it and other processors such as GPUs.
Back to the textCUDA — A software platform for general-purpose computing on NVIDIA GPUs. Programs built for CUDA are not guaranteed to run unchanged on other GPUs.
Back to the text