Image and video generation
Animate a photo locally with Wan-Animate-2: what to check first
It is easy to imagine a family portrait waving or a character drawing dancing. Then you discover that the photo is only half the input. Wan-Animate-2 takes a reference image and a video that supplies the motion. Before choosing a machine, decide what motion footage you can provide and what size of result you actually need.
One input supplies appearance, the other motion
The reference photo should show the face, clothing and background you want to preserve. The driving video supplies gestures and expressions. Wan-Animate-2 directly consumes that video rather than requiring a separate pose-extraction product as the input. The official example also supplies a factual description of the character's appearance and background. Do not expect a single straight-on portrait to remain perfect through an extreme dance and every camera angle. Start with a short motion segment that resembles what the photo can support, then inspect the face and hands.

Choose the driving clip before the hardware
Write down the action you want: a wave, speaking expression or full-body dance. Then inspect the driving clip's subject scale, camera angle and range of motion. Frequent cuts and occlusion make the result harder to judge. Trim a short section without scene changes before attempting a long sequence. If the output looks wrong, compare the photo and video viewpoints and check whether the prompt consistently describes the person and background before blaming compute. A better GPU cannot replace a suitable motion reference.
Keep the repository path and ComfyUI workflow distinct
The official Wan-Animate-2 repository supplies inference scripts, a checkpoint download and a command-line example with paths for the reference image and driving video. Its download example names Wan-AI/Wan2.2-Animate-2-14B. Separately, ComfyUI publishes a Wan Animate 2 workflow template with nodes for the image, motion video, text and saved output. They are not interchangeable installation instructions. Match your ComfyUI version and required nodes and files, and do not transfer the repository's distributed settings into ComfyUI by assumption. Pick one route and complete its provided example first.

A 14B label does not size the GPU
Video spans a sequence of frames. Weights alone do not determine memory: the reference photo, driving frames, intermediate data and output size also matter. The official repository says its defaults are tuned for 720p on eight A800 GPUs and reports a 480p test on two A800s. That is a starting point, not evidence that consumer cards with a similar aggregate memory capacity will run at the same speed. A ComfyUI template may offer different precision and cache settings, but their effects on fit, quality and repeatability need separate checking. Seek a completed run on the same runtime and setup before purchasing.
Finish and save a short clip first
Read the inputs, generate the frames, save a real video file and play it back before calling the first test complete. An on-screen preview followed by a failed save is not a finished workflow. The published ComfyUI template chains motion-transfer blocks for longer footage, using an 81-frame segment in its example. Completing one block therefore does not prove that a long video will continue automatically. Check how the face, clothes and motion carry across segment boundaries. A saved short clip gives you a stable baseline for changing resolution, length or extra blocks.

Choose hardware around the finished sequence
A finished video includes choosing material, rerunning failed sections and watching the saved file for awkward frames. Decide whether your routine is a few seconds of character motion or a long narrated piece before asking what hardware it needs. The site's video-generation experience gives a sense of estimated waits across devices; those figures are not measured Wan-Animate-2 timings. Generic text-to-video and image-to-video workloads differ from transferring motion from a separate driving clip. First establish a runtime and one finished short example. Then distinguish a memory limit from excessive waiting or inconsistent output before comparing machines.