This ComfyUI tutorial workflow turns a single still image into a short video with synchronized, auto-generated audio using the MiniMax H3 family via the FastVideo FastH3 8-Step V2 checkpoint. The pipeline is simple and fast: LoadImage reads your source, ImageScaleToTotalPixels (assisted by GetImageSize) rescales it to a GPU-friendly pixel budget while preserving aspect ratio, and the MiniMaxH3ImageToVideo node (id: 4c314f31-ecda-4b08-ae98-faaba1bf613f) renders both video frames and a native audio track in one pass. SaveVideo then writes an MP4 output that includes the model’s audio when available.
FastH3 8-Step V2 is a DMD2-distilled variant of MiniMax H3 that cuts inference to just 8 sampling steps, delivering much faster turnaround than the base model while maintaining audio/video synchronization. The workflow uses the minimax_h3_video_vae_fp16.safetensors VAE and the FastVideo/FastVideo-FastH3-8-Step-V2 checkpoint, giving you a practical, single-image-to-video path that’s ideal for animating portraits, product shots, and quick social clips without extra audio post-processing.


