Sign In

Testing Wan2 S2V on RTX4090: Render Times, Quality, and MultiTalk Comparison

1

Testing Wan2 S2V on RTX4090: Render Times, Quality, and MultiTalk Comparison

I’ve been testing the new Wan2 S2V workflow on an RTX4090 (14B_fp8). Here’s a breakdown of the results, render times, and limits I’ve found, plus a direct comparison with MultiTalk. Spoiler: S2V shows promise with audio reactivity, but still struggles with FPS and long-term consistency.

Tests #Wan2 #S2V – 13s in 4/5 (512x672 – 208f @16fps)
Model: 14B_fp8 – RTX4090 24GB VRAM + 64GB RAM – SageAttention – Same Seed

Prompt: a man is passionately singing and playing piano of an emotional song, lens flare

The workflow provided by ComfyUI is a real mess: it has to be significantly adjusted each time depending on video length and content.

Capture d'écran 2025-08-30 130404.png


It’s also difficult to automate some calculations (77-frame batches): since audio almost never matches perfectly, the last batch rarely hits 77 frames. But adjusting just that last batch adds no real value.
I still added a few Sage / Torch optimizations.

Test results:

  1. Lora / 4 steps / Cfg 1 → 120s (5.91s/it)
    LipSync OK, but little expression and very few movements.

  2. Lora Strength x1.5 / 4 steps / Cfg 1 → 87s (5.61s/it)
    Lipsync OK. More movements, but way too frantic — almost Parkinson-like after a few seconds. Image quickly gets “burned” (contours).

  3. Lora Strength x0.5 / 8 steps / Cfg 2.5 → 414s (16.78s/it)
    Good compromise: solid lipsync, decent movements, though a bit static at the piano.

  4. No Lora / 20 steps / Cfg 6 → 15 minutes
    Good lipsync, good expressions, but not much difference from version 3 for a much longer render time. (Funny detail: it rains inside his mouth).

👉 Key notes:

  • Using a Lora reduces the required steps and speeds up rendering.

  • But it comes at the cost of movement, which becomes more static.

  • S2V is interesting because it reacts more to the input audio (soft or loud parts), with characters adapting their gestures slightly to the sound.

👉 The best result is still version 3, with a refined prompt:
A man passionately sings and plays a poignant song on the piano. His hands move from left to right across the piano keys. lens flare
→ Solid facial expressions, accurate lipsync, good balance between quality and render time.

Weak points:

On longer exports (55s / 874 frames), consistency collapses: lipsync disappears and the image degrades fast.

  • S2V: 11 minutes (874 frames @16fps) with Lora, 4 steps, Cfg 1 → impossible to render correctly with other parameters (including version 3 above).

  • MultiTalk: 12 minutes (1365 frames @25fps) for a stable result, precise lipsync and consistent image.

Compared to #MultiTalk / InfiniteTalk, the gap is clear:

  • MultiTalk runs at 25 FPS vs S2V locked at 16 FPS.

  • More stability and precise lipsync (S2V loses 10 frames every second).

  • Much more consistent, especially with singing.

  • Downside: MultiTalk sometimes produces overly frantic and exaggerated movements. Still, it remains clearly ahead of S2V at this stage.

Conclusion: S2V is promising, especially in how it reacts to audio, but it’s still too limited by FPS and lacks long-term consistency. MultiTalk remains the leader, even if it needs some smoothing on extreme movements.

1