Sign In

General Previewing Tips for Locally and Remotely Generated "Long" Video Generation.

0

General Previewing Tips for Locally and Remotely Generated "Long" Video Generation.

[work in progress, published by popular demand first release June 2026, svipro stuff not yet available]

Abstract: This guide is 100% white in text and content allowing anyone with the will to generate long videos learn about using lightx2v and svipro

Introduction

These guidelines are general but the details and figures are limited to the ComfyUI workspace.

If you don't care about the reasons, skip to Summary.

My general guide for the concepts state here:

https://civitai.red/articles/27997/basic-understanding-of-video-generation-with-wan22-lightx2v-and-svipro

It is a general consensus that "previewing" in lower formats is not possible by the nature of how videos are generated. Well: yes and no. With WAN and the current possible parameters unless you use controlnet or similar systems you cannot transport those magic-moment videos from low to high generation. But you can define a strong enough prompt and prompt adhesion to generate consistent videos across resolution and steps. Proof is my own gallery.

As discussed in the general WAN video generation guide and how video is generated: the basic quality-speed parameters available in WAN video generation can lead to quite different results depending on frame size, steps and lightx2v strenght. Specially when it comes to prompt adhesion.

Other parameters which affect video generation speed may be: RAM, sagge attention and FP accumulation. But those are parameters that can't be tuned to generate approximated preview.

RAM

Affects in terms of capacity of the model, number of loras XY size lenght size. In this guide all those parameters are considered fixed except for the XY since model, loras and length are part of the process.

Sage Attention

We will only discuss long video generation with Sage Attention fixed in on or off, so its not a parameter to play with. Different Sage Attention techniques result in "completely" different videos. Nothing suggest that local video generation with WAN degrades enough to disable it and the boost time is within the order of magnitude of the video generation. Sage Attention reduces quality but I think my portfolio demostrates quality is high enough to oversee its effects by the time being. [reference TBD]

FP Accumulation

As FP Accumation is not generally available, has demonstrated two key points: degrades a lot long WAN videos, video generation is not stokastic. If you want high quality videos is not an option. [reference TBD]

imagen.png

Example of parameters you should know when generating a video. dequant and patch are not part of this guide: FP_Accumulation and SageAttent

Once discussed what are considered non-parameters lets move on the actual parameters and how you can "preview" video generation.

Note: There are techniques already implemented in ComfyUI to see the results almost online by decoding the latent while the video is being generated (June 2026). These methods, though, do not let you generate faster videos just preview the current generated video. Unless you have a powerful machine and big fast SSD do not turn this setting on since generates massive files which may further slow down your process. A 6 chunk 704x1280 video may require 10-12GB of HDD space in temporal files.

Goal

We want to be accurate enough on "low res" generation to then cook videos at 704x1280 (or 16:9) pixels with 14 or 16 steps which usually means hours of generation. These "low res" videos may take 700 to 1500 seconds to generate (10 to 30 minutes).

We will keep the workflow as close as the final production we want to achieve and then change parameters to speed up the process acknowledging the consequences.

Workflow

First: reduce the size.

Reducing the size will produce worse movement detail, faster movements or distorted movements.

Limit to multiple of 32 video sizes. Since the latent is composed of those elements (actually 16) the RAM usage and computation time is proportional to the latent space not the pixel space. So adding 2 pixels in any dimension out of those multiples creates an entire row or column or latent elements. [some corrections have been made regarding this issue in some specific workflows]

Non multiple of 32 and 16 videos have half cooked latent elements which cost you to generate in time and space and are the usual cause of AI slop since the model does not fully process those elements and may change the expected behavour of motion.

Preview at 480x832 (or 16:9 equivalent). 512x832 max. Go down to 384x512 if your hardware is limited.

Be aware of the consequences.

Since we have changed the number of pixels the following parameters will make your video behave differently: CFG, lightx2v.

CFG is more exagerated when pixels are reduced.

lightx2v may lead to two paths: extreme motion or low motion. Extreme motion in fluids, cloth, limbs or small elements. Low motion in persons or big movements.

Notes

  • reduced size and high step count can help you know the final quality of the prompt adhesion.

  • with the same seeds reduced size may generate different main-motions than bigger sized videos.

Second: reduce the number of steps.

Reducing the steps will procude far worse movement details: facial expressions, unfinished or unstarted prompted actions, fluid explosions or exagerated motions.

Limit yourself to a 3 high 3 low combination or even better 3+1 high 3 low. (note 3+1 means total high = 4 --> 1 on non light and 3 on light) check this guide.

Steps is a fundamental parameter for speed and prompt adhesion, more than CFG.

CFG forces the initial motion, with non lighx2v this initial force gives fluidity and prompt adhesion and its very important.

Steps and lighx2v strenght are intimated linked: more steps less lightx2v. The slope is very steep.

<4 steps you can go as high as 1.5-1.6

5-6 means 1.1-1.2

7 steps = 1.0 strenght approximately.

>8 steps means 0.9-0.8. Yes! This is very important. Use lightx2v with values less than 1.

But there is a caveat: the loras. Loras change the motion speed or reduce it. So for a perfect workflow where each chunk has its own loras, length and CFG lightx2v should be chunk indepedent based on the number and strenght of the loras.

The typical rule of thumb is the sum of loras strenght should not exceed 1.0 to 1.5 (depends on the type of lora). If you are close to that number and are motion loras (actions, not styles) lightx2v tends to need a little more humpf 0.1-0.3.

imagen.png

Example of chunk parameters for each chunk you select loras HIGH, LOW, CLIP if required, length, a seed (may be shared or individual) and the particular strengh of lightx2v. Note that this must be separated because the same loras are used in the "non-lightx2v" steps.

Other effects

Long videos have a side effect: prompt is squezed to the end. Still looking for the reason but after 8-10 chunks prompt seems to be delayed and you'll notice your last chunk barelly follows the prompt and its still doing the last prompt. A quick solution, add a short final chunk of 45 to 65 frames where then only prompt is the second half of the last prompt starting with "the action must continue as follows".

Interface

Current ComfyUI implementations (as of July 2026 and earlier) have a "quite good" memory handling system for loras and models. It is not part of this guide discuss this issue since memory optimization in video generation is a whole topic.

· It is worth mentioning, though, you should use the same models for previewieng and for the final video, specially in the clip, how your prompt is tokenized and how loras inject its effects is a keypoint. Do not use pruned utm clip tokenizer, instead load it into CPU (select the device in the node -> detault change it to CPU), is something is cheap right now are CPUs (since PCs aren't selling because of ram) so there is a surplus. Check if swaping your CPU to the max available for your socked its expensive, probably it isn't.

· Now Models are loaded one after another. i.e. comfyUI runs one model, unloads it, then the other (I mean high and low). If can't keep them on RAM loads it from your HDD. Buy a fast M.2 pcie hard drive in terms of reading! 7000MB/s drives on reading are usual, don't worry about writing.

· If you're a 100% on lightx2v, they have their own interface which is quite faster than ComfyUI.

Lightx2v

Note: Remember that 480p and 720p are "around" sizes. Since its better to use 32 pixel multiples even if latent is made of 16pixels units, 480x832 is the resolution, and 704x1280 or 736x1280. Use the usual 4:3 16:9 or 21:9 ratios (or portrait equivalents). High / Low => 1030 High and 480p_r64 low.

1030/480p_r64

This combinations works well in both 480p and 720p. 480p its WAN2.1.

Remember that 1030 doesn't have camera prompting (in high) so 480p will be your link to camera prompting. 480p uses "the viewer" and must be moved like a person:

"The viewer walks backwards" "The viewer moves up".

Strenght 1.2/1.1 for 720p

Strenght 1.05 1.02 for 480p.

1022/1022 (fixed rank)

This combination gives worse motion and better image detail for the same steps. Camera prompting its almost impossible. If your video has low general motion (tiktoker like) these are good combinations. I generally don't like it.

Strenght 1.3/1.2 for 720p

Strenght 1.13/1.1 for 480p

1030/1022 (fixed rank)

Works as alternative to the 480p version, gives better image quality but worse prompting. In this case youre forced to use a triple sampler with some high steps lora-less. (2-3 steps with CFG 2-3)

this can lead to extreme motions in some cases.

Strenght = Completely random, from 1.0 to 1.5 depending on loras.

260412 / 260412 (both in rank64 or rank128)

This combination gives excepcional detail and good motion. If you want even more motion use r64 instead of r128. Using one or the other depends on the available RAM since R128 its 2GB more per lora. Usually 2GB is the spare required for the latent.

Camera can be prompted both by "the viewe" but my latest experiments show that camera can be used as "the camera".

This is important because will not mix "he" with "the viewer". As 1030/480 may do.

The only problem is the amount of ram required for the r128 ram size.

Strenght 1.2/1.1 for 720p. = Depending on the loras CFG 0.9/0.9 for 720p helps.

Strenght 1.05 1.02 for 480p.

Summary

Values valid for: WAN 2.2 and SVIPro --> lightx2v 1030 high, 480p_r64 low. + the Possibility to have non-light steps.

To summarize: when previewing you may be tempted to change CFG or lightx2v, its fine. These are the guidelines:

LOW STEP | LOW RES: higher lightx2v (1.5-2) lower CFG (1.5-1.8) for non lightx2v. The higher the number of non-lightx2v the higher CFG.

What you loose: prompt adhesion (some actions will not be performed because the lack of steps). Some fluids will appear in great quantities, cloth will fly, etc. Some actions will not be performed.

same video on its final form

HIGH STEP | HIGHER RES: lower lightx2v (1.1-0.9) and higher CFG (2.0-2.5) for non lightx2v. The higher the number of steps the lower the lightx2v strenght values. The higher the number of non-lightx2v steps the higher its CFG (up to 3).

What can happen: your lightx2v strenght is too high and you get the low motion effect or no motion. Reduce the lightx2v.

You can test the correct values specific of your lora combination by generating a single chunk with another prompt (more to come in this matter).

[table here]

0