Sign In

Create AI Influencer Ad Videos with Nano Banana and Kling in ComfyUI

1

Create AI Influencer Ad Videos with Nano Banana and Kling in ComfyUI

You need a UGC-style promotional video. A creator holding your product, talking about it on camera, ready to post as a Reel or Short.

No filming. No actors. No editing.

Model photo and product image in. Talking influencer video out.

Run it now on Floyo!

Why This Workflow

Single-pass text-to-video struggles with precise product placement. The product ends up in the wrong hand, at the wrong angle, or hallucinated entirely.

This workflow solves it by splitting the job in two. Stage 1 places the product accurately in a photo. You approve it before any video runs. Stage 2 animates only what's already correctly placed. Clean output every time.

  • product placed naturally in the model's hands or scene

  • you approve the photo before committing to video generation

  • Kling animates with natural movement and lip-synced dialogue

  • vertical output ready for Reels, Shorts, and social ads

  • no filming, editing, or animation skills required

How It Works

Stage 1: Nano Banana (Promotional Photo)

Upload your model image and product image. Write a prompt describing how the product should appear in the scene. Nano Banana combines both into a realistic ad photo, the model holding or presenting the product naturally.

Run it multiple times. Each generation produces a different pose, framing, and placement. When you find one that looks right, save that image. That's what feeds Stage 2.

Stage 2: Kling (Talking Video)

Upload your approved photo into the final image input. Write the dialogue prompt, what you want the influencer to say. Kling animates the image with subtle body movement, natural expressions, and lip-synced speech. Output is a short vertical video ready to post.

Key Inputs

Model Image

The person whose look carries through both stages. Clean, well-lit shot with a simple background.

Works well with:

  • studio or neutral background portraits

  • clear upper body or full body shots

  • forward-facing or slight angle poses

Product Image

The item you want featured. Clean product shot on a neutral background for the most accurate placement.

Works well with:

  • single product on white or simple background

  • clearly visible product with readable details

  • skincare, supplements, fashion, tech accessories, food products

Stage 1 Prompt

Describe how the product should appear in the scene.

Examples:

  • "holding the serum bottle at chest height, warm studio lighting, lifestyle photography, natural smile"

  • "presenting the product toward camera, clean background, beauty editorial style, soft light"

  • "product placed on table beside model, natural daylight, café setting, candid lifestyle shot"

Stage 2 Dialogue Prompt

Write exactly what you want the influencer to say. Keep it short, 1 to 3 sentences works best for social content length.

Examples:

  • "says: I've been using this every morning and my skin has never looked better. Seriously, try it. Link in bio."

  • "says: This is the one product I actually notice a difference from. Three weeks in and I'm not stopping."

  • "says: Found this and I'm obsessed. The texture is incredible — you need to try it."

What This Is Great For

E-commerce and product marketing: Generate promotional videos for product pages, social ads, and email campaigns without a production shoot.

Social content at scale: Each Stage 1 run produces a different pose and framing. Each Stage 2 run can carry different dialogue. Test multiple angles and scripts from the same two input images.

Small brands without production budgets: Professional UGC content requires model fees, a videographer, and post-production. This workflow produces the same format from two images and a prompt.

Multilingual campaigns: Change the Stage 2 dialogue prompt for different languages. Same visual, different speech.

What to Watch Out For

Get Stage 1 right before moving to Stage 2. The video only animates what's already in the photo. A weak product placement in Stage 1 produces a weak video. Regenerate Stage 1 freely, it's fast.

Keep dialogue short. 1 to 3 sentences is the sweet spot. Longer scripts can affect lip-sync accuracy and produce clips that run too long for the format.

Input image quality is the ceiling. Blurry, low-resolution, or poorly lit source images reduce quality in both stages. Clean, well-lit inputs produce the best output.

Products with complex geometry or very small packaging may need more Stage 1 regenerations to place accurately.

1