Sign In

Qwen Image 2.1 Consistency LoRA — Edits That Don't Move

Download

1 variant available

bf16 SafeTensor

qwen-image-2.1-consistency.safetensors

BF16, good balance • 152.05 MB

Verified:

Type
LoRA
Stats

319

Reviews
Published

Sep 27, 2026

Base Model

Qwen 2.1

Training
Steps: 1,500
Usage Tips
Strength: 1
Hash
AutoV2
4F44ADA1BE
default creator card background decoration
Followers - 36

36

Likes - 214

214

Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.

02_crt_watercolor.jpg

A LoRA for Qwen Image 2.1 edits: the edit happens on the original's frame, so nothing moves.

Ask Qwen Image 2.1 for a watercolor, a comic or an anime version of a picture and it comes back a few percent taller, sometimes nudged sideways, different for every seed. A waistband lands 40 px lower, a horizon jumps 25 px, a face no longer lines up with the original. On a local edit ("make the cardigan navy") it also repaints things you didn't ask about: hair strands, signs, texture. That breaks anything that stacks the edit on the original: masks, stitching, before/after sliders, video frames.

With this LoRA the same prompt makes the same edit, on the original's frame. No trigger word: load it and write your edit instruction as usual.

What it fixes

Held-out pictures it never trained on, same prompt and seed, with and without the LoRA:

  • Restyles (36 edits: watercolor, comic, anime, oil, ink...): worst corner off, median 24.3 px without the LoRA, 1.6 px with it. 75 % land under 3 px, against 3 % without.

  • Local edits on people (18 recolors and removals): pixels outside the edited thing that change noticeably, 12.8 % without, 7.5 % with. PSNR outside the edit 27.7 → 31.6 dB. Every edit still happened.

  • The look stays Qwen's. Two seeds of plain Qwen differ from each other by about as much as the LoRA's restyles differ from plain Qwen's.

The gallery shows the drift up close: the dashed line is where a feature sits in the original, the arrow is how far it moved.

How to use it

  1. LoraLoaderModelOnly right after the model loader, strength 1.0 (lower lets some drift back).

  2. Text Encode Qwen Image 2.1: your picture as image_1, resolution 0, the edit instruction as the prompt.

  3. KSampler on the encoder's latent output: 25 steps, CFG 1, euler / simple, denoise 1. Sample on that latent: a latent of any other size makes Qwen zoom by the size ratio, and no LoRA can undo that.

  4. VAE Decode → Split Image with Alpha (the Qwen 2.1 VAE decodes RGBA).

My free Qwen Image 2.1 Edit workflow is already set up this way; add the LoRA after its model loader.

Limits

  • Tested at about 1 MP with the settings above. Not yet tested with 4- or 8-step turbo LoRAs, at 2 MP, or above CFG 1.

  • Comic and anime restyles redraw every outline, so some shapes still wobble a few pixels on their own.

  • It keeps the picture in place; it doesn't make weak edits stronger.

Training

  • ai-toolkit (qwen_image_2) on the Comfy-Org INT8 convrot base, the same weights ComfyUI runs. Rank 32, alpha 32, AdamW8bit, learning rate 1e-4, batch 1, 1500 steps.

  • 950 edit pairs from 257 pictures. Each pair started as a real Qwen Image 2.1 edit; I measured its drift and took it out, so every target sits on its source's frame. Many pairs run backwards (the edit is the source, the untouched original is the target), and local edits keep the original's pixels outside the edited thing. Every edit was checked by eye.

The full numbers, the step 2000 file (tighter alignment, slightly paler paintings) and the figures are on the Hugging Face model card.