Download
2 variants available

13
79
License:
Qwen Research License AgreementQwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.
Extend a picture on any side with Qwen Image 2.1 and my own outpaint LoRA (Optional), and the prompt writes itself. Qwen3-VL describes your picture the same way the LoRA's training captions were written, and Stitch Inpaint puts your picture back on top of the result.
The first run grows the record-shop crop from the sample inputs on all four sides at once, from 539 × 669 to 896 × 1152, in about 7 seconds on an RTX 5090.
What's inside
Load Image + Pad: drag any border outward, on as many sides as you like. Flat gray (#808080) fill, a 1 MP canvas in steps of 32 and a 32 px seam feather: the setup the LoRA was trained on.
Auto prompt: Text Generate runs the Qwen3-VL 8B encoder that Qwen 2.1 already loads (no extra download) and describes your picture with the instruction that wrote the training captions. The final prompt is the outpaint instruction, then
Scene:, then that description: the same shape as the training captions. Show Text displays it before it goes to Qwen 2.1.Extra detail (optional): one line for what your picture doesn't show yet. It goes to the captioner, which writes it into the description. The sample's line is
Also include in the prompt: she wears a short solid black pleated skirt.Clear it or rewrite it for your own picture.Encode + sample: Text Encode Qwen Image 2.1 takes the padded canvas as
image_1at resolution 0; 25 steps, CFG 1,euler/simple, fixed seed, with my outpaint LoRA at strength 1. The LoRA is optional: turn its row off, or skip the download, and the workflow runs plain Qwen 2.1.Stitch + save: Split Image with Alpha drops the alpha channel that Qwen 2.1's VAE decodes, Stitch Inpaint puts your picture back over the result with a 32 px blend and tone match, Compare slides the result against the padded canvas, and Save Image writes a PNG with the workflow embedded.
Quick start
Update ComfyUI. You need a build with Text Encode Qwen Image 2.1 (mid-September 2026 or newer; I tested on 0.37.0).
Install AusBoss nodes from ComfyUI-Manager (search "AusBoss", registry id
ausboss-nodes) and restart. You need 2.2.0 or newer, which is what Manager installs.Unzip
sample_inputs.zipintoComfyUI/input/(it holdsausboss_record_store_crop.pngandausboss_rain_street_crop.png).Drag the workflow in and download the files from the card. The LoRA is optional but recommended; without it, the LoRA Loader warns once and the workflow runs plain Qwen 2.1.
Queue. For your own picture, drag the pad edges on Load Image + Pad and clear or rewrite the extra-detail line.
Models
qwen_image_2.1_int8_convrot.safetensors →
models/diffusion_models· 7.26 GBqwen3vl_8b_int8_convrot.safetensors →
models/text_encoders· 9.35 GBqwen_image_2.1_vae_bf16.safetensors →
models/vae· 676 MBoptional, recommended qwen-image-2.1-outpaint.safetensors →
models/loras· 159 MB
BF16 versions of the Qwen files are on the same page: Comfy-Org/Qwen-Image-2.1. The LoRA's model card, with training details and a held-out comparison, is at ausboss/Qwen-Image-2.1-Outpaint-LoRA.
Custom nodes
Only ComfyUI-AusBoss, 2.2.0 or newer (Manager installs it): free and open source. Text Encode Qwen Image 2.1, the reference cache, Text Generate and Split Image with Alpha are core ComfyUI.
Manager can't find AusBoss? In the classic Manager window, if the Channel box is empty, pick default, then click Install Missing Custom Nodes again. See what to click.
About the LoRA (optional)
I trained it with ai-toolkit on 924 before/after pairs from 231 pictures: my own set plus openly licensed images (Flickr photos from CommonCatalog CC-BY, public-domain museum art from PD12M, and anime from anime-with-caption-cc0). Each picture was cropped four ways, drawn from six layouts (all sides, one side, opposite sides, a corner, three sides, a small window), and padded with flat gray. The caption is the outpaint instruction, usually followed by Scene: and a Qwen3-VL description of the whole picture. Rank 32, learning rate 1e-4, 1500 steps. It's free on Hugging Face, direct download, no login.
Settings that worked, and what didn't
Keep the outpaint instruction first and unchanged:
Outpaint the image: replace the solid gray areas with a seamless continuation of the scene, keeping the existing picture unchanged.It's the LoRA's trigger.Flat gray #808080 and a 1 MP canvas: what the LoRA was trained on. Load Image + Pad's budget scales a small picture up and a big one down so the canvas lands on 1 MP. I tried 1.5 MP with the LoRA: no reliable gain (closer to the original in 6 of 10, clearly worse in 2), so it stays at 1 MP. A big photo comes back at 1 MP.
No noise mask: pinning the known area with Set Latent Noise Mask drew a visible rectangle at the seam on Qwen 2.1. The LoRA holds the picture in place on its own, and Stitch Inpaint puts it back.
Describe the picture you have: I also tried letting Qwen3-VL look at the gray canvas and describe the finished picture. It invented things (a piano in the record shop) and landed further from the real pictures, so the caption describes your picture and the extra-detail line carries anything new.
Why the LoRA: Sometimes Qwen Image 2.1 outpaints just fine on its own, and sometimes it needs help. I tested 20 pictures that I cropped myself, so the original is the answer key, with a small and a large border, with and without the LoRA (same prompt, same seed). With a small border, plain Qwen 2.1 shifted or rescaled the picture in 14 of 20, so the new area didn't line up with the original at the seam; with the LoRA, none did. With a large border, plain Qwen 2.1 gave up on 2 of 20 and left a gray smear around the picture; the LoRA painted a real scene in all 20. When plain Qwen already does fine, the LoRA changes little, and sometimes the plain render looks a bit better (last gallery image).
Weak spots: a flat backdrop (a plain studio wall, a clear sky) can get a soft vertical smudge on the new side, and a very small window (under about 15 % of the canvas) leaves a lot to invent. Extend in two passes, or try another seed.
Speed
Measured on an RTX 5090 (32 GB) with the models loaded: 6.9 s for the sample, caption included, and 6.0 to 7.2 s across 15 pictures at 1 MP.
I haven't tested smaller cards yet. ComfyUI offloads what doesn't fit, so expect it to run slower rather than fail.
About the gallery
The two PNGs are straight outputs of this workflow from the sample crops (record shop seed 26092503, rain street 20260925); drag one into ComfyUI to load its exact settings, extra-detail line included. The full pictures were made with Qwen Image 2.1 and cropped for this; the LoRA never saw them in training. The last image compares plain Qwen 2.1 and the LoRA at the same prompt and seed on three of my pictures: one where the LoRA is better, one where the plain render arguably looks better, and one where they just differ. The cover is cut from the real render.
License & credits
Qwen Image 2.1 is © Alibaba Qwen, released under the Qwen Research License: non-commercial use only; commercial use needs a separate license from Qwen. The outpaint LoRA is mine and follows that license when used with the model. The Qwen3-VL 8B encoder and the model files are Comfy-Org's repackage, under the same license. Training pictures: CommonCatalog CC-BY, PD12M (public domain / CC0) and anime-with-caption-cc0, plus my own. The workflow and the AusBoss nodes are free.
Changelog
v1.0 (2026-09-25): first release.
Questions or bugs: GitHub issues or the comments here. Post what you make with it; I read everything. I post new workflows on X @Zanzibased and GitHub.

