Name a thing in your picture. You get it as a transparent PNG, and your picture with it gone. Nothing to mask or paint.
Why use this workflow
A background remover gives you the cut-out. An object remover gives you the picture without the thing. This gives you both from one word, and the two line up:
You type a name, not a mask. "the dog", "the woman on the left". SAM 3 finds it.
You get two layers. The thing on a see-through background, and your picture with the thing taken out. They are the same size, so you can stack them, move the thing, or put something else behind it.
Its shadow and reflection go too. Qwen Image 2.1 takes the thing out together with what it left on the ground or the table.
The rest of your picture is not repainted. Only the thing, a band around it, and what changed with it are taken from Qwen's picture. Every other pixel is your own.
It stays in place. It runs with my Consistency LoRA, so Qwen's picture lines up with yours.
How to use it
Load your image. Optional: drag the blue handles to crop it first.
In What to lift out, name one thing.
Optional: in What to leave in its place, say what should be there instead. Leave it as it is to just take the thing out.
Run.
The cut-out and the picture without it are saved as two PNG files. Before and after slides between your picture and the result. The request sent to Qwen shows in Show Text.
Faster, optional
The Viggle Turbo row is switched off. Download the file below, switch the row on and set Steps to 8. A run takes about half the time. In my 18 test runs the results looked the same as at 25 steps.
What you need
AusBoss nodes 2.6.2 or newer. In ComfyUI-Manager, search "AusBoss". Already have them? Update first.
ComfyUI 0.38 or newer.
The Qwen Image 2.1 models:
qwen_image_2.1_int8_convrot.safetensors →
models/diffusion_models· 7.26 GBqwen3vl_8b_int8_convrot.safetensors →
models/text_encoders· 9.35 GBqwen_image_2.1_vae_bf16.safetensors →
models/vae· 676 MB
SAM 3, which finds the thing: sam3.1_multiplex_fp16.safetensors →
models/checkpoints· 1.75 GBMy Consistency LoRA: qwen-image-2.1-consistency.safetensors →
models/loras· 159 MB
Optional, for the speed row: Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r128.safetensors → models/loras · 680 MB. The workflow runs without it as long as its row stays off.
Tips
Name the thing the way you would point at it. By what it is, its color, or where it stands: "the dog", "the woman on the left".
Solid things work best. See-through parts keep what was behind them: the cut-out of a car has the street in its windows, and a bicycle keeps the wall between its spokes.
If a piece of it stays, run it again. The seed changes each run.
A long, hard shadow can leave a faint mark. In my test with a dog on a beach at sunset, a soft streak stayed where the shadow was.
Hair is cut along its outline. Loose strands are not in the cut-out, and a few thin ones far from the head can stay in your picture.
A thing partly hidden comes out as the part you can see. An ear cup behind hair is cut along the hair.
What goes with the thing may go missing. A headphone cable, or a book or a bag a person holds, can be taken out of the picture without being in the cut-out. Naming both ("the woman and her bag") did not work in my tests: then only the bag came out.
Something held in a hand can take the hand with it. In my test with a coffee cup, the hand went too. Typing "her empty hand resting near her chin" in the second box kept a hand there.
To change what someone wears, use my Outfit Swap workflow. This one takes a thing out. The second box is for small fills, like "an empty wooden bench".
A close-up leaves little of your picture. When the thing fills half the frame, what was behind it is Qwen's guess. If taking it out changes the light, the whole result is Qwen's picture.
Your picture is scaled to about 1.5 megapixels first, and both results come out at that size. You can change it in the first box.
About 12 seconds a run on an RTX 5090, about 6 with the speed row.
Credits
Qwen Image 2.1 is by Alibaba Qwen, under the Qwen Research License: non-commercial use only. SAM 3 is by Meta, under the SAM License. The optional speed LoRA is Viggle Turbo v0.3 by Viggle, under the Qwen Research License: research use only.
v1.0 (2026-10-05): first release.
Questions or bugs: GitHub issues or the comments here. More from me on X @Zanzibased and GitHub.



