Sign In

Copita: Train an Anima / SDXL Style LoRA from a Single Image

6

Aug 1, 2026

(Updated: 21 days ago)

tool guide
Copita: Train an Anima / SDXL Style LoRA from a Single Image

Hello — I'm Io Sakaki of Studio Masakaki. I research AI illustration in Japan: nearly 300 long-form, hands-on investigations into LoRA training, prompting and retouching so far, written for several thousand paying readers. Copita is the tool that grew out of that research — an application that trains a style LoRA for SDXL or Anima from a single reference image.

Copita creates a matching comparison image from one reference image, then extracts the difference between the two images as a LoRA. If only the art style differs, the result is a style LoRA. If only the line weight differs, the result can work as a concept-slider LoRA similar to a Bold LoRA. Training time depends on the settings, but an RTX 3060 generally takes a little under an hour with the lightweight preset, while an RTX 4080 typically finishes in 30–40 minutes.

00220-2862861084.png

This guide covers installation and basic operation, then goes deeper into every major parameter and the settings that help you get the best result for your GPU. If you run into a problem, see Troubleshooting and FAQ near the end.

Everything in this guide was made with Copita — including Anima Bold v1, the free line-weight slider LoRA on my Civitai page. If you want to judge the output before reading further, download it and move the slider yourself. This article doubles as the story of how it was built.

https://civitai.com/models/2823195

What Copita Can Do

Copita is a combined Anima/SDXL training application that uses the copy-machine method to create a style LoRA from one image. Internally, it extracts the difference between two images. This makes it useful not only for style LoRAs, but also for adjustment LoRAs that increase or decrease a property—such as line thickness or eye size—as if you were moving a slider.

The difference from conventional training on tens or hundreds of repeated dataset images is summarized below. The copy-machine method makes it easy to create a LoRA and isolate a particular feature, but it is not good at learning concepts that cannot be expressed as a clear difference between two images.

004_en.png

Copita makes this copy-machine method easy to use. Load a reference for the concept you want to learn on the left, follow the on-screen steps, and Copita helps you obtain a suitable comparison image. With a pair like the one below, for example, you can create a copy-machine LoRA that makes line art thicker.

005_en.png

Features shared by both images—the character's appearance, clothes, pose, composition, and background—are less likely to be learned. The parts you changed—linework, coloring, shading, texture, palette, brushwork, and so on—are extracted and learned as the LoRA. In other words, LoRA quality depends heavily on how well you can make an image pair in which only the intended concept differs.

Copita also includes an experimental LoRA Fusion feature that merges two to four LoRAs you created previously. Its purpose is to combine LoRAs of the same kind that are individually imperfect—for example, each changes the pose in a different way—so the intended effect remains while unrelated side effects are reduced. The method and results are examined later in this guide.

006_en.png

Copita is designed for the following environment:

・Windows 10 or 11

・NVIDIA graphics card; tested with 12 GB of VRAM or more

・Python 3.10 or 3.11; Python 3.11 is recommended

・At least 20 GB of free storage is recommended

Even if your system-wide default is Python 3.12 or 3.13, Copita can prepare its own private Python 3.11 environment during first-time setup. It does not change the system-wide Python installation or the Python environments used by Forge Neo or Stability Matrix.

Copita is an original Studio Masakaki application for authorized purchasers. You may not share, reproduce, or redistribute the application, its ZIP archive, included files, article-only distribution links, or download URLs without permission.

By downloading, extracting, launching, or using Copita, you agree to its Terms of Use. The same terms are included as a text file with the application. Please read them before use.

Read the Terms of Use and get Copita on Gumroad

https://iosakaki.gumroad.com/l/copita

Installing Copita

1. Download Copita

After reading and accepting the Terms of Use, download the latest Copita International Release ZIP from the product page and extract it anywhere convenient on your PC.

Download Copita on Gumroad

https://iosakaki.gumroad.com/l/copita

The international release begins with v1.2. It includes the revised preprocessing for Anima images, improved training accuracy, layer-specific learning controls, stability improvements, and bug fixes.

2. Run setup_first_time.bat

The extracted folder contains two BAT files:

・setup_first_time.bat: performs initial setup; normally used only once

・start_copita.bat: launches Copita; use this every time afterward

image.png

For the first launch only, double-click setup_first_time.bat. Copita automatically builds the required environment, downloads the necessary files, and prepares sd-scripts for LoRA training and the model used for tagging. This step takes a little while, so let it run without closing the window. Setup took about eight minutes on the test machine.

010_en.png

When the screen displays Setup finished as shown below, setup is complete. Press any key to close the window.

011_en.png

3. Run start_copita.bat

Double-click start_copita.bat. After a short wait, your browser should open automatically and display Copita. If it does not open, manually visit http://127.0.0.1:7865, which is shown in the console window. You can usually open it by holding Ctrl and clicking the address.

012_en.png

To exit Copita, close both the browser tab and the console window that opened when Copita started.

Basic Workflow

This is Copita's main screen. Use the tabs at the top to switch between the main Difference LoRA Training Mode, the experimental LoRA Fusion Mode, and Training History, where you can review previous Copita training runs.

014_en.png

Let us train a style LoRA in Difference LoRA Training Mode.

1. Load the Style Image

On the left, under Style Image, load one image in the style you want Copita to learn. If you want it to learn facial features and linework, a bust-up composition is usually better than a full-body image where the face is small. A near-frontal pose, simple clothing, a smile or neutral expression, and a simple background generally produce a more stable LoRA. We will use the following image for this example.

image.png

Choose a PNG or JPG with an aspect ratio from roughly 1:1 to 3:4. Copita automatically downsizes oversized files internally, so you do not need to resize them first. The two images may have different pixel dimensions, but their aspect ratios should match as closely as possible.

What Makes a Good Input Image for the Copy-Machine Method?

Suppose you create a style LoRA from a busy character illustration. Copita may learn not only the desired palette, brushwork, face, and way of drawing hands, but also the water droplets, dust, visible breath, cropped legs, ahoge, and other details as part of “the style.” Prompts can control some of those features, but at high LoRA strength many characteristics of the original illustration may appear together.

To avoid introducing unrelated concepts, use a reference in which the desired features are large and clear and unnecessary details are minimized. Avoid sweat, wet skin, speech bubbles, floating splashes, emphasis symbols, and other features unrelated to the target concept—especially features that are hard to express with Danbooru tags.

・To change full-body proportions or body type, use an image that shows the full body. If you do not want to learn the coloring, train from a monochrome image and include monochrome in the caption.

・To exclude as many elements as possible besides the character's appearance, a white background is effective. Be aware that the LoRA will not learn how the style handles backgrounds and may develop a preference for white backgrounds.

2. Create a Comparison Image That Differs Only in the Target Concept

After you load the first image, Copita displays a text prompt on the right. This is an example prompt for asking ChatGPT or another image-editing AI to create a comparison image that Copita can learn from.

Before using any service for this purpose, confirm that its terms permit generated output to be used for AI training. Some models explicitly prohibit it.

image.png

What Is a Comparison Image?

To train a style LoRA, the comparison image must show the same person, clothes, pose, composition, and layout as the style image on the left, while using a different art style. In practice, converting the reference into a photorealistic image often works well. The subject remains conceptually the same, but the linework, coloring, and rendering change completely, which makes the original style easier to isolate.

Copita automatically proposes a prompt for obtaining this comparison image. Paste the original image and the prompt into an AI service that supports image editing to create a result like the one below. Again, confirm beforehand that the service permits images generated for LoRA training.

image.png

Copita provides three prompt templates: Normal, Special body type, and Concept slider. Use Normal for an ordinary style LoRA. Use Special body type when proportions or body shape are part of the style—for example, a chibi character—because the comparison must change the style and body proportions substantially while keeping the depicted subject aligned.

020_en.png

Use Concept slider when you want a LoRA that adjusts a property such as line thickness or eye size. The template asks for a comparison image that keeps the person, pose, expression, clothes, hairstyle, composition, and background in the same positions while changing a specified feature. Replace the blanks with the concept and target change you need.

For example, to make a LoRA that thickens linework, prepare a pair like this and request that the line art be made thicker.

021.png

Compare the two images carefully. If clothing changed, the face angle shifted, fingers broke, the aspect ratio changed substantially, or extra margins made the person smaller, those unrelated differences may be learned and reduce LoRA quality. The accuracy of this pair directly determines the quality of the result, so revise it until unnecessary differences are minimized. If you ask ChatGPT to create the comparison image, requesting multiple candidates and choosing the best one can also improve quality.

3. Load the Comparison Image and Tag the Pair

Next, use AI tagging to describe what appears in the images.

Copita provides Tag Images and Tag in More Detail. Start with ordinary tagging. If it produces too few tags or misses clothing, accessories, sweat, or background elements, try detailed tagging. Detailed tagging also increases the chance of false positives, so always review the result yourself.

image.png

The pixel dimensions appear under each input image. Check them before continuing. Different sizes are acceptable as long as the aspect ratios match.

Tagging Tips

Always review detected tags and remove incorrect tags, duplicates, and anything that describes style or subjective quality. Copita automatically removes tags in categories such as the following:

artist names, series names, character names, artist_name, watermark, signature, text, bad_hands, monochrome, and other style or concept tags that describe how an image was drawn rather than what it depicts.

For example, if white shirt is present, the more general shirt is unnecessary. If red bow is present, red ribbon is redundant. Remove artist tags and style tags such as anime screenshot or watercolor_(medium) if they appear.

Given the following detection result:

1girl, solo, long_hair, looking_at_viewer, sailor uniform, blush,full-face_blush, simple_background, shirt, skirt, blonde hair,white shirt, blue skirt, pleated skirt, long_sleeves,white_background, closed_mouth, twintails, school_uniform,serafuku, white_shirt, yellow hair, flat color

remove sailor uniform, blush, shirt, skirt, yellow hair, and flat color. More specific or normalized tags—serafuku, full-face_blush, white shirt, blue skirt, and blonde hair—already cover the first five. Keep both blue skirt and pleated skirt, because they describe different concepts. Remove flat color because it is exactly the sort of stylistic property the LoRA should learn.

Conversely, if an unwanted property such as monochrome is not detected, add it manually. Otherwise, the result may become a style-plus-monochrome LoRA.

The most important rule is that the two captions must be identical. From the model's point of view:

🤖 “Both images were drawn from exactly the same prompt.”

🤖 “But when I compare them, this one concept is very different.”

🤖 “I should learn how to reproduce only that difference.”

This makes it easier to isolate a clean concept as a LoRA. Any uncaptioned feature—such as spray or visible breath—may instead be learned together with the target concept.

025_en.png

4. Enter the Training Settings

Once the pair and tags are ready, move to Training Settings. Select SDXL or Anima as the base model type, click the folder button, and choose the parent Models folder that contains your usual checkpoints, LoRAs, and related model files.

image.png

Enable Remember Models Folder if you want Copita to select the same folder automatically next time.

For your first Anima training run, select anima_baseV10.safetensors as the base model. You can instead train against the checkpoint where you plan to use the LoRA, but a merged model such as WAI ANIMA may already impose a strong anime rendering style. That style can overpower the LoRA and make its effect difficult to evaluate, so a merged model is better reserved for later experiments.

image.png

For Anima training, you must also specify the text encoder (qwen_3_06b_base.safetensors) and VAE (qwen_image_vae.safetensors) that you normally use for generation. Copita loads them automatically when they can be found under the selected Models folder. You can also save these as your Anima model settings so the same files load automatically next time.

If you are unfamiliar with text encoders or VAEs, refer to my introductory guide to Anima and Forge Neo (Japanese — my main publication is on FANBOX; machine translation works well).

Next, give the LoRA a name and choose its output folder. If you selected a Stability Matrix Models folder, Copita should automatically detect its Models/LoRA folder.

Copita warns you if a LoRA with the same name already exists in the destination. Adding a suffix such as v1 is useful. Once you reach names like mylora_v12, however, it can be difficult to remember the source pair and target architecture, so choose a name that identifies both the image pair and whether the LoRA is for SDXL or Anima.

image.png

5. Choose a Training Preset

Choose one of the three presets. Anima and SDXL use different values behind presets with similar roles. For a first attempt, select the Normal Training Set.

image.png

If Normal places too much load on the GPU or takes too long, use the completion-focused Stable Training Set. If Normal does not reproduce the face or rendering strongly enough, try the Detailed Training Set. Detailed is recommended when you need to learn small design features such as eyelashes or fingernails.

Anima LoRA training is heavier than SDXL training. On an RTX 4080, an Anima LoRA generally takes about 30–40 minutes with Normal and 45–60 minutes with Detailed. On an RTX 3060, Stable took about 40–45 minutes in testing, while Normal took roughly two hours. You can reduce the number of steps if training takes too long, but step count directly affects the result. Experiment to find the best balance for your hardware.

Open the Training Settings tab to inspect or fine-tune the values selected by each preset. Every parameter is explained later in this guide.

image.png

6. Start Training

When everything is ready, click Start Training. Beforehand, close Stability Matrix, Steam, and any other application using significant VRAM.

image.png

Training proceeds in three stages:

1. Train a copy-machine LoRA for the left image.

2. Train a copy-machine LoRA for the right image.

3. Extract the difference.

Although the interface begins with only one pair, Copita performs two complete LoRA training runs, so the process takes time. More steps increase the duration of each run. A larger batch can shorten it by processing more work simultaneously, if your GPU has enough VRAM.

Track progress in the Training Log at the bottom. You can stop and restart at any time with Stop Training. The finished LoRA is written to the destination you selected and also copied to Copita\output as a backup.

image.png

If training takes an abnormally long time, it may have exceeded available GPU memory. Open Task Manager and check VRAM use. If VRAM is full, stop training, close programs consuming VRAM, reduce batch or dim, and try again.

033_en.png

In this example, VRAM was full, so Stability Matrix was closed to free memory.

When training completes successfully, the green progress bar reaches 100% and Copita displays the saved LoRA location. You can now test the result in Forge Neo or another compatible interface.

image.png

7. Review Training History

The Training History tab automatically records the style image, comparison image, LoRA name, training settings, and whether the run completed or was stopped.

image.png

The LoRA Generation Example field starts empty. Drag in a representative output image to keep a visual note of what the LoRA produces. This becomes especially useful after you have created many LoRAs and can no longer remember which image pair produced each one.

Testing the Finished Style LoRA

RTX 4080 with Detailed Training

For this test, I used the pair below to train a style LoRA on an RTX 4080 with 16 GB of VRAM and the Detailed Training Set, then applied it in Anima.

スクリーンショット 2026-08-01 224153.png

Because the LoRA was trained directly on Anima's base model, anima_baseV10.safetensors, it should produce a broadly similar effect in most Anima-derived checkpoints. Applying it to WAI ANIMA will inevitably pull the result toward WAI ANIMA's characteristic face and rendering. In that case, the point of applying this LoRA is to blend the two styles.

The following XYZ Plot compares strengths. The LoRA is strongest at the top and becomes weaker toward the bottom. The natural-language prompt asks for a girl in a straw hat standing in a sunflower field.

image.png

The bottom row has no LoRA applied. The composition remains mostly stable, while the reference style begins to show clearly between strength 0.5 and 0.75.

At strength 0.8, random prompts produce results like the following.

image.png

Four different characters, four different settings — and the same palette, rendering and eye design in all of them. That transfer is exactly what a style LoRA is for.

If the Style Is Not Reproduced Well

If a test generation does not resemble the style, the checkpoint's native rendering or your quality tags may be overriding it. Remove style-related instructions such as 4K, ultra detailed, bold line, or watercolor_(medium) and try again. If that does not help, see Troubleshooting and FAQ.

Training Presets Explained

This section explains the parameters behind each preset. Once you are comfortable with Copita, try adjusting them yourself. For a broader explanation of learning rate, dim, and alpha, you can also consult Studio Masakaki's general LoRA-training guides. (Japanese, on FANBOX)

Differences Between Presets

Selecting one of the three presets automatically applies the values shown below.

image.png

Anima Presets

Anima Normal Training Set is the standard configuration: steps 300, LR 0.00016, batch 2, dim 8, alpha 1, and difference strength 1.25. It is tuned to learn the style reliably in as little time as practical.

Anima Stable Training Set is for systems where Normal stops or becomes extremely slow: steps 200, batch 1, dim 8, alpha 2 to help compensate for fewer steps, and difference strength 1.25. You can enable lowram manually to reduce VRAM use further, at the cost of speed.

Anima Detailed Training Set is for cases where the face or rendering is too weak with Normal: steps 400, LR 0.00012, batch 2, dim 16, alpha 8, and difference strength 1.25. It also enables Prioritize Color, described below, so color and rendering are learned more strongly.

SDXL Presets

SDXL Normal Training Set is the standard configuration: steps 500, LR 0.0001, batch 2, dim 16, alpha 16, and difference strength 1.0. SDXL is lighter to train than Anima, so this preset uses somewhat more steps and larger dim/alpha values to capture the style clearly.

SDXL Stable Training Set is for systems where Normal runs out of VRAM or stops partway through: steps 300, LR 0.0001, batch 1, dim 16, alpha 16, and difference strength 1.0. Reducing the batch to one limits VRAM use while retaining a minimum useful amount of SDXL training.

SDXL Detailed Training Set is for cases where Normal does not capture the face, coloring, or line habits strongly enough: steps 600, LR 0.00008, batch 2, dim 32, alpha 32, and difference strength 1.0. It can pick up finer details, but it is also more likely to learn unwanted habits from the input, so cleaning the pair and tags becomes especially important.

What Each Parameter Means

steps

This is the number of training iterations used to create each copy-machine LoRA. Copita intensely trains one image until it behaves like a copier, does the same for the second image, then extracts the difference. Reaching that copy-machine state requires a sufficient number of steps.

More steps generally make the style easier to learn but also take longer. Too many can memorize the source content, lock the face or composition, and damage fine detail. Copita accepts 100–1600 steps.

LR — Learning Rate

Learning rate controls the strength of each training update. A higher value produces effects sooner, but an excessive value can make lines and colors rough and introduce unrelated habits. The preset value is normally the right starting point.

Think of a high LR as cramming overnight: you finish the material quickly, but what you learn is less precise.

batch

This is the number of images processed at once. The copy-machine method has only one source image, so a larger batch processes the same image simultaneously to drive intentional overfitting. batch 2 is often faster but uses more VRAM. Use batch 1 if memory is insufficient.

Copita allows 1–4, but values above two may fail to improve speed and can even slow it down. In most cases, choose one or two based on available VRAM.

dim — Dimension/Rank

Think of dim as the size of the container available for storing a style or concept. A larger value can capture finer facial and rendering details but also makes it easier to learn irrelevant features.

A complex magical-girl outfit may require a high dim to reproduce its details. A LoRA intended only to thicken linework should not need that capacity; a high value may learn many concepts besides line thickness and make the result harder to use.

A useful shorthand is: high dim is a detail-obsessed overthinker; low dim is a simple generalist. Copita allows values from 4 to 64.

alpha

alpha works with rank/dim to control how the LoRA is learned and scaled. Roughly, a smaller alpha relative to dim tends to produce a gentler effect, while a larger one tends to act more strongly. Because it also interacts with learning rate and stored weights, it should not be understood as merely another inference-strength slider.

There is no single correct ratio. Many LoRAs published on Civitai use an alpha around one-eighth to one-half of dim. Copita allows 1–64.

Difference Strength

This controls how strongly the right-image copy LoRA is subtracted from the left-image copy LoRA. A higher value makes the extracted difference stronger, but excessive values can destabilize the result. The default is 1.25, and the available range is 0.6–1.8. Because you can also adjust LoRA strength during generation, you usually do not need to change this immediately.

Save Three LoRA Variants at 100-Step Intervals

This saves intermediate LoRAs in addition to the final version. Sometimes a model from before full training produces a more natural effect, so the variants make comparison easier.

At most three are retained. Even if you train for 800 steps, Copita saves the 600-, 700-, and 800-step versions. More precisely, it keeps the two most recent 100-step checkpoints before the final result, plus the final LoRA.

Prioritize Color

This experimental option addresses a tendency of the Anima copy-machine method to learn linework more reliably than color and rendering. Copita internally converts only the right-side comparison image to grayscale before training, leaving the style image in color. This makes the source palette stand out more strongly as part of the difference. Try it when you want stronger color or rendering transfer.

048.png

You can keep caption tags such as red eyes and blonde hair unchanged. There is no need to add monochrome only to the right caption. By continuing to assert that the subject has blonde hair and red eyes while the right image is grayscale, the model effectively sees a very large color difference and learns the left image's color more strongly.

lowram — Low-VRAM Mode

This reduces VRAM consumption at the cost of speed. It is intended mainly for systems with 8 GB of VRAM or less. Leave it disabled when you have sufficient memory.

Layer-Specific Learning

This Anima-only v1.2 feature lets you set separate learning-rate multipliers for three module groups: Self-Attention, Cross-Attention, and MLP.

image.png

The default is 1 / 1 / 1. If you are unsure what to use, keep the defaults. See Studio Masakaki's dedicated layer-specific-learning guide for advanced usage. (Japanese, on FANBOX).

Tips for Choosing dim and alpha

dim determines how much detail the LoRA can represent, but a bigger value does not automatically create a heavier, more capable, or better LoRA. The basic rule is: use a larger value for a complex or compound difference and a smaller value for a simple, single difference.

A style LoRA needs to encode linework, coloring, eyes, shadows, palette, and other information. dim 8, used by Anima Normal, or dim 16, used by Anima Detailed, is often manageable. dim 32–64 may help with a very complex style that must be learned strongly, but it also captures sweat, droplets, composition, camera angle, pose, expression habits, and defects. This can produce an image that resembles the source yet differs from what you actually wanted.

A concept-slider LoRA such as a line-thickness control has a much simpler difference. Composition, facial structure, clothing details, wrinkle texture, and camera behavior are not only unnecessary—they are harmful if learned. A lower capacity is therefore preferable. dim 8, or even dim 4 / alpha 1, may work. If capacity is too small, however, the effect can be weak or break at high inference strength, so no one setting is universally optimal, especially for a newer architecture such as Anima.

If a slider LoRA is weak or has unwanted side effects, improve the image pair or increase steps before increasing dim. Raising dim can make it learn even more irrelevant information and worsen the problem.

Training a Concept-Slider LoRA

Next, let us train a LoRA that adjusts one concept like a slider. In SDXL, well-known examples include Bold LoRAs that thicken linework and Flat LoRAs that reduce detail for a flatter look. Applying these with negative strength can reverse their purpose: a Bold LoRA becomes a line-thinning LoRA, and a Flat LoRA can increase the amount of detail.

051.png

The copy-machine method is well suited to a single concept, but it will not necessarily create a perfectly clean Bold LoRA on the first attempt. The high-quality SDXL Bold LoRA used as a reference here was specially developed by Tsukisuwa Nana through repeated adjustment and likely repeated merging. Do not assume that every first training run will reach that quality.

First Bold LoRA Example

Let us train a Bold LoRA for Anima. Here is the style/comparison pair.

052.png

The intended result is a LoRA that makes lines thicker when applied positively, so the thick-line image goes on the left and the thin-line image goes on the right. Reversing them would create a LoRA whose positive direction makes lines thinner.

One way to make the pair is to generate the thin-line source in SDXL, use it as a reference with ControlNet Anytest v3, and apply an existing Bold LoRA at strength 1.5 to create the thick-line version.

In this case, asking SDXL for monochrome still produced a slight sepia tint, so the image was converted to zero saturation in Clip Studio Paint. A few elements marked by the pink arrows had been added during the bold-line conversion and did not exist in the original, so they were removed by hand to keep the difference as pure as possible.

053.png

You can also create the comparison image by giving the left image to ChatGPT or another editor and asking it to make the lines thicker without shifting the line art. Use whichever method is most convenient.

After loading the pair, run tagging. monochrome was missing from the detected tags, so it was added manually:

1girl, solo, long_hair, looking_at_viewer, simple_background, shirt,white_background, closed_mouth, collarbone, upper_body, medium_hair,expressionless, portrait, monochrome

054_en.png

Detailed Training finished in about 45 minutes. We can now apply the result in Anima.

055_en.png

The following outputs use different LoRA strengths. At the top, negative strength makes the lines thin. They grow progressively thicker toward the bottom.

056_h.png

The intended effect is clearly present, and negative strength correctly thins the linework. However, the result does not isolate the difference as cleanly as Tsukisuwa Nana's Bold LoRA; it also affects art style and composition. dim 16 from Detailed Training appears to have provided enough capacity to learn irrelevant features besides line thickness.

Reducing Side Effects in a Slider LoRA

The previous pair was created with SDXL, so its style was far from Anima's base model. That mismatch may have contributed to the composition changes. The effect also required a strength near 2, which is inconvenient. For a second attempt, we will use WAI ANIMA as the training base and create both images in the same WAI ANIMA style, while exaggerating the line-thickness difference between them.

First, this deliberately generic image — what Japanese users call a “masterpiece-face,” the default look a merged checkpoint falls back to when prompted with only quality tags — was generated in WAI ANIMA with:

1girl, solo, cowboy shot, monochrome, black hair, safe, black eyes,white t-shirt, bob cut, looking at viewer, white background,simple background

057.png

The image was then given to ChatGPT with the following instructions to produce one thick-line version and one thin-line version, increasing the separation between the two endpoints.

058_en.png

Load the pair and tag it. monochrome and lineart were added manually. Because positive application should thicken the linework, the thick image is on the left and the thin image is on the right.

060_en.png

This time, select waiANIMA_v10Base10 as the base model and train with the Normal preset at dim 8. Lowering dim from 16 to 8 should reduce unrelated style learning.

061_en.png

Here is the result.

062_h.png

It still affects composition, but the change appears milder than with dim 16. Because the line-thickness gap in the pair is larger, the lines also become more decisively bold.

To confirm whether higher dim really increases composition side effects, the next test used dim 64 / alpha 32 and 600 steps.

063_h.png

As expected, it learned too much besides line thickness and did more than alter the composition—the image broke down.

For a simple concept-slider LoRA, a low-dim model that is deliberately “simple-minded” appears to work better. The following result is the final test at dim 4 / alpha 1 and 300 steps.

064_h.png

The results support three practical conclusions:

・A simple slider LoRA often works well around dim 4–8.

・Excessively high dim learns unrelated features.

・A single difference-training run is unlikely to create a perfect slider LoRA.

As LoRA strength rises in the charts, the camera pulls back, the straw hat grows wider, and the whole palette darkens. Can we create a cleaner Anima slider LoRA closer to the carefully refined SDXL Bold LoRA?

Experimental LoRA Fusion

This is why I developed LoRA Fusion, available from the tab at the top of Copita.

image.png

LoRA Fusion merges and compresses two to four LoRAs with the same intended effect to create a more stable result. It combines them with user-defined weights and approximates them as a LoRA at the selected output dim.

This is not a magic process that mechanically extracts only what all inputs have in common. However, averaging LoRAs that have different side effects can make their shared intended effect more prominent while reducing unrelated changes.

For example, each Anima Bold LoRA made from the previous pair affected composition or the subject. If you train the same concept from several different pairs, each output will probably contain different side effects, while the desired “make the lines thicker” effect should be common to all of them. Fusion aims to keep that common effect and dilute the side effects by combining several prototypes.

How to Use LoRA Fusion

Using it is simple:

1. Choose two to four LoRAs from the selected LoRA folder.

2. Assign their weights.

3. Enter an output name and select one or more output dim values.

4. Click Start Fusion.

Fusion does not retrain from scratch, so it finishes quickly. Type part of a LoRA name in the filter box to narrow the list.

067_en.png

The weights do not need to total one. Copita normalizes them internally, so leaving every value at 1 is fine when you are unsure. Fusion is not limited to purifying a Bold LoRA. You could combine a face-rendering LoRA, a full-body style LoRA, and a coloring LoRA at a ratio such as 1:1:1 or 2:3:1 to build a style closer to your preference.

For example, 1:0.5:0.5 and 2:1:1 produce the same normalized ratio.

068_en.png

You cannot mix an SDXL LoRA with an Anima LoRA. Choose LoRAs for the same architecture and, preferably, the same base model. LoRAs with substantially different structures may fail to merge or may produce an unexpected result.

Bold LoRA Fusion Test

For this test, I created Bold LoRAs from several different pairs. Seven prototypes, v1 through v7, were trained; v2, v4, v6, and v7 produced the most usable effects and became the candidates for Fusion.

In LoRA Fusion Mode, specify the folder containing your LoRAs. If Difference LoRA Training Mode already remembers an output folder, that folder appears automatically.

Select two to four LoRAs. Copita lists them in selection order as Weight 1, Weight 2, and so on. Enter each weight and click Start Fusion. You cannot select more than four.

071_en.png

The source ranks can differ. In this example, v02 is dim 16, v04 is dim 8, and v06 and v07 are dim 4; they can still be merged. You may select multiple output ranks. Choosing 8 and 16 produces two files as shown below. This screenshot shows a fusion of v2, v4, and v6.

072_en.png

Here is an example generated with a fusion LoRA that combines v2, v4, v6, and v7 at 1:1:1:1.

073_h.jpg

With the earlier standalone LoRAs, raising the strength also pulled the camera back, widened the hat and darkened the palette. The fusion version suppresses those side effects: the framing and colors stay close to the no-LoRA row while the lines still thicken.

For a clearer comparison, v06 appears in the left column and the fusion LoRA in the right, at the same strengths and seed. Watch the framing and the sunflowers around the figure: in the right column they stay almost unchanged from -2 to +2.

074_h.png

Because weaker LoRAs are included and averaged, the sunflower and cloud outlines do not become quite as thick. However, the left column keeps drifting as the strength rises — the camera pulls back and the framing shifts — while the right column has much less effect on composition.

Because v4 and v6 were trained from the same pair at different ranks, I removed v4—the one with the stronger style side effect—and fused v02, v06, and v07 instead.

075_h.jpg

This version thickens the linework clearly while reducing harmful composition changes. The inputs were mixed at 1:1:1, but their measured quality suggests that further weight tuning could improve the result.

With a different prompt, strengths from -1 to +1 can now adjust line thickness with almost no composition change. LoRA Fusion remains experimental, but it appears genuinely useful for slider-style LoRAs.

image.png

One possible future direction for Copita is a sequential-training mode that accepts multiple pairs and produces several slider LoRAs in one batch. You could prepare several pairs with different line weights, tag them together, leave training running for a few hours, then fuse the resulting LoRAs. For now, the same workflow can be carried out one pair at a time.

The finished slider that came out of this chapter is published on my Civitai page as Anima Bold v1, free to download. If you build something good with Copita, I would love to see it shared there too.

https://civitai.com/models/2823195

Troubleshooting and FAQ

Q. I applied the finished LoRA, but it has no visible effect.

First, confirm that you are using the latest international release. Then raise the LoRA strength during generation to around 2 temporarily and check whether the target effect appears at all.

If the intended result—such as a closer style or thicker linework—appears gradually at higher strength, Copita did detect the difference, but the effect is not strong enough to overcome the checkpoint's ordinary prompt response.

If the output keeps reverting to generic AI-anime or “masupikao” rendering, quality tags may be painting over the learned style. Remove style-changing quality tags such as 4k and anime screenshot, then test using only the Copita LoRA as the style instruction.

A strongly fine-tuned checkpoint such as WAI ANIMA will inevitably pull the output toward its native anime rendering; that is what the checkpoint was designed to do. Training on anima_baseV10.safetensors and applying the result to the same base model reveals the learned effect with the least interference.

Q. The finished LoRA is too weak.

If the effect appears above strength 2 but is too weak at 1, remake the image pair with a more exaggerated and clearly visible difference. For a style LoRA, raising dim or alpha can also help, but it increases the risk of learning unrelated details as part of the style.

If the result is still weak, the difference may be conceptually difficult to read, or the tags may be incomplete or overly complicated because the subject itself contains too many elements.

Q. Applying the LoRA causes unexpected side effects.

If you wanted only the style but the expression or composition also copies the reference—or unrelated objects appear—the likely causes are excessive dim and alpha, misalignment between the two images, or tags that should have been removed from both captions.

Lower inference strength or remake the comparison image.

Anima itself also tends to draw features that were not requested with Danbooru tags. Generate once with the same seed and prompt but without the LoRA. If the same behavior remains, it comes from the base model rather than the LoRA you trained.

Q. My LoRA makes hands fall apart.

This can happen when a style LoRA is trained from a pair that does not show hands, especially when the target has a rough sketch-like style. The model must infer how unseen parts should look. After learning rough linework, it may decide to draw hands roughly as well. Hands are already difficult for image models, so finger counts can quickly become ambiguous.

To reduce this problem, use a pair that includes clearly drawn hands or separate the goals. First train the overall style from clean linework, then train another LoRA that converts that style into rough linework. A single “rough style LoRA” may become a “rough style plus broken hands LoRA”; separate style and roughness controls are easier to use.

Q. A neutral expression resembles the reference, but detailed instructions make it less similar.

The copy-machine method can extract only the surface appearance visible in one illustration and its comparison image. It cannot learn how the character smiles or becomes angry, how they move, or how the back of their head is drawn if those things never appear in the source.

For rich variations, conventional training on tens of dataset images is still more appropriate. Choose the training method based on the LoRA you want, or combine several separate LoRAs to approach the intended style.

Q. The result is not merely dissimilar—the image becomes distorted when the LoRA is applied.

The difference implied by the pair and tags may have failed to converge and instead diverged. This seems especially likely when trying to learn style from a copyrighted character while keeping the character tag in the caption—for example, training from an illustration and photorealistic conversion of Asuka with 1girl, solo, souryuu asuka langley, ....

From the model's perspective, the situation is something like:

🤖 “Understood—this is how I should draw Asuka from now on. But then how should I draw 1girl in general?”

The process may have extracted a difference that works only under very narrow prompt conditions.

Pairs for style extraction work best when they are simple and require as little conditioning as possible. If you must train from a copyrighted character, consider making the comparison image with image-to-image in the base model—using that model's typical face—instead of asking an LLM to make a photorealistic conversion.

Q. Copita does not open automatically when I launch it.

Some browsers do not open automatically. In the console window, locate http://127.0.0.1:7865 and open it manually. Holding Ctrl while clicking the address usually works.

Q. I get “Python environment is missing” when launching Copita.

First-time setup did not complete correctly. Run setup_first_time.bat before start_copita.bat.

Q. First-time setup failed partway through and Copita says sd-scripts is not set up.

Run setup_first_time.bat again. Setup reuses components that were already completed before the failure. If retrying still does not work, extract a fresh copy of Copita and run setup again from the beginning.

Q. Windows says, “This file was blocked,” when Copita starts.

On a Windows 11 system with Smart App Control enabled, running start_copita.bat may produce an error such as:

ImportError: DLL load failed while importing arrays:This file was blocked by application control policy.

This is not a Copita bug. Windows may not yet recognize the newest release of the internally required pandas component as trusted, so the import is blocked. Replace it with a slightly older stable release:

1. Close the Copita console window if it is open.

2. In File Explorer, open the folder containing start_copita.bat.

3. Click the address bar, erase its contents, type cmd, and press Enter.

4. In the new Command Prompt window, paste the following line and press Enter:

Copita\.venv\Scripts\python.exe -m pip install "pandas==2.2.3"

5. Wait until Successfully installed pandas-2.2.3 appears.

Close Command Prompt and launch start_copita.bat again.

Q. The model I want to train does not appear in Copita.

Check that the Models path is correct. Select the parent folder that contains subfolders such as StableDiffusion, TextEncoders, and VAE—for example, Stability Matrix's data/models folder. Copita can also detect files stored directly in the selected Models folder when those subfolders do not exist.

Q. No LoRAs appear in LoRA Fusion Mode.

In Difference LoRA Training Mode, go to LoRA Name and Save Location and make Copita remember the folder that contains the LoRAs you want to merge. Then return to Fusion Mode and click Refresh LoRA List.

Q. I followed the instructions, but training never completes.

If installation succeeded but training stops midway or fails a preflight check, the GPU may not match the selected settings. Switch to the Stable preset, set batch to 1, lower dim to 8, and enable lowram if necessary.

Q. My Copita screen looks different from the screenshots in this guide.

The screenshots follow v1.2 as closely as possible, but some minor labels or layouts may differ slightly between builds. Use the English labels in the application as the authority when a cosmetic difference appears.

Q. Training takes too long.

Copita's presets aim to keep training practical, but GPU performance imposes a real limit. Fewer steps shorten the run but reduce LoRA accuracy. Adjust steps and dim according to how much fine detail you need and what your GPU can handle.

Q. I do not know why training failed.

Copita saves a text log under Copita\runs\ui_jobs. Showing that file to an LLM may help identify the cause. The Stable preset has been tested with 12 GB of VRAM, but background software and local configuration can still affect stability, so adjust settings for your system.

Q. Can I use Copita if my system-wide default is Python 3.13?

Yes. Copita v1.2 searches for a compatible Python 3.11 or 3.10 installation and, if none exists, prepares a private Python 3.11 environment inside the Copita distribution folder. It does not change your system-wide Python or the environments used by Forge Neo or Stability Matrix.

Closing Notes

That concludes this guide to Copita, the copy-machine LoRA trainer for Anima and SDXL.

Copita is not a magic wand that creates a perfect LoRA in one attempt. But with a carefully made comparison image and clean captions, it can help you build a practical stylistic foundation for your own work. One of its main goals is to help creators carry a style they previously expressed in SDXL over to Anima without assembling a large dataset.

Anima has a much shorter history than SDXL, and it remains to be seen whether it will become an equally dominant ecosystem with a similarly rich LoRA library. I hope Copita makes it easier to move existing creative assets and styles into that new environment.

Thank you for reading!

About the author

I'm Io Sakaki (Studio Masakaki). I write long-form, hands-on research on AI illustration from Japan — nearly 300 investigations so far into LoRA training, prompting, retouching and video models, with every experiment documented, including the failures.

image.png

Civitai: https://civitai.com/user/studiomasakaki — free LoRAs built with Copita, including Anima Bold v1. Follow me here for future English releases.

FANBOX: https://studiomasakaki.fanbox.cc — my main publication. Articles are in Japanese, but machine translation reads well, and this is where new research appears first.

Gumroad: https://iosakaki.gumroad.com/l/copita — Copita for international users. One purchase includes future updates.

6