Sign In

Anima Generation Guide - Part 1 - General (Model, a bit of history)

4

Jul 31, 2026

(Updated: 3 days ago)

generation guide

00008-3997414095.png

Hello, fellow latent explorers!

It was rather interesting to see that Anima model is so underrated currently. I was asked couple of times how do I make my images, so I did a quick search for guides on it and.. well... the situation is bad. Most guides are either lacking or straight on LLM hallucinations (they usually stumble on pony tags baked in it). So I decided to make my own writeup, explaining obvious stuff and providing my tips and trick to make your images look good while pointing out some common misconceptions. I am assuming that if you are interested in this model - you're a 2D enthusiast and will be able to read between the lines (since I'll keep this guide PG-13).

I will make it a series of articles, because it:

  1. It takes time, and I don't have a ton of it available.

  2. It is published exclusively on Civitai since it is the platform of my choice for finding free stuff (except for yesterday's update, we'll see how it goes), and articles here tend to fail in case a lot of images are attached.

All the parts will be stored in a collection here.

A bit of history

Common misconceptions stem from directly comparing older models, so I have to provide some context in case reader has no idea how anime got its own subniche in image generation.

First real breakthrough in local image generation happened with the "release" (search more on that yourself) of Stable Diffusion 1.5 model by Stability AI. This was THE NAME for a long time, so you can stumble on it nowadays here and there with everything related to local generation. At certain point 2D based finetunes of it started emerging, but they were mostly limited both by original model's capability and by lack of really extensive compute put in. There were sophisticated finetunes (Hassaku (SD1.5) was my favorite), but I will not delve deeper in there.

Later Stability AI released their second praised by community model SDXL (Stable Diffusion XL). Some 2D finetunes of it existed, but all was forgotten after Pony Diffusion V6 XL came out. This model featured a rather sophisticated base of the dataset, but ultimately provided widest variety at the time, since author included some danbooru in it, and basically broke one of the text encoders of the model (SDXL used CLIP-L and CLIP-G) to force "understanding" of danbooru tags to the model. Effect was great and really appreciated by community. Basically danbooru tag worked really well with CLIP's token based encoder due to similar separators and nature, enabling unprecedented control at the time. This model had ups and downs, but it gave us a lot of good finetunes, loras, and horrendous so called "pony score" line that is plaguing generations and continuing to pollute LLM training data to this day (basically training went wrong so you HAD to place

score_9, score_8_up, score_7_up, score_6_up, score_5_up, score_4_up,

in front of any prompt to get anything remotely good).

Separate similar project but based on only anime dataset was cooking at the same time: Kohaku-XL It was left at the state it is in, because local resources were just not fitting the scale, even for SDXL, a rather small model by modern standards. It was later picked up by another author and Illustrious-XL was released. This is a completely different branch of SDXL finetunes. Version 0.1 was basically a barely usable pretrain. But community picked it up and did wonders to it. Technically it had a different text encoder part broken (compared to pony) and more focused anime on 2D art.

A lot of good finetunes were made and another big project that I consider worth noting came from it - NOOB-AI. That set of models was finetune over Illustrious 0.1. It had it's own quirks, but I consider that the best model of the time.

Later Illustrious-XL 1.0 was released, so now full need for some sort of branching genealogy tree to figure out what was finetuned on what, merged by whom and when etc.

But basically all those models are technically a finetuned SDXL with all loras working to some degree between them.

Those are the biggest projects. There were other ones, but I will not cover them here.

Common naming conception between some models is:

  • Base model - either a full pretrain on new architecture or a big finetune (dataset of millions of images)

  • Finetune - further training of the base model on smaller handpicked dataset (usually aesthetically tuning defaults of the model)

  • Merge - block merge between different models and/or addition of some loras (check how loras work somewhere else).

After that the scene was rather stale. New models with better architectures were released, but compute requirements went through the roof. There was Flux1d, SD3.5, Flux2, Qwen image etc. The problem is that only big corpos could afford it, and they never were interested in a sophisticated 2D arts finetune, since it is a rather niche stuff. Another issue is that all the released models were heavily finetuned already (in the way of being placed on the rails). This often prevented any further finetune of the model, meaning that it had to "break" and reconverge at certain point that increased compute requirements even further.

This made community mostly focus on loras that could not fundamentally change the model. The only big project worth checking out here is Chroma.

So here we were, at a rather stale state of an endless stream of SDXL finetunes like Chenkin-noob-XL, Illustrious 3, 4, whatever etc.

Until Anima-preview dropped.

That is why in a lot of parts of this guide I'll focus on comparing it with "generalized SDXL derivative", since it is from were I came and bulk of the readers will come from. Also there are fundamentally not that much difference between them. They are all inept in comprehending sentences, have limited vocabulary, cannot do complex text, require extensive use of extensions and controlnets to push the model etc. You will never be able to even control your composition like this with those models without specific tooling or a lot of rerolling:

image.png

What is Anima?

All the official info is on the model card:

https://huggingface.co/circlestone-labs/Anima

You SHOULD read it yourself. Basically most of the info is right there.

What is not there? Mostly technical stuff, so I'll interpret it for you.

Anima is a deep full finetune of Nvidia Cosmos-Predict2-2B-Text2Image but with a trick in the sleeve. The huge (9GB) T5 text encoder is replaced by Qwen 3 0.6b. Since you cannot simply replace that without ruining a model - special adapter layer is baked into the model. It adapts encoder output of Qwen llm to match existing T5 format and semantics. This is not unheard of, similar stuff was made with SDXL, replacing a CLIP with LLM (Check ELLA, SDXL-T5, Rouwei-Gemma project). Funnily enough there was a reverse project, adding CLIP to a model that does not have it. Despite being really small - Qwen 3 0.6b does a really decent job here, and fits way better into RAM.

Besides that - model is finetuned on ridiculous amount of 2D images (mostly from danbooru database, but with some additions, it has pros and cons, more on that in prompting guide).

Currently folders have few versions. Basics are explained in the model card (direct quote):

Versions

  • Anima-Base

    • The pretrained, unrefined base model. Maximum flexibility, diversity, and style adherence.

    • LoRAs should be trained using this version.

  • Anima-Aesthetic

    • Fine-tuned for better consistency and a higher quality default art style.

    • v1.0b is an alternate version that is just an aesthetics full finetune, without additional style adjustment and stabilization loras merged in like 1.0. I personally think 1.0 is better.

  • Anima-Turbo

    • Distilled version for fast generations.

    • Use at CFG 1 and 8-12 steps.

    • The distillation process also increases stability and gives the model a strong default style, but reduces diversity.

What is not explained is that in the folder there are also preview versions. They are a history already, do not waste time on them.

With this guide series I'll focus on Base version. Why? This is the most diverse version, and if you will be able to tame this beast - working with any further finetune will be just like removing couple of checkboxes from your todo list. All the images used in this guide series will be generated either with only base model, or with the use of my own loras. Btw check them out, they are good.

The mindset

If you are coming from your favorite whatever finetune of previous architectures - this is a base model. It is deliberately unrefined, so that community could train on it. One of the most common complaints I see - I do not like this model's style - is completely irrelevant since Anima-base is not about style, it is supposed to be as diverse as possible, for you to be able to steer effortlessly it in any direction with further lora training. And it achieves those things. Yet there are fundamental things worth talking about and I'll cover them in next articles.

Embrace that it is a new architecture, it was trained differently and on a completely different base. Start learning from scratch. Some stuff just works differently. Some new things are possible. And some stuff is lost unfortunately.

This was supposed to be a short prologue, but I got carried away.

Part 2 is planned have UI setup that I use and basic T2I parameters.

4