Sign In

Stay Ready: Challenge Judging Under Test Pressure

0

Stay Ready: Challenge Judging Under Test Pressure

If you take part in Civitai Challenges regularly, you probably know the feeling: you put real time into an entry, try to hit the theme, polish the details — and then the judging result can sometimes feel strangely one-sided, inconsistent or simply hard to understand.

I’ve had that feeling myself and I know I’m not the only one.

Sometimes several very different images end up with almost the same score. Sometimes an entry that seems to fit the Challenge perfectly gets pushed down unexpectedly. And in some cases the results can feel so odd that it is difficult to tell whether we are looking at an intentional judging decision or simply a system that is reaching its limits.

That is why the current development work in Civitai’s public GitHub repository is interesting.

Nothing here has been officially announced or confirmed as an upcoming feature. But judging by the current commits and testing around Challenges, it looks like Civitai is experimenting with a different way of ranking entries.

The current individual judging would apparently stay. Images would still be reviewed on their own, receive scores and feedback and still need to pass the theme checks.

The bigger change seems to happen after that.

Instead of relying mainly on absolute scores like 8.8, 9.0 or 9.1 to decide which images move forward, Civitai is testing systems that compare strong entries directly against each other.

One of those experiments is called Rolling Swiss.

The basic idea is actually pretty simple: groups of four Challenge images are compared with each other and the judge ranks them from strongest to weakest. Over time, strong images increasingly get compared with other strong images.

So instead of asking only:

“How good is this image from 0 to 10?”

the system can also ask:

“Which of these images actually fits this Challenge better?”

That could help with one of the most frustrating parts of the current system: tied or nearly identical scores. The development notes even mention a test where ten entries received exactly 9.00, which makes it extremely difficult to create a meaningful ranking from those scores alone.

The direction currently being tested looks roughly like this:

Individual judging → theme check → comparisons between entries → ranking → finalists → final winner selection

For Rolling Swiss, the existing final AI judge would still choose the winners from the strongest finalists. So this is not currently a complete replacement of the judging system. It looks more like a new layer designed to make the road to the finals less dependent on tiny score differences and random tie-breaking.

There is also a more aggressive Pairwise Ladder approach being tested, where direct head-to-head comparisons can eventually decide the podium itself.

For those of us who enjoy the Challenges, I think the important part is not whether one exact system wins these tests.

It is that the judging system is clearly being looked at.

Because right now, a lot of participants feel a little left alone with results that can be difficult to understand. You submit, you get a score and often there is very little visibility into why one image ended up above another. When that happens repeatedly, it can start to feel less like competition and more like rolling dice against a black box.

These experiments would not magically make judging perfect. They also do not currently solve every issue — especially the fact that large Daily Challenges may still only have a portion of all submitted entries actually judged.

But comparing good entries against each other, instead of trusting decimal scores alone, feels like a meaningful direction.

So for now: no announcement, no promise and no guarantee that this exact system will ship.

But there is definitely movement in the code.

And as someone who actually likes taking part in these Challenges, I’m glad to see the judging itself being put under some pressure.

Stay ready. Challenge judging may finally be getting the attention many of us have been asking for.

Current Development References

Everything described above is based on public development work in Civitai’s GitHub repository. These links are not official confirmation that the feature will ship in this exact form — they simply show what is currently being built and tested.

I would finish the article with one short line underneath:

Again: this is ongoing public development work, not an official Civitai announcement. The final system may change significantly before — or if — it reaches production.

-- Moonbear 🐻

0