Likely metadata blackout.
Everyone is going to simply shove their models, their data, and their prompts; through a metadata blackout system and watermark remover - something that will just burn the "ai watermark" from the images, all while simultaneously introducing a false watermark to imply something is ai trained data but really isn't. The task is what 30 lines of python with pillow, and 5 lines with pytorch.
So, essentially all those ais you trained to spot and check, all those little tools you built to extract and test and compare metadata, all those simple string concatenators and stringification tokenizers, NSFW checklists on all that; are going to be left only to those who are unaware - since it's pretty obvious we can just remove metadata and your system will only flag once in a blue moon. The computation required for an accurate classifier is steep; and the amount of images flooding here, is high.
People will just purge the metadata, and an entirely NEW problem will appear. No metadata, hard to find useful tools, and very difficult to work with HIDDEN subsystems in this environment.
None of which will matter for anything except AI training analysis for future AI bots; since no matter what we see, it's going to need tagging. Most likely because it's going to just have it's metadata expunged before it's uploaded, so scraping and training classifiers or detectors, is just going to get that much harder in the future.
You cannot simply get good information with untagged classifiers. They are too generic, and require too many training steps, with too many specific layers dedicated to classification; and those classifiers only fit within certain parameters, so they cannot be reused for other tasks.
No data for new models!
Simply put, all of our classifiers are going to get worse over time, and the ai generated models are going to get more corrupted and more difficult to assess AI or not based on the quality and output from model to model.
You cannot classify a new model, if that new models data is misclassified and misinterpreted into a new system based on old taggers and old classifiers assuming old systems.
Simply put, the accuracy will drop substantially, and the future data that is CURRENTLY FLOODING OUT, will simply start to dwindle over time; making it much harder to detect AI generation in the future, and much easier to dodge your bots and censors or whatever you devise for the system.
Especially when the competition decides they won't actually implement something of this nature.
What this really means.
You won't have the data you need, to scan the images for AI detection, since you have banned the ability to actually scan for them in newer models.
Ergo; shooting yourselves in the foot.
If you think this is enough data for classifiers - you're sorely mistaken.
CLIP on it's best of days - is nowhere near 100% accurate. Let alone the diffusion models based on it. They are comically incorrect more often than not.
We need MORE IMAGES with GOOD TAGGING AND DATA. Laion 5b wasn't even enough to scratch the surface of the pixel potential in that environment, let alone actually provide a useful utility.
WE CANNOT classify everything with an error chance. It's going to fail given enough time. Cascade failure, not simple failure. We're talking impossible to improve due to the core simply having foundational flaws grades of failure.
We need SO MUCH MORE DATA it's not even funny. Maybe even trillions of images to get a full actual depictive environment within that little 1024x1024 zone.
We built additional layers, systems, subsystems, analysis mechanisms, and considerably more; but the outcome stands, that we will simply run out of data before we can make anything truly substantial beyond the topical.


