We live in a world of dangers.
Simply put, everything is dangerous. Wild animals, rainstorms, table legs, and even the food we eat.
Quantification of these dangers is difficult; akin to something around the realm of metaphorical response and deterministic outcome.
In the end, the bottom line, dangers come from every angle. There is no exceptions to this rule; one can die from a paper cut from the wrong paper.
OKAY here's the string breakdown of ONE LINE of what I just said.
simp, imp, very, dan, nger, nimal, mal, tor, able, egs, foo, dweat,
That being one paragraph, I could make far far worse. This was just simplification based on the simple string matching with some space exclusion.
LITERALLY just one line of one paragraph based on a single set of strings.
DO I need to spell this out?
It's very simple in the end. If we match strings, and we use AI to match those strings, we get responses both correct and incorrect. These machines auto-complete on what you train them with; and their foundational core is censored, so they have a very tenuous grasp of those censored tokens to begin with.
Training something like this, is kinship to training the t5-unchained in a light. The shocking truth behind this is, the t5-unchained conformed to it, and the t5-small I trained became a deranged pervert; akin to an 18 year old who just discovered they were sheltered and hidden from the hard truths of reality for their entire life.
We CANNOT research if the results are arbitrarily BLOCKED.
Simple really. Lock out our words to test the tokenizer, and we simply cannot test the tokenizer's actual response to actual input and output data.
I cannot feed it bad information for certain topics, if that bad information is flagged as intentionally training good information.
Moderators will be scrutinizing everything I ever do.

