David is
A system of AIs. Roughly 8400 different micro-sized AIs all with their own opinions.
This is not a fixed size. David can be added to, have his layers frozen, new elements added, elements removed, and everything will still work. Nothing is dependent on any other structure; and yet they are all cooperative. A systemic harmony.
David has no token limit, because David does not handle traditional token representations. David consumes a representation that can be any token count as long as the representation is valid in the spectrum of the expert teaching it.
The evolution of this model will have a full rope structure with a full classification gating system for words along with representative meaning; which is inherently different than VIT-CLIP out of pure need and not out of choice. VIT models discard geometry given enough epochs and this makes them noncompliant with the goal. Bert has more kinship.
Here are some of the formulas responsible for making David work
https://github.com/AbstractEyes/lattice_vocabulary/blob/master/symbolic_lattice_formula_map.md
David has four feature sizes; 512, 768, 1024, and 1280.
clip_base_laion_patch32
clip_large_patch14
david's internal representation
clip_bigG
This repo you can see an earlier variant of David and the utilization with Imagenet.
https://huggingface.co/AbstractPhil/rose-geoclip
Here is the vocabulary David uses, and I've come up with a very elegant and fast way to extract from it for use that I will be releasing very soon with David.
https://huggingface.co/datasets/AbstractPhil/geometric-vocab
David was finetuned on these features here;
https://huggingface.co/datasets/AbstractPhil/imagenet-clip-features-orderly
Which were originally hosted here;
https://huggingface.co/datasets/AbstractPhil/imagenet-clip-features
I prepared the features ahead of time using the images and the exact words meant for imagenet use, giving distinct shared space representations that represent weak lexicality with clip and fair complexity with vit.
David learns from these features as a student. David doesn't care what patch is used, what sort of behavior the models represent, or what the prompt was to summon something. David simply sees a feature and learns the patterns. These patterns are often not related to patches at all shown by the statistics. They are more representative of bloating valuations and grouped heatmaps, which is not what I expected but it's not unwelcome.
David sees a geometric crystal representing something, it could be anything, and then David learns whatever that representation is in whatever complexity it is. David does not care of your simple systems, nor your tokenizers, or any of that. David uses the geometric-vocab.
David Sees
Clip features. You feed David clip features, a dimensional directive, and a size, and the output will be warped.
David learned ALL of the shapes for all of them, so you can say... only want CLIP_L - you can gate the others. You don't need anything else, so you only need to grab the projection of David's space to the clip_l dimensional expectation - and this is trained with final layer features not clip_skip 2. David does not have layers, so turning off those layers simply doesn't exist. Skipping layers is not a thing. David however does have many knobs for tuning, and those are trained specifically in. Baked in the core, actually.
David Generalizes
David does not care what your feature has. David simply extracts what David learned based on the input feature or features, and then the comparator subsystem will warp the output to an encoder shaped system with classification bias.
In other words, your output feature is perfect for diffusion use - since you can pass a fully trained system of geometric-vocabulary to David - and David can understand roughly 33,000 of them by design.
The output feature will pass classification accuracy tests; likely in the upper route of 80+% accuracy as shown by image-net and cifar100. Cifar100 David breached 94%.
David Doesn't Just Warp
David learns patterns - highly deterministic patterns that fit within specific subspaces of reason - entirely compartmentalized to symbolism utilizing; anchor, need, relation, target, and observer paradigm - which is a canonical form I use.
David collapses the meaning of a 5 dimensional spectrum of crystals from [batch, 5, 1024] into a meaningful series of [5, 512], [5, 768], [5, 1024], [5, 1280] which can then be collapsed further or extracted as singular forms via the extract_features function built clean into David.
David is the natural extension of a geometric represented VIT-CLIP housed entirely in a highly condensed and curated geometric-representative space, trained as a student from the very masters that exist. Chaos and noise are specifically curated, even encouraged for certain toggles and allowed for training. This allows for additional utilization and strengthening in terms of classification capability and deterministic outcome.
David Benefits From
Additional training and a larger series of geometric potentials.
Proper gating and respect to the teacher/student hierarchy for training - simply transformers CE on this will never work. This is a cross-contrast variant that requires a very specific form of curated losses and deterministic elemental control that will be released with David's paper.
At least one feature. David is trained with many but requires a feature from an expert to warp the feature through classification analytical warping - pulverizing through rose loss alucard walking is canonically what I called this process in diffusion. You can find the repo for this here;
The geometric-vocabulary that comes with David. It will auto-launch based on the loader script that loads the packaged vocabulary with David.
A tokenizer that fuzzy matches or directly matches expected tokens; or manually selecting valid tokens. This process here isn't done; David does not see tokens like any other model. Fabricating tokens for runtime wouldn't work, so only similar tokens to expectation would be allowed. Meaning David quite frankly may not have the information you need yet. This will improve over time and with a vocabulary curation, classification selection, model expansion, and including additional training.
David Does Not Need
A full retrain to expand David's capability. Each internal structure is independent of the core structure; and thus everything can simply be frozen while new pieces are added - and those new pieces will form just like they are supposed to, independent of the core.
Handholding. David will work without a classifier nor a geometric vocab token - just by giving David a prepared feature. David DOES NOT NEED TEXT. David was trained as a classifier AND feature structure warping system.
Giving David JUST a feature will allow you to classify that feature with a fairly high degree of accuracy based on one of David's clusters. This will allow that particular system to cut through and be identified; potentially incorrectly as David is not perfect, but it will alter the feature based on that valuation or simply yield the classification.


