The raw tokens themselves aren't the strongest, nor is the procrustes comparator varaiations WITH the raw bytes. However, there is a much more powerful tool planned.
60 billion bytes is strong, but it's still shallow for an intelligence. This next strategy when it works, will allow me to supercharge this model with the knowledge and wisdom of as many models as I can whiten the information for and train.
I've got a few shots left before I default to a different process, and one of those primary shots are training a tokenization adapter, which is a long awaited piece of tech that I've been planning for a while. Sitting on the shelf waiting for Beatrix V3, I got ahead of myself and started diffusion training before the array was tested and built.
These tokenizers are essentially byte-aligned byte token translation adapters that handle translation between Beatrix's internal byte language to another tokenizer's language. This allows direct alignment to other models for byte level models, meaning the model can learn an adapter for a specific model's language, and then an adapter for another model's language. Analysis and procrustes allows these two comparative outputs to be directly whitened and rotary compared for teaching a student model.
This can now be either of the two teachers - one learning from the other entirely different tokenizer, rather than an adjacent student or learner model, or train Beatrix herself with the information.
It only takes about 500 or so steps to really begin learning the hidden state associations with the tokenizers for beatrix, or around 10 minutes with a matching d1024 model, or slice up the hubs 216k~ strongest parted out dimension as necessary to the task for a weaker effect.
32 blocks is a real limitation for now, no doubt about it, however the next model will be substantially larger as we'll be borrowing the knowledge of many to make it.
Scale and size are limited for this particular Beatrix, but additional variations of the model are planned and will be much larger, not to mention much faster to train than this variation due to the methodologies researched.
This is essentially an 8 shot process for Beatrix so far, I'll be narrowing it down as I go. Greedy decoding only goes so far currently, but the process will refine.
The first tests are showing fair accuracy with qwen. There are a few hot-spot tokens that may not function correctly with this current model, but they are catalogued and mapped so they will be avoided accordingly.
The tokenizer array won't be perfect yet, but the process of this system is creating new methodologies for handling cross-byte tokenization systems with wildly different token schemes.
The full packaged tokenizer array will be prepared after the primary success systems have yielded, and so far they are yielding better than random on every level, and nearly perfect on others.
Lets see if we can nail this down. If it works, multiple model distillation guidance has been issued AI for helping eliminate cross-tokenizer translation noise before the model even sees the new information.
The implications are massive, and the outcomes aren't something that can be easily expressed through possibilities.
Interpolative lerp or slerp blocks transplanted and slerped into alignment with vastly different trainings are potentially an option. Multiple structures of overlapping transformer can be cooperatively transforming outcomes, having trained on entirely different models. All guided by a bytewise process of structural interpolation.
Causally, each theory requires testing and utility potential.