Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
AbstractPhil 
posted an update 6 days ago
Post
120
Mini-Beatrix-3 AbstractPhil/mini-beatrix-3 is released into automodel testing phase.

The arms have been implemented:
1. Trunk only, pure output no extensions.
2. 4 arm mode, the pretrained arms retrained and added.
3. 8 arm mode, 4 arm mode + 4 new arms from scratch.
4. 9 arm mode, 8 arm mode + 1 caption specific arm

Each arm is issued a very specific task to curate behavior, with the primary 8 the model is curated within a series of expected parameters. With this 9th arms inclusion, I'm hoping to bring a bit of stability to the experimental diffusion slider system thanks to photography, cc12m captions, llama captions, and booru image tagging curation.

The model's entire run stayed stable as shown by the tensorboard readouts, the model is now ready for post-train testing.

Additionally, as the 8 arm spectrum updates, the diffusion router is in experimentation phase with Sana's 256 position form, and Anima's CircleStone lab release.

AbstractPhil/geolip-beatrix-anima
AbstractPhil/geolip-beatrix-sana

The earlier tests are building sliders to determine how effective Beatrix is at managing both Sana and Anima's mood scales through various slider techniques.

The results are mixed for now, however there are some very positive and yielding results that ought to allow not only Beatrix to control Anima, but also other models to control Anima in a similar intelligent experimental fashion.

This is the first experimental line of conditioning through sliders, leading hopefully to full edit processing given a bit of time.

So, Beatrix will be editing images.

Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.

These will be the primary test agency for the utilities and diffusion experimentation's introduction.

Slider probes went very well. The core model alone with or without the arms yields a roughly 0.72 slider strength. The earlier probes showed negative or unresponsive training to Anima, now train effectively enough to merit continued research - those same conditionings working with (ugly) Sana sliders. The unfinished Beatrix V3, Beatrix 2s, and Beatrix V1 all showed minimal or little response to the training while the completed V3 shows additional slider accuracy meriting 10x over random chance. These sliders have an overall low training cost for each Beatrix V3 guided sliders.

Now the attention mechanism is to be included into the equation. The hubs themselves house a great deal more useful information, so lets see how effective they are.

The raw tokens themselves aren't the strongest, nor is the procrustes comparator varaiations WITH the raw bytes. However, there is a much more powerful tool planned.

60 billion bytes is strong, but it's still shallow for an intelligence. This next strategy when it works, will allow me to supercharge this model with the knowledge and wisdom of as many models as I can whiten the information for and train.

I've got a few shots left before I default to a different process, and one of those primary shots are training a tokenization adapter, which is a long awaited piece of tech that I've been planning for a while. Sitting on the shelf waiting for Beatrix V3, I got ahead of myself and started diffusion training before the array was tested and built.

These tokenizers are essentially byte-aligned byte token translation adapters that handle translation between Beatrix's internal byte language to another tokenizer's language. This allows direct alignment to other models for byte level models, meaning the model can learn an adapter for a specific model's language, and then an adapter for another model's language. Analysis and procrustes allows these two comparative outputs to be directly whitened and rotary compared for teaching a student model.

This can now be either of the two teachers - one learning from the other entirely different tokenizer, rather than an adjacent student or learner model, or train Beatrix herself with the information.

It only takes about 500 or so steps to really begin learning the hidden state associations with the tokenizers for beatrix, or around 10 minutes with a matching d1024 model, or slice up the hubs 216k~ strongest parted out dimension as necessary to the task for a weaker effect.

32 blocks is a real limitation for now, no doubt about it, however the next model will be substantially larger as we'll be borrowing the knowledge of many to make it.

Scale and size are limited for this particular Beatrix, but additional variations of the model are planned and will be much larger, not to mention much faster to train than this variation due to the methodologies researched.

This is essentially an 8 shot process for Beatrix so far, I'll be narrowing it down as I go. Greedy decoding only goes so far currently, but the process will refine.

The first tests are showing fair accuracy with qwen. There are a few hot-spot tokens that may not function correctly with this current model, but they are catalogued and mapped so they will be avoided accordingly.

The tokenizer array won't be perfect yet, but the process of this system is creating new methodologies for handling cross-byte tokenization systems with wildly different token schemes.

The full packaged tokenizer array will be prepared after the primary success systems have yielded, and so far they are yielding better than random on every level, and nearly perfect on others.

Lets see if we can nail this down. If it works, multiple model distillation guidance has been issued AI for helping eliminate cross-tokenizer translation noise before the model even sees the new information.

The implications are massive, and the outcomes aren't something that can be easily expressed through possibilities.

Interpolative lerp or slerp blocks transplanted and slerped into alignment with vastly different trainings are potentially an option. Multiple structures of overlapping transformer can be cooperatively transforming outcomes, having trained on entirely different models. All guided by a bytewise process of structural interpolation.

Causally, each theory requires testing and utility potential.

·

The tokenizer translation systems work just as expected, everything can simply exist as an arm.

https://huggingface.co/AbstractPhil/beatrix-tokenizers

One arm for translation for example Beatrix to Sentencepiece, another arm for a specific trained version of a sentencepiece model to generate similar and synthetic behavior to that model. The more data, the more similar Beatrix can represent the final hidden state of that model.

The token translation systems show recall at 99.5% accuracy at even the single direction overwhelmed fractal final layer, which means they work.

The similarity rankings vary from model to model, but they are considerably higher than placebo. With the quiet mechanism, Beatrix does not forget what she knows while she learns these arms. Bert ranking at >80% as I have the most possible Bert data compacted into the most useful way. The qwen from anima is on the list, as well as the T5, and multiple other models such as GPT2. A uniform qwen anima extraction with cc12m would take roughly 2.7 tb of data, so I'll need to operate at runtime which is slower, but we can't let a training cycle dominate the entire process with data movement.

With this, we have something that ought to procrustes rotate where she needs to rotate, and then we can flood her student with knowledge. Not just teach information, flood information.

Beatrix herself doesn't need to know the information, she just needs to be aware of what her own internal weights are similarly representing in comparison to what the expected outputs are meant to form. With that we rotate, whiten, and procrustes analyze the points. Suddenly, MSE and InfoNCE will provide exactly what the model needs to be an interpreter between two experts and a single student.

With that I'm updating the Abstract Powered Org to include updated information and support official releases, rather than just sitting there gathering dust.

The recall stability is an issue for v3. I've pinned down the causal factor for the measurements being inaccurate, and it's due to the BF16 memory being a large contributor to the cascade rounding errors. Even just converting the pretrained attention in the v3 model to fp32, the model regains some of her capability. It's just not enough to merit calling this model complete, not yet.

It's always something isn't it... I'll be renting a big group of cards to train a proper 2s control completion first, and then based on the outcome of the softmax on fp32 - if she destabilizes that is the cue for splat attention completely. If not, we're going to make v4 a hybrid, 4 layers splat per 1 layer softmax similar to how qwen operates.

The model's delta attention and memory may be getting rounded to death, and the viability of a solution to this isn't a very costly time sink, just a vram sink.

I'll be renting 16 48 gig cards to train the control variant of 2s's finished finetune in a day or two, and then depending on the outcome we'll be preparing the v3 control variant for 3 in a full fp32 attention spectrum.

Most of these arms that currently exist were prepared in a matter of a few days just on my single card, we have learned a great deal about what makes the model function and makes the model not function, and additionally we have learned a great deal about the potential of the tokenizers and their capacity within the model itself.

We will have many more answers with the completed 2s control, and if the result yields the v3 control will follow within less than a few days.

In this post