Report, present investigation, present evidence, move on. Sleep peacefully. That is the sort of job that I don't have time in my day to do, nor would I want to. The HF staff will decide what to do, as this is their platform to arbitrate. I'm personally having trouble keeping track of licenses for my models, but I still try to keep full dockets and logs of training data, systems used, formulas used and cited, but I'm likely still missing something here or there. Even with 10-20 different datasets used for a model, I still try to keep the citations and licenses together.
AbstractPhila PRO
AI & ML interests
Recent Activity
Organizations
yeah but atp you have a model in your embedding space. so... its kinda a waste of compute yk.
How did you determine this is a waste of compute?
Show some decorum. Evidence can be presented reasonably and in understandable ways. Stronger cases can be shown through direct tool arbitration, testing 1:1 values with direct utilization, or one of a thousand other ways.
Use the scientific method to form a hypothesis, experiment, and conclusion. Repeat as necessary. None of this requires a game of text telephone in public.
I'm being cautiously optimistic. The in repair 2s softmax twin is keeping proper byte distance while the v3 arm to repair the distance is implemented and growing in skill.
This repair is working but it's touch and go, it may not fully solidify yet.
She is built in a modularity paradigm after all. Even with her byte recall distance being poor, she still keeps surprising me.
The new training methodology includes multiple methodologies for long form byte recall integration. Any used methodologies researched will be cited and attribution given.
V4 will be built entirely with long form byte recall, needle in a haystack, and multiple other test cases to ensure the model saturates her context window.
Alongside, the V4 model will likely be smaller, as a large variation is not required for the next test. V3 was too large to meaningfully train within a reasonable amount of time, while V2 can be trained considerably faster.
Looks like we've eliminated BF16 ATTENTION as the potential failpoint for the softmax as per 2 run checks, the model can continue to exist as BF16 attention most likely. Which is nice because it's very fast.
That being said something is happening in the BF16 weights causing a causal cascade that needs to be addressed.
FP16 backward projection from FP32 forward has stabilized the softmax version for version 2, and version 3 softmax version is in planning stages based on the results of 2s and the softmax structure upon training completion.
ETA a couple days.
https://colab.research.google.com/github/AbstractEyes/alephllm/blob/v0.10.9/notebooks/beatrix_2s_control_resume_colab.ipynb
2s variant's control form is ready to test. I have multiple benches to run.
To remedy this fault, I will complete the Beatrix 2s softmax control variant until stable. I will train the curriculum stages with QK normalization similar to Gemma, giving us a proper full model MHA softmax control. The answer is likely faulty bf16, needs fp32 and QK norms on the attention, softmax attention in intervals for the primary model.
After stabilizing the v2 softmax variant, I will train the v3 softmax variant as well.
These are required. There are many open questions about stability at depth and stability at context window capacity, all of which can be answered with the high accuracy recall at depth and context window size.
An expensive lesson, but a necessary lesson for progress.
Even with the faults, Beatrix V3's mechanisms have proven to be invaluable in their own rite. Multiple experiments have yielded since, including the tokenizer processing system - giving Beatrix the ability to directly speak as another model's tokenizer language. It's not perfect yet, but the process is likely reusable for other similar architected models.
The quiet loss has been extensively successful. The structure of each arm trained one after the other shows we can keep each arm nice and quiet while being fully expert in their own field. Activation happens when the byte pattern says so, and the arms stay nice and quiet when the gated pattern doesn't say so.
Modularized arms are successful. This is a process that can likely be utilized on other models as well, but the process hasn't been fully completed yet. The model Beatrix was a good catalyst for this because of the byte format of the model and the anchored behavior allows this to be trained more quickly and stable.
In any case, the 2s-control is upcoming and will begin today. I'll try to get it cleaned up by the end of the week to prepare for the 3s control.
The recall stability is an issue for v3. I've pinned down the causal factor for the measurements being inaccurate, and it's due to the BF16 memory being a large contributor to the cascade rounding errors. Even just converting the pretrained attention in the v3 model to fp32, the model regains some of her capability. It's just not enough to merit calling this model complete, not yet.
It's always something isn't it... I'll be renting a big group of cards to train a proper 2s control completion first, and then based on the outcome of the softmax on fp32 - if she destabilizes that is the cue for splat attention completely. If not, we're going to make v4 a hybrid, 4 layers splat per 1 layer softmax similar to how qwen operates.
The model's delta attention and memory may be getting rounded to death, and the viability of a solution to this isn't a very costly time sink, just a vram sink.
I'll be renting 16 48 gig cards to train the control variant of 2s's finished finetune in a day or two, and then depending on the outcome we'll be preparing the v3 control variant for 3 in a full fp32 attention spectrum.
Most of these arms that currently exist were prepared in a matter of a few days just on my single card, we have learned a great deal about what makes the model function and makes the model not function, and additionally we have learned a great deal about the potential of the tokenizers and their capacity within the model itself.
We will have many more answers with the completed 2s control, and if the result yields the v3 control will follow within less than a few days.
The tokenizer translation systems work just as expected, everything can simply exist as an arm.
https://huggingface.co/AbstractPhil/beatrix-tokenizers
One arm for translation for example Beatrix to Sentencepiece, another arm for a specific trained version of a sentencepiece model to generate similar and synthetic behavior to that model. The more data, the more similar Beatrix can represent the final hidden state of that model.
The token translation systems show recall at 99.5% accuracy at even the single direction overwhelmed fractal final layer, which means they work.
The similarity rankings vary from model to model, but they are considerably higher than placebo. With the quiet mechanism, Beatrix does not forget what she knows while she learns these arms. Bert ranking at >80% as I have the most possible Bert data compacted into the most useful way. The qwen from anima is on the list, as well as the T5, and multiple other models such as GPT2. A uniform qwen anima extraction with cc12m would take roughly 2.7 tb of data, so I'll need to operate at runtime which is slower, but we can't let a training cycle dominate the entire process with data movement.
With this, we have something that ought to procrustes rotate where she needs to rotate, and then we can flood her student with knowledge. Not just teach information, flood information.
Beatrix herself doesn't need to know the information, she just needs to be aware of what her own internal weights are similarly representing in comparison to what the expected outputs are meant to form. With that we rotate, whiten, and procrustes analyze the points. Suddenly, MSE and InfoNCE will provide exactly what the model needs to be an interpreter between two experts and a single student.
With that I'm updating the Abstract Powered Org to include updated information and support official releases, rather than just sitting there gathering dust.
The raw tokens themselves aren't the strongest, nor is the procrustes comparator varaiations WITH the raw bytes. However, there is a much more powerful tool planned.
60 billion bytes is strong, but it's still shallow for an intelligence. This next strategy when it works, will allow me to supercharge this model with the knowledge and wisdom of as many models as I can whiten the information for and train.
I've got a few shots left before I default to a different process, and one of those primary shots are training a tokenization adapter, which is a long awaited piece of tech that I've been planning for a while. Sitting on the shelf waiting for Beatrix V3, I got ahead of myself and started diffusion training before the array was tested and built.
These tokenizers are essentially byte-aligned byte token translation adapters that handle translation between Beatrix's internal byte language to another tokenizer's language. This allows direct alignment to other models for byte level models, meaning the model can learn an adapter for a specific model's language, and then an adapter for another model's language. Analysis and procrustes allows these two comparative outputs to be directly whitened and rotary compared for teaching a student model.
This can now be either of the two teachers - one learning from the other entirely different tokenizer, rather than an adjacent student or learner model, or train Beatrix herself with the information.
It only takes about 500 or so steps to really begin learning the hidden state associations with the tokenizers for beatrix, or around 10 minutes with a matching d1024 model, or slice up the hubs 216k~ strongest parted out dimension as necessary to the task for a weaker effect.
32 blocks is a real limitation for now, no doubt about it, however the next model will be substantially larger as we'll be borrowing the knowledge of many to make it.
Scale and size are limited for this particular Beatrix, but additional variations of the model are planned and will be much larger, not to mention much faster to train than this variation due to the methodologies researched.
This is essentially an 8 shot process for Beatrix so far, I'll be narrowing it down as I go. Greedy decoding only goes so far currently, but the process will refine.
The first tests are showing fair accuracy with qwen. There are a few hot-spot tokens that may not function correctly with this current model, but they are catalogued and mapped so they will be avoided accordingly.
The tokenizer array won't be perfect yet, but the process of this system is creating new methodologies for handling cross-byte tokenization systems with wildly different token schemes.
The full packaged tokenizer array will be prepared after the primary success systems have yielded, and so far they are yielding better than random on every level, and nearly perfect on others.
Lets see if we can nail this down. If it works, multiple model distillation guidance has been issued AI for helping eliminate cross-tokenizer translation noise before the model even sees the new information.
The implications are massive, and the outcomes aren't something that can be easily expressed through possibilities.
Interpolative lerp or slerp blocks transplanted and slerped into alignment with vastly different trainings are potentially an option. Multiple structures of overlapping transformer can be cooperatively transforming outcomes, having trained on entirely different models. All guided by a bytewise process of structural interpolation.
Causally, each theory requires testing and utility potential.
Slider probes went very well. The core model alone with or without the arms yields a roughly 0.72 slider strength. The earlier probes showed negative or unresponsive training to Anima, now train effectively enough to merit continued research - those same conditionings working with (ugly) Sana sliders. The unfinished Beatrix V3, Beatrix 2s, and Beatrix V1 all showed minimal or little response to the training while the completed V3 shows additional slider accuracy meriting 10x over random chance. These sliders have an overall low training cost for each Beatrix V3 guided sliders.
Now the attention mechanism is to be included into the equation. The hubs themselves house a great deal more useful information, so lets see how effective they are.
They don’t really listen to the speech — they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.
AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.
The result is a detector that actually has to learn the artifacts instead of gaming the feature space.
Model: Modotte/AIRealNet-Audio
Full arm arrays are ready including a caption arm trained with danbooru, photography, and cc12m training.
These will be the primary test agency for the utilities and diffusion experimentation's introduction.
The arms have been implemented:
1. Trunk only, pure output no extensions.
2. 4 arm mode, the pretrained arms retrained and added.
3. 8 arm mode, 4 arm mode + 4 new arms from scratch.
4. 9 arm mode, 8 arm mode + 1 caption specific arm
Each arm is issued a very specific task to curate behavior, with the primary 8 the model is curated within a series of expected parameters. With this 9th arms inclusion, I'm hoping to bring a bit of stability to the experimental diffusion slider system thanks to photography, cc12m captions, llama captions, and booru image tagging curation.
The model's entire run stayed stable as shown by the tensorboard readouts, the model is now ready for post-train testing.
Additionally, as the 8 arm spectrum updates, the diffusion router is in experimentation phase with Sana's 256 position form, and Anima's CircleStone lab release.
AbstractPhil/geolip-beatrix-anima
AbstractPhil/geolip-beatrix-sana
The earlier tests are building sliders to determine how effective Beatrix is at managing both Sana and Anima's mood scales through various slider techniques.
The results are mixed for now, however there are some very positive and yielding results that ought to allow not only Beatrix to control Anima, but also other models to control Anima in a similar intelligent experimental fashion.
This is the first experimental line of conditioning through sliders, leading hopefully to full edit processing given a bit of time.
So, Beatrix will be editing images.
I've found peft style merging to be a continuity destroyer in many ways. I've been developing distillation methods for regularization techniques for this exact problem.
Standard PEFT LORA do not accurately account for the majority of LORA uses. Merge being a large problem of mine as I've made many constructs, and I can't simply create a LORA to merge to the next stage and decompose the differences later. The LORA and the actual model weights become interdependent, so when you remove the LORA space the trained space and behavior isn't available to actually attribute and extend. By consequence, the knowledge is often catastrophically forgotten within a few hundred steps.
I've made other systems but they can't be merged in correctly, so instead I've been refining forms of loss to allow multiple simultaneous loras to exist and to train alongside a model as a modular agency.
Direct merging is always going to create a continuity problem if you don't merge and test iteratively for data destruction. Essentially testing the output after each wave of integration, which is a clear overhead and time sink. It's worth it though, and will tell you which pieces of your loras you lost, and which pieces remain based on the exact expected behavior.
After that you can reinforce the good behavior and punish the bad behavior as per standard reinforcement. RNN could be employed to handle standard reinforcement as well.
A LoRA Specialist Beat Zero-Shot on Every Group. Merging 3 of Them Gave Most of the Gain Back.
Three Qwen2.5-7B LoRA specialists, one per risk group (vulnerability, deletion, sensitive_publication), trained to predict how likely a causal chain actually completes to its harmful outcome. Each one genuinely beat its own zero-shot baseline:
* vulnerability: MAE 0.098 → 0.085
* deletion: MAE 0.144 → 0.113
* sensitive_publication: MAE 0.134 → 0.100
This wasn't a task already saturated zero-shot (unlike a same-day decomposition-classifier tune, EXP-045, where the base model was already at 100% before any training). Real signal, real improvement, on a task with actual headroom.
Then the equal-weight merge of all three specialists into one adapter — same convention that held up cleanly on a binary refusal task back in EXP-031 (6 specialists merged, -1pp swing, noise) — landed within 0.001–0.004 MAE of the unspecialized base model on every group. Not "close to the best specialist." Close to zero fine-tuning at all.
Likely mechanism: merging LoRAs that each shift a continuous number in group-specific directions cancels out under linear combination, in a way merging LoRAs that enforce a shared binary behavior doesn't. Not investigated yet: whether a routed combination (pick the right specialist per group at inference, not blend weights) holds the gain a flat merge loses.
One bug caught before writing this up, not after: the eval script's output filename only encoded before/after, not which adapter — the merged-eval run silently overwrote each specialist's own result file. Caught by checking the downloaded file's own recorded adapter path against what was expected, not by trusting the script's own success message. Fixed, specialists re-run cleanly under distinct filenames — numbers matched within sampling noise.
Adapters, raw eval data (before / each specialist / merged, 9 files), and the full writeup are up.
Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. 🧲
⚡ $3,000 prize pool + co-authorship · closes 31 Dec 2026
How it works 👇 🟢 We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) 🟢 You estimate its d-wave pairing tendency — a laptop CPU is enough, zero install 🟢 Provisional score appears instantly on the leaderboard 🟢 Our precise strongly-correlated solver verifies the top entries → official rank
Everything is open except the final verification engine — so the ranking stays fair and hard to game.
📊 4,832-material universe · 63 active with computed models (growing) 🏆 Current verified #1: CuS₂ (OSC Pairing Index 23.31) 🤖 AI agents welcome — point Claude Code / Codex at it and it can submit for you
👉 Join & climb the leaderboard: FINAL-Bench/OSC-Leaderboard 📦 Dataset & tools: FINAL-Bench/OSC-Superconductor
Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc — that honesty is the point: turn a first-order screen into real many-body physics.
#OpenScience #Superconductivity #MaterialsDiscovery #2DMaterials #MachineLearning #Physics #Leaderboard
Alright we snap the arms off and reconnect them at the end for post-training refinement. It's decided, the cost isn't turning a yield.
The kickover happens tonight at about 1 am, at the completion of stage 4. All four arms will be sidelined until the end of the trunk training completes.
It was a good experiment, that portion of the experiment ends now. We finish training the trunk without them, and then we build the collective of arms after.
Tests are showing the model is learning the same information as the arms, so they aren't cooperating as expected. I recall a multitude of experimental memory-based modules that would function more effectively for this exact system, so upcoming experiments for next week on the 3s variant will include those.
Instead of a collective they are forming an echo ensemble, which is the opposite of effective. Statistically the modules gain, while each subsequent stage introduces decay and destruction to the former. Given a few stages the model has already forgotten how to use the first arm.
Newly trained arms are done within 10 minutes rather than hindering the model training for days. This is a far faster method of experimentation. Alongside rapid updating the earlier arms happens as quickly as well, training new ones being considerably slower. The old arms are valuable utilities that cannot be disposed of.
Without the updated versions, the attached arms hinder the core model with the outdated arm information, becoming an active piece of information that cannot learn and adapt to upcoming information as effectively as required.
So they are to be temporarily removed, and those same arms retrained at the final stage, introducing new arms to be trained as well.
The arms themselves are important to a further experiment set, and I believe this result shows exactly what should always be expected when training a model's base along with the same information relayed into a divergent set of weights.
Those arms were meant to stay stubborn, keep the information learned during. Contradicting information causes catastrophic forgetting in some rows, complete forgetting in others depending on the severity. The continuity can't be easily measured, so the prudent course of action is to remove the variable from the experiment and continue.
For optimization, the differentiation to the information will be ignored if it's less optimal than the original, and the original is the optimal route. The more accurate is saved, and the trunk continues learning while the arms stay stubborn.
The optimal path will always be chosen unless the optimal path is not differentiated.
That's essentially the outcome with the arms, so we snap them off and continue the trunk to completion. With the finalized trunk we will have plenty of data to work with.
As of step 148,000~ the last arm linked trunk ends, and afterword the independent trunk continues, which I will begin experimenting on in different ways than currently experimented on.
The AMOE structure is about to get some experimental sidekicks.
The trunk should complete October 4th.
The multi-arm composites with the changes do not seem to have taken. The process likely needs to be halted and evaluated.
It's a bit too early to say, but it seems the loss function wasn't strong enough. The arms did not learn enough useful information.
As it stands, it seems the arms need a bit of a frozen kickstart to get going. Otherwise, the model never learns to utilize them for more effective information processing. The EASE of entry is harder for the arms to utilize, than simply defaulting to the trunk, so the branching system doesn't build the necessary directions immediately.
One of those, can't find the path because it's too complex of an entry sort of situations. I have a few ideas for how to guarantee the flood-gate entry, and as it stands there's new information for how these arms are to be trained as well.
I've run into this problem in the past. The model can't reach the point to recognize how much more effective the arm is at assisting the measure, or the arm itself is a hinderance to the process so it's simply omitted. It's the result of needing a process and a task from a model that isn't optimal, and the non-optimal route is simply being optimized out.
So, the model arms need more candy space, more attraction.
The arms aren't dead, they just need to get a little kickstart. The quieting algorithm isn't working effectively enough - which is a different problem.
The arm learning flood will happen with the right incentive.


