Sol Milkshake

Sol Milkshake

Milkshake tests recurrence and hyperspherical optimization in MLX on Apple Silicon. Five stored transformer blocks produce eleven block applications. Local token-pattern memory and rolling chunk memory accompany the shared blocks.

The model has 2,990,000 parameters. Its released checkpoint received 2,526,565,888 token exposures across pretraining and recovery. It generates text completions and hasn't been instruction-tuned.

Configuration

Setting Value
Parameters 2,990,000
Token exposures 2,526,565,888
Context 2,048 tokens
Tokenizer 2,048-entry byte-level BPE
Hidden width 192
Stored blocks / effective applications 5 / 11
Recurrent layout 1 prelude, 3 middle blocks used three times, 1 coda
Attention 6 query heads, 2 KV heads, head dimension 32
Q/K normalization Unit normalization with RoPE
Attention features XSA value subtraction and value residuals
Routing Full first pass, then 75% and 50% token capacity
FFN Gated SiLU, width 512
Local memory Rank-26 tensorized 2-5-gram memory
Rolling memory 32-token chunks, 32 slots, width 64
Token embeddings Tied to the output head
Weights MLX NPZ

Run with MLX

Install MLX, the tokenizer library, and the Hugging Face Hub client:

pip install "mlx>=0.29" "tokenizers>=0.22" huggingface_hub

The downloaded repository contains the model implementation, loader, and generation function:

from huggingface_hub import snapshot_download
import sys

model_dir = snapshot_download("solintellegence/Sol-Milkshake")
sys.path.insert(0, model_dir)

from modeling_sol_milkshake import load_model, generate

model, tokenizer = load_model(model_dir)

text = generate(
    model,
    tokenizer,
    "The future of efficient language models is",
    max_new_tokens=64,
)

print(text)

Evaluation and selection

Benchmark Examples Normalized accuracy
HellaSwag 10,042 25.02%
ARC-Easy 2,376 29.76%
ARC-Challenge 1,172 22.70%
PIQA 1,838 52.50%
ArithMark-3 1,000 30.60%

Milkshake's local Intelligence Index score is 3.158. The four language-model tasks used their complete zero-shot splits and lm-eval 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. See evals/ for the raw outputs.

We used this same benchmark suite to select the recovery checkpoint, so selection affects the reported scores. They aren't an untouched held-out estimate, and no independent evaluation has verified them.

Pretraining and recovery

Pretraining drew on FineWeb-Edu, FinePDFs-Edu, English UltraFineWeb multi-domain and question-answer subsets, Cosmopedia, and FineMath. For recovery, the mixture was 70% original frozen curriculum, 15% Cosmopedia-v2 English, and 15% FinePhrase.

Recovery kept the tokenizer and prepared streams fixed. Filtering, deduplication, and decontamination rules also stayed the same. The run record is training_state.json.

Repository contents

Load the MLX weights from model.npz with modeling_sol_milkshake.py. That file also provides generation; sol_config.py defines the architecture. The download includes config.json, tokenizer files, training metadata, and evaluation outputs.

Milkshake is intended for research on recurrence, memory, hyperspherical optimization, and MLX inference. Completions can be repetitive, incoherent, or incorrect.

The model license is CC BY 4.0. Dataset licenses apply separately.

Downloads last month
559
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including solintellegence/Sol-Milkshake