Instructions to use maximorulli/nanochat-adamW with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use maximorulli/nanochat-adamW with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="maximorulli/nanochat-adamW", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("maximorulli/nanochat-adamW", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("maximorulli/nanochat-adamW", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use maximorulli/nanochat-adamW with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "maximorulli/nanochat-adamW" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maximorulli/nanochat-adamW", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/maximorulli/nanochat-adamW
- SGLang
How to use maximorulli/nanochat-adamW with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "maximorulli/nanochat-adamW" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maximorulli/nanochat-adamW", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "maximorulli/nanochat-adamW" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "maximorulli/nanochat-adamW", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use maximorulli/nanochat-adamW with Docker Model Runner:
docker model run hf.co/maximorulli/nanochat-adamW
Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string
Nanochat d26 AdamW Full
This repository hosts maximorulli/nanochat-adamW, a d26 Nanochat decoder-only causal language model with custom Transformers remote code for loading the architecture and tokenizer.
Loading
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("maximorulli/nanochat-adamW", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("maximorulli/nanochat-adamW", trust_remote_code=True, torch_dtype="auto")
Architecture
- Decoder-only causal Transformer.
- 26 layers, hidden size 1664, 13 attention heads, 13 key/value heads.
- Context length 2048 and vocabulary size 32768.
- RoPE positional encoding, RMSNorm/pre-norm blocks, QK normalization.
- ReLU-squared MLP.
- Untied token embedding and LM head.
- Alternating value embeddings / ResFormer-style value paths; 13 value-embedding layers.
- Residual and x0 lambda scalar parameters.
- Linear layers without bias.
- Logit softcap:
15 * tanh(logits / 15). - Total parameters: 1,681,790,292.
- Estimated training FLOPs/token: 6.573818e9.
Tokenizer
The tokenizer is a custom RustBPE/tiktoken-style tokenizer with GPT-4-like splitting and 32768 vocabulary entries. It was trained with a 2,000,000,000 character cap and a 10,000 document cap.
Special tokens:
| Token | ID |
|---|---|
| `< | bos |
| `< | user_start |
| `< | user_end |
| `< | assistant_start |
| `< | assistant_end |
| `< | python_start |
| `< | python_end |
| `< | output_start |
| `< | output_end |
Token byte statistics: min 1, max 19, mean 6.6029, std 2.8250.
Base Checkpoint Metrics
| Metric | Value |
|---|---|
| Hub repo | maximorulli/nanochat-adamW |
| Variant | AdamW full |
| Checkpoint step | 7226 |
| Optimizer | adamw_only |
| Training tokens | 7,577,010,176 |
| Validation BPB | 0.861691430098635 |
| Batch size | 1,048,576 tokens |
| CORE estimate | not reported |
| MFU | 70.77% |
| Training FLOPs | 4.980988e+19 |
| Training time | 938.86 minutes |
| Peak memory | 47998.42 MiB |
- A later AdamW continuation checkpoint is outside the scope of this release.
Datasets
Base/tokenizer/pretraining:
karpathy/fineweb-edu-100b-shuffle, parquet shardsshard_00000.parquetthroughshard_01822.parquet.- Train split: all shards except the final validation shard.
- Validation split: final shard.
- Packing: BOS-aligned best-fit packing to 2048 tokens.
- Effective token utilization: 100% packed utilization; approximately 35% of overlong rows cropped at
T=2048.
SFT provenance documented for this training series:
HuggingFaceTB/smol-smoltalk, train and test.- Custom identity conversations, used twice during SFT.
cais/mmlu,auxiliary_trainfor SFT training andall/testfor validation/evaluation.openai/gsm8k,main/trainfor SFT training andmain/testfor validation/evaluation.- Synthetic
SimpleSpelling(size=200000, split=train). - Synthetic
SpellingBee(size=80000, split=train). - Synthetic
SpellingBee(size=256, split=test).
Evaluation datasets:
allenai/ai2_arc: ARC-Easy test and ARC-Challenge test.cais/mmlu: all/test.openai/gsm8k: main/test.openai/openai_humaneval: test.- Synthetic SpellingBee test.
- Tokenizer evaluation categories: news, Korean, code, math, science, FineWeb-Edu train, FineWeb-Edu validation.
Tokenizer Evaluation
Compared with GPT-2:
| Category | Bytes | GPT-2 tokens | GPT-2 ratio | Ours tokens | Ours ratio | Diff |
|---|---|---|---|---|---|---|
| news | 1819 | 404 | 4.5 | 403 | 4.51 | 0.2% |
| korean | 893 | 745 | 1.2 | 797 | 1.12 | -7.0% |
| code | 1259 | 576 | 2.19 | 620 | 2.03 | -7.6% |
| math | 1834 | 936 | 1.96 | 1025 | 1.79 | -9.5% |
| science | 1112 | 260 | 4.28 | 258 | 4.31 | 0.8% |
| fwe-train | 4208518 | 900364 | 4.67 | 892476 | 4.72 | 0.9% |
| fwe-val | 4776536 | 1031472 | 4.63 | 1026947 | 4.65 | 0.4% |
Compared with GPT-4:
| Category | GPT-4 tokens | GPT-4 ratio | Ours tokens | Ours ratio | Diff |
|---|---|---|---|---|---|
| news | 387 | 4.7 | 403 | 4.51 | -4.1% |
| korean | 364 | 2.45 | 797 | 1.12 | -119.0% |
| code | 309 | 4.07 | 620 | 2.03 | -100.6% |
| math | 832 | 2.2 | 1025 | 1.79 | -23.2% |
| science | 249 | 4.47 | 258 | 4.31 | -3.6% |
| fwe-train | 874799 | 4.81 | 892476 | 4.72 | -2.0% |
| fwe-val | 1004142 | 4.76 | 1026947 | 4.65 | -2.3% |
Related SFT And Downstream Metrics
These related downstream/SFT results are included for context and should not be read as direct base-checkpoint metrics unless explicitly stated.
Muon-full-derived SFT, source d26, SFT step 501:
| Task | Score |
|---|---|
| SFT validation BPB | 0.27659193826474665 |
| ARC-Easy | 0.6439 |
| ARC-Challenge | 0.5017 |
| MMLU | 0.3881 |
| GSM8K | 0.1039 |
| HumanEval | 0.122 |
| SpellingBee | 0.9922 |
| ChatCORE | 0.3772 |
Eval settings: dtype=bfloat16, temperature=0, max_new_tokens=512, top_k=50, batch_size=8.
AdamW-source SFT, base model step 7226, SFT step 501:
- SFT validation BPB: 0.38646146285770866.
- ChatCORE/downstream task metrics: not reported.
Muon step-2000-source SFT, base model step 2000, SFT step 80:
- SFT validation BPB: 0.3852505442606574.
- Stop reason: reached target validation BPB.
- ChatCORE/downstream task metrics: not reported.
Metric Availability Notes
- CORE estimates are not reported for these base checkpoints.
- Metrics are reported only when they are part of this release.
- A later AdamW continuation checkpoint is outside the scope of this release.
Intended Use
This is a small research/training artifact intended for experimentation with Nanochat-style training, optimizer comparisons, tokenizer behavior, and lightweight causal language modeling. It has not been aligned for broad production use.
Limitations
The model may produce incorrect, unsafe, biased, or low-quality text. The base checkpoints are not instruction-tuned chat assistants. Context length is 2048 tokens. Generation through the included Transformers implementation is portable but does not include the original training code's custom Flash Attention path.
- Downloads last month
- 8