Instructions to use brivangl/qwenar-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brivangl/qwenar-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="brivangl/qwenar-0.6b", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("brivangl/qwenar-0.6b", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from brivangl/qwenar-0.6b: direct link, hf CLI and curl.
- Browser
- Download file 14.2 kB
-
https://huggingface.co/brivangl/qwenar-0.6b/resolve/main/README.md
- Command line
-
hf download hf://brivangl/qwenar-0.6b/README.md
-
curl -L -o README.md https://huggingface.co/brivangl/qwenar-0.6b/resolve/main/README.md
license: apache-2.0
library_name: transformers
pipeline_tag: feature-extraction
language:
- en
tags:
- sentence-embeddings
- sentence-autoencoder
- text-reconstruction
- sonar
- large-concept-model
- dora
- relora
base_model:
- perplexity-ai/pplx-embed-v1-0.6b
- Qwen/Qwen3-0.6B-Base
datasets:
- MedRAG/pubmed
- MedRAG/textbooks
- DKYoon/SlimPajama-6B
- PleIAs/SYNTH
brivangl/qwenar-0.6b
A sentence is compressed into one 1024-dimensional vector, and that vector alone is decoded back into the sentence.
At its best checkpoint this model reconstructs 30 of 32 held-out probe sentences character-for-character from a single float vector.
The idea is SONAR's and the Large Concept Model's: a sentence embedding lossless
enough to decode, so the vector can stand in for the text. The construction is
different β instead of training a seq2seq model from scratch, qwenar connects
two off-the-shelf open checkpoints with a single learned linear bridge and
adapts both with DoRA. The whole learned connector is about 30 lines of code.
Code, training pipeline and full evaluation: https://github.com/IvanDrokin/QwenAR
Requires
transformers>=5.15, and older versions fail silently. This model was trained on 5.15.0. Under transformers 4.x it raises nothing and produces fluent, plausible text that is simply wrong: teacher-forced cross-entropy goes from 0.0052 to 3.2614 and next-token accuracy from 100% to 56%.
Usage
from qwenar.pipeline import Qwenar # pip install qwenar
qw = Qwenar.from_pretrained("brivangl/qwenar-0.6b")
vecs = qw.encode(["Serum ferritin remained within the reference range."])
print(vecs.shape) # torch.Size([1, 1024])
print(qw.decode(vecs)) # back to the sentence
print(qw.roundtrip(["..."])) # both at once
Or as a plain transformers model, with no extra dependency:
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("brivangl/qwenar-0.6b", trust_remote_code=True)
enc_tok = AutoTokenizer.from_pretrained("brivangl/qwenar-0.6b", subfolder="encoder_tokenizer")
dec_tok = AutoTokenizer.from_pretrained("brivangl/qwenar-0.6b")
batch = enc_tok(["Serum ferritin remained within the reference range."],
return_tensors="pt", padding=True)
emb = model.encode(batch["input_ids"], batch["attention_mask"]) # (1, 1024)
ids = model.generate_from_embeddings(emb, max_new_tokens=64, do_sample=False)
print(dec_tok.batch_decode(ids, skip_special_tokens=True))
Two tokenizers, because the encoder and the decoder come from different
checkpoints. The decoder tokenizer sits at the repository root so plain
AutoTokenizer.from_pretrained(repo) gives you the one generation needs; the
encoder tokenizer is in encoder_tokenizer/. Decoder targets must be
right-padded β Qwenar enforces this.
Architecture
sentence βββΊ perplexity-ai/pplx-embed-v1-0.6b
Qwen3 with the causal mask DISABLED β bidirectional
mean-pool over non-pad tokens; no tanh, no INT8
β
vec (1024,) βββ this is the artifact
β
EmbedToPrefix: Linear(1024 β KΒ·d_model) β GELU
β view(B, K, d_model) β RMSNorm K = 2
β
prefix (B, 2, d_model)
β
Qwen/Qwen3-0.6B-Base + DoRA(r=32, Ξ±=64) on q,k,v,o,gate,up,down
β
teacher-forced next-token cross-entropy
labels = -100 on the prefix and on padding
1194M parameters total. The encoder is adapted too (DoRA r=16, at a 10Γ lower learning rate), not frozen.
Training data
Sentence-level English, from a 22.83B-token corpus. The published run saw about 13.0B supervised tokens β roughly 57% of one epoch; it never completed a pass over the data.
| source | share of train tokens |
|---|---|
MedRAG/pubmed |
54.2% |
DKYoon/SlimPajama-6B |
30.0% |
PleIAs/SYNTH |
15.7% |
MedRAG/textbooks |
0.2% |
Over half the training signal is PubMed abstracts. These proportions were
not a design target β they are what the source directories happened to contain.
Sentences were segmented with SaT (sat-3l-sm) and capped at 96 tokens, so the
model has never seen a longer input.
The corpus itself is not released, and the dataset revisions were not recorded, so it cannot be rebuilt byte-for-byte. Full provenance, including what is missing, is in DATASET.md.
Training procedure
| Optimizer steps | 759,748 |
| Supervised tokens | 13,013,262,296 |
| Hardware | 4 GPUs, bf16 mixed precision |
| Wall clock | 329h 52m |
| Batching | token budget (4096 tokens/batch), not a fixed batch size |
| Adapters | DoRA β encoder r=16, decoder r=32/Ξ±=64 |
| ReLoRA | 7 merge + reinit cycles, every 100k steps |
| Learning rates | bridge 1e-4, decoder LoRA 1e-4, encoder 1e-5, cosine |
ReLoRA periodically merges the adapters into the backbone, zeroes lora_B,
clears the optimizer state and restarts the schedule. One consequence matters
for anyone picking a checkpoint: validation loss spikes after each merge, so
the last checkpoint is not the best one. These weights are step 700,000,
taken at the end of a converged cycle.
Evaluation
Reconstruction cross-entropy on held-out sentences, and the batching invariants.
best val/loss (= val/ce) |
0.0126 at step 700,000 |
| teacher-forced cross-entropy on the probe set | 0.0052 |
| next-token accuracy | 285/285 |
| exact reconstruction, 32 probe sentences | 30/32 |
| encode / decode batch invariance | ~6e-07 |
Batch invariance is worth stating explicitly: encoding or decoding a sentence alone gives the same result as doing it inside a padded batch, to float precision. Batching is a performance choice here, not a behavioural one.
The failure mode is consistent and worth knowing: the sentence frame survives while dense numerals and rare nomenclature scramble.
src: Musculoskeletal injury causes pain and when chronic can affect mental
health, employment and quality of life.
gen: Musculoskeletal injury causes pain and when chronic can affect mental
health, employment and quality of life.
src: arbazone 14, 1,3-thiazolines 15a-c and 4-thiazolidinone 16 exhibited
significant COX-2 inhibition ...
gen: arbazone 14, 13, 1-thiazolones4a-b and 15-cidine 2-thiazinoleic strain 60
exhibited significant ...
1024 floats carry a sentence's structure and content, but not arbitrary high-entropy strings.
Evaluation on general-domain text (SlimPajama, added 2026-09)
Out-of-domain evaluation added after the SlimPajama models were trained: the same battery that scores them, run on this checkpoint unchanged.
| teacher-forced cross-entropy (whole split, token-weighted) | 0.0141 |
| perplexity | 1.0142 |
| next-token accuracy | 99.58% |
| exact reconstruction, 8,000 sentences (greedy) | 96.21% |
| character error rate (Levenshtein / length) | 0.69% |
| word error rate | 1.00% |
| cosine(enc(src), enc(reconstruction)) | 0.9988 |
| reconstruction truncated at decoder max length | 0.00% |
| exact reconstruction, 32 stress probes | 26/32 |
| throughput, 8ΓH100 bf16 (encode / decode) | 15,098 / 1,016 sentences/s |
By sentence length β reconstruction degrades past ~48 tokens, and the 4B encoder holds up markedly longer:
| length (qwen3 tokens) | n | brivangl/qwenar-0.6b-montevideo exact / CER | brivangl/qwenar-4b-montevideo exact / CER | brivangl/qwenar-0.6b exact / CER |
|---|---|---|---|---|
| 0-15 | 3,064 | 100.0% / 0.00% | 100.0% / 0.00% | 99.9% / 0.00% |
| 16-31 | 3,426 | 99.7% / 0.02% | 99.9% / 0.00% | 99.0% / 0.09% |
| 32-47 | 1,282 | 97.0% / 0.40% | 99.3% / 0.07% | 87.6% / 1.65% |
| 48-63 | 208 | 79.3% / 3.77% | 97.1% / 0.10% | 57.2% / 10.45% |
| 64-79 | 14 | 7.1% / 31.85% | 28.6% / 13.36% | 7.1% / 44.94% |
| 80-95 | 3 | 0.0% / 40.72% | 33.3% / 5.71% | 0.0% / 52.92% |
| 96-111 | 3 | 0.0% / 42.76% | 0.0% / 24.60% | 0.0% / 50.54% |
Prose vs. code (rule-based detector; code is 0.7% of the split):
| content | n | brivangl/qwenar-0.6b-montevideo exact / CER | brivangl/qwenar-4b-montevideo exact / CER | brivangl/qwenar-0.6b exact / CER |
|---|---|---|---|---|
| prose | 7,941 | 98.7% / 0.23% | 99.6% / 0.04% | 96.3% / 0.65% |
| code | 59 | 81.4% / 3.98% | 94.9% / 0.92% | 78.0% / 7.15% |
The failure mode is the published one: the sentence frame survives while dense code syntax, identifiers, URLs and formula symbols scramble.
src: def flatten(xs): return [y for x in xs for y in (flatten(x) if isinstance(x, list) else [x])]
gen: def flatten(xs): return [y for x in xs for y in (flatten(x) if isinstance(x, list) else [x])]
src: fn parse(input: &str) -> Result<Vec<u32>, ParseIntError> { input.split(',').map(str::trim).map(str::parse).collect() }
gen: fn parse(input) : result &<&Str<Vec<UInt32>, parseErrorIndex> { inp.split(',|map').str.parse((map).str.parse()::trim())
(Levenshtein 58, CER 49.15%)
src: SELECT u.id, COUNT(o.id) AS orders FROM users u LEFT JOIN orders o ON o.user_id = u.id WHERE u.active GROUP BY u.id HAVING COUNT(o.id) > 3;
gen: SELECT u.id, COUNT(o.id) AS orders FROM users u JOIN LEFT orders o ON u.o_id = u.user WHERE u.id AS ACTIVE GROUP BY u.id hOrderBy u.id; > 0)
(Levenshtein 46, CER 33.09%)
src: func main() { ch := make(chan int, 8); go func() { ch <- 42 }(); fmt.Println(<-ch) }
gen: func main() { chan := ch.make, int8} { go func 4(); } ch := 23() <- fmt.Println; (ca+)
(Levenshtein 34, CER 40.48%)
Comparison with the SlimPajama-trained successors
| model | encoder | d_emb | K | params | training data | steps | corpus seen | CE β | tok acc β | exact β | CER β | WER β | cos β | probes β |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| brivangl/qwenar-0.6b-montevideo | pplx-embed-v1-0.6b | 1024 | 2 | 1194M | SlimPajama 9.2B sentences | 600,000 | 14% | 0.0047 | 99.85% | 98.61% | 0.26% | 0.39% | 0.9995 | 27/32 |
| brivangl/qwenar-4b-montevideo | pplx-embed-v1-4b | 2560 | 4 | 4629M | SlimPajama 9.2B sentences | 1,000,000 | 9% | 0.0012 | 99.96% | 99.59% | 0.05% | 0.10% | 0.9998 | 28/32 |
| brivangl/qwenar-0.6b | pplx-embed-v1-0.6b | 1024 | 2 | 1194M | PubMed / SlimPajama-6B / SYNTH mix | 750,000 | 57% ΒΉ | 0.0141 | 99.58% | 96.21% | 0.69% | 1.00% | 0.9988 | 26/32 |
All numbers on the same held-out SlimPajama split (200,000 sentences), same greedy decoding, bf16. CE / tok acc: teacher-forced over the whole split. exact / CER / WER / cos: free-running reconstruction of 8,000 sentences β CER = character Levenshtein / source length, WER = word Levenshtein / source words, cos = cosine between encoder embeddings of source and reconstruction. probes: exact reconstructions of 32 fixed stress sentences (long sentences, code, numbers, URLs).
ΒΉ this model; the successors (brivangl/qwenar-0.6b-montevideo, brivangl/qwenar-4b-montevideo) were trained on the 9.2B-sentence
SlimPajama corpus and are evaluated in-domain here. For general web text prefer
them; this checkpoint remains the stronger choice only for biomedical text, on
which the successors were not evaluated.
Limitations
Read this before drawing conclusions from the numbers above.
- Reconstruction is the only thing measured. There is no retrieval, STS, clustering, or downstream evaluation anywhere in this project. The model was trained purely to reconstruct, with no contrastive objective, so the embedding space is optimised to be decodable, not to be semantically well-shaped. Do not assume these vectors are good sentence embeddings for similarity tasks β nothing here tests that.
- Heavily biomedical. ~54% of training tokens are PubMed abstracts. Expect worse reconstruction on general text than the headline number suggests.
- English only. SONAR's defining property is a shared multilingual space; this model has no multilingual training or evaluation whatsoever.
- Sentences only, β€96 tokens, by construction of the training data. Behaviour on longer inputs is untested.
- One run, no ablations. Nothing isolates the contribution of K=2, DoRA vs LoRA, ReLoRA vs a single cycle, or an unfrozen vs frozen encoder. These are the hyperparameters that were used, not ones shown to be necessary.
- Less than one epoch. The run saw ~57% of the corpus.
- Trained on web and biomedical text, so it reproduces whatever biases and inaccuracies those carry. It is a reconstruction model: it will faithfully re-emit harmful or false input text.
License and attribution
Apache-2.0. Built from
perplexity-ai/pplx-embed-v1-0.6b (MIT) and
Qwen/Qwen3-0.6B-Base (Apache-2.0); the
bidirectional encoder code is derived from the former. Third-party code and
model licenses: THIRD_PARTY.md.
Implements ideas from SONAR, Large Concept Models, ReLoRA, DoRA and LoRA+.
Citation
@software{qwenar,
author = {Drokin, Ivan},
title = {qwenar: a sentence autoencoder built from open checkpoints},
year = {2026},
url = {https://github.com/IvanDrokin/QwenAR}
}