Commit History

card: note GLM-5.1·2.7bpw build was removed from the Hub
c0bd329
verified

avlp12 commited on

docs: link parent collection
6bdc55c
verified

avlp12 commited on

Card+chart: sync 2.56 sibling to its re-measured clip+DWQ main (3.698/2.054, tasks 0.638n/0.812/0.744; wikitext gap −25%)
8f46693
verified

avlp12 commited on

Remove 76 orphaned old-main shards (replaced by 77-shard clip+DWQ build; old build preserved on dwq-8bit-noclip)
8a7010a
verified

avlp12 commited on

Chart render
e609fad
verified

avlp12 commited on

Chart: 2.777/1.835
444a06e
verified

avlp12 commited on

Card: clip+DWQ update (2.7774), branch map, router-KD verdict
dbb0672
verified

avlp12 commited on

main: anchor-guarded clip-search + 8-bit-teacher DWQ re-run — wikitext 2.7774 (prev 2.814 on dwq-8bit-noclip)
b0bc16b
verified

avlp12 commited on

Card: correct capacity-vs-teacher claim — 2.56 bpw tested, did NOT benefit from 8-bit teacher (sweet spot finding)
3511c9d
verified

avlp12 commited on

Chart source: 8-bit-teacher DWQ PPL
6d0f23b
verified

avlp12 commited on

Chart: 8-bit-teacher DWQ PPL
ef70c13
verified

avlp12 commited on

Card: DWQ-retuned against an 8-bit teacher (strided PPL wikitext 2.851->2.814, code 1.841->1.832)
7fdadbd
verified

avlp12 commited on

Add files using upload-large-folder tool
fc05ac1
verified

avlp12 commited on

Add files using upload-large-folder tool
c1a9803
verified

avlp12 commited on

Note MTP figures predate the DWQ retune (plain decode format-invariant; single-request MTP stays roughly neutral on the cleaner DWQ'd target)
9b81b66
verified

avlp12 commited on

DWQ-retuned benchmarks + chart: strided wikitext 2.946->2.851, code 1.893->1.841, tulu PPL 3.766->3.603 (KL vs 4.5bpw ref -24%); refresh 2.56 comparison to DWQ numbers; add DWQ section
f7839f1
verified

avlp12 commited on

DWQ-retuned weights (layerwise DWQ vs 4.5bpw teacher, 45% ZH mix): strided wikitext 2.946->2.851, code 1.893->1.841; tulu PPL 3.766->3.603; KL vs 4.5bpw ref 0.244->0.186 (-24%). MTP re-attached; pre-DWQ preserved on pre-dwq branch.
bf279b9
verified

avlp12 commited on

Upload README.md with huggingface_hub
3b299fe
verified

avlp12 commited on

Model card: long-context MTP fixed by the fork's small-L gather path (0.26x -> 0.85-0.98x); tuning guidance
489a50d
verified

avlp12 commited on

Model card: long-context MTP works (acceptance validated >2048; speed ~neutral there); corrected earlier misdiagnosis
eb822d5
verified

avlp12 commited on

Model card: MTP numbers measured on this build (perf-neutral single-request on 3-bit experts; +11% on a 4-bit sibling)
b4b2728
verified

avlp12 commited on

Model card: document the MTP context limit (dense regime <=2048 tokens; auto-fallback beyond)
6134c74
verified

avlp12 commited on

Add native MTP (nextn) layer 78 for self-speculative decoding (+4.5GB shard; backward-compatible — stripped by loaders without MTP support). Update model card.
324158a
verified

avlp12 commited on

Card: add evidence-based 'why int8 not fp8' for MLA-KV (int8 cos 0.99998 vs fp8 0.994-0.9997 on real latent; 100-200x lower MSE)
8876b85
verified

avlp12 commited on

Card: specify exact fork/PR + IndexShare load symptom (stock mlx-lm/omlx: Missing 285 indexer params)
f1faf6a
verified

avlp12 commited on

Upload README.md with huggingface_hub
e7589dd
verified

avlp12 commited on

Upload assets/recipe.svg with huggingface_hub
4e5e11b
verified

avlp12 commited on

Upload assets/recipe.png with huggingface_hub
8deb56a
verified

avlp12 commited on

Upload README.md with huggingface_hub
541863d
verified

avlp12 commited on

Upload README.md with huggingface_hub
61fb34d
verified

avlp12 commited on

Upload README.md with huggingface_hub
aed8dbb
verified

avlp12 commited on

Upload folder using huggingface_hub
f4a9c76
verified

avlp12 commited on

Upload folder using huggingface_hub
29f2653
verified

avlp12 commited on

Upload README.md with huggingface_hub
d758546
verified

avlp12 commited on

Upload README.md with huggingface_hub
38d6785
verified

avlp12 commited on

Upload folder using huggingface_hub
49f8742
verified

avlp12 commited on

Upload README.md with huggingface_hub
c899bc7
verified

avlp12 commited on

Upload folder using huggingface_hub
2f1d21c
verified

avlp12 commited on

Upload README.md with huggingface_hub
01fab14
verified

avlp12 commited on

Upload folder using huggingface_hub
2d6984b
verified

avlp12 commited on

initial commit
eb5b3df
verified

avlp12 commited on