Arki05 commited on
Commit
d556012
·
verified ·
1 Parent(s): 80349a8

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -45,3 +45,15 @@ BLS-Mini-Code-1.0-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
45
  BLS-Mini-Code-1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
46
  BLS-Mini-Code-1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
47
  BLS-Mini-Code-1.0.imatrix filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
 
 
 
45
  BLS-Mini-Code-1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
46
  BLS-Mini-Code-1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
47
  BLS-Mini-Code-1.0.imatrix filter=lfs diff=lfs merge=lfs -text
48
+ North-Mini-Code-1.0-BF16-00001-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
49
+ North-Mini-Code-1.0-BF16-00002-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
50
+ North-Mini-Code-1.0-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
51
+ North-Mini-Code-1.0-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
52
+ North-Mini-Code-1.0-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
53
+ North-Mini-Code-1.0-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
54
+ North-Mini-Code-1.0-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
55
+ North-Mini-Code-1.0-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
56
+ North-Mini-Code-1.0-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
57
+ North-Mini-Code-1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
58
+ North-Mini-Code-1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
59
+ North-Mini-Code-1.0.imatrix filter=lfs diff=lfs merge=lfs -text
BLS-Mini-Code-1.0-BF16-00001-of-00002.gguf → North-Mini-Code-1.0-BF16-00001-of-00002.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4070c9914714241ca3e1facc069c8a337afd0406761ed97a6f41098457467e8f
3
  size 44796475616
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e037609a4170dddde94277bf3497d8e1450068aae4bbdbcdcb023137cb397b2e
3
  size 44796475616
BLS-Mini-Code-1.0-BF16-00002-of-00002.gguf → North-Mini-Code-1.0-BF16-00002-of-00002.gguf RENAMED
File without changes
BLS-Mini-Code-1.0-IQ2_M.gguf → North-Mini-Code-1.0-IQ2_M.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4389183c1268066a65e673351995b18736aacd4e6005667ec8715ce38f167e39
3
  size 10258257728
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:76bf6be2729ee57dc8e2889861e0eca93a82a8165e25cb1d433d9212f02f8ee7
3
  size 10258257728
BLS-Mini-Code-1.0-IQ2_XS.gguf → North-Mini-Code-1.0-IQ2_XS.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6f02052c3140c3dad78e05f7d436ea698b1269a53fe5e736fcd0e83ab6229b8f
3
  size 9208223552
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a0c1f6cc368f4712c721a4ebc4c04ca5d26939d6d6f6fed82857a9eb62a128af
3
  size 9208223552
BLS-Mini-Code-1.0-IQ2_XXS.gguf → North-Mini-Code-1.0-IQ2_XXS.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4d461356a56907cb57f10b2ea8862396d3b92e017b7a76982e4ed9fe3bc1dfca
3
  size 8306022208
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:57a5f7b8f571388e23bddb3f481ed4cdd905e428a4b3aaf4a0d1843a34b2cb7f
3
  size 8306022208
BLS-Mini-Code-1.0-IQ3_M.gguf → North-Mini-Code-1.0-IQ3_M.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1d4c034a4179540646537386906d63afe692b1d4994ff4238daee42ca2486bd2
3
  size 13560125248
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c9983a8c53abef1ea7f63525702748d4cb6443b2f809e115aee91d01925db4eb
3
  size 13560125248
BLS-Mini-Code-1.0-IQ4_XS.gguf → North-Mini-Code-1.0-IQ4_XS.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6ab79c61ff36ad7280e588887df870f460fae4821cd3e7a403bde442f7cac2a1
3
  size 16412456768
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ed667972eb5553f4286cd509f3b5f282f4a031d705680c43bbcfef65435ff7c0
3
  size 16412456768
BLS-Mini-Code-1.0-Q4_K_M.gguf → North-Mini-Code-1.0-Q4_K_M.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b55a5fefbfb3e26bdd135badc63deeb99c1b732ea67da8895e8ab319a6f6ed60
3
  size 18593978176
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1dae9d226358088d3518b4fe7badd7c081cd1f64d11b5541187c788532c018e6
3
  size 18593978176
BLS-Mini-Code-1.0-Q5_K_M.gguf → North-Mini-Code-1.0-Q5_K_M.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:672791ad83e7117dc9960294d957d0dfc4a8142c9dd0a5ec75e2b79f489975ee
3
  size 21727778624
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:24c2a4360e5af87f859b65e9517cfe691929b98daa40cd21ab26d15e6efc2c91
3
  size 21727778624
BLS-Mini-Code-1.0-Q6_K.gguf → North-Mini-Code-1.0-Q6_K.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:81986d9a4fc5131fde0f4ed30fc5be6bca0d5a9564015269fd36cce5b17c26f5
3
  size 25057441600
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e1aa31df862f134858e59cfbd88bf73abf379978235a505fbc1da12d5c8f013b
3
  size 25057441600
BLS-Mini-Code-1.0-Q8_0.gguf → North-Mini-Code-1.0-Q8_0.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6ba4a1c8bafca1f70321efddb55517fdd59a6803b914be505808d8ac4cf695bc
3
  size 32437286720
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f3fec34f80131923eeb5ce40e53fc1bcc1eb0830362a2afe94556fe719360e87
3
  size 32437286720
BLS-Mini-Code-1.0.imatrix → North-Mini-Code-1.0.imatrix RENAMED
File without changes
README.md CHANGED
@@ -1,5 +1,6 @@
1
  ---
2
- base_model: CohereLabs/BLS-Mini-Code-1.0
 
3
  pipeline_tag: text-generation
4
  library_name: gguf
5
  tags:
@@ -7,30 +8,40 @@ tags:
7
  - moe
8
  - code
9
  - reasoning
 
10
  - gguf
11
  quantized_by: Arki05
12
  ---
13
 
14
- # BLS-Mini-Code-1.0 — GGUF
15
 
16
- GGUF quantizations of [CohereLabs/BLS-Mini-Code-1.0](https://huggingface.co/CohereLabs/BLS-Mini-Code-1.0),
17
  a 30.5B-total / ~2.9B-active sparse MoE code model by Cohere (`cohere2moe`
18
  architecture: Command-R7B-style hybrid SWA/full attention with NoPE on global
19
  layers, parallel residual blocks, 128 fine-grained experts with sigmoid top-8
20
- routing, reasoning-by-default chat format).
 
 
 
 
 
 
 
 
 
 
21
 
22
  > **Status / requirements:** needs llama.cpp with `cohere2moe` support —
23
  > [PR #24260](https://github.com/ggml-org/llama.cpp/pull/24260) (not yet merged).
24
- > Build that branch until it lands. The upstream model repo currently ships
25
- > **no license**; these files inherit whatever terms Cohere attaches to the
26
- > original weights.
27
 
28
  ## Quants
29
 
30
  All quality numbers are measured against the **bf16 model as ground truth**.
31
  The headline table uses **wikitext-2 (test)** — the only evaluation set that is
32
  fully held out from the imatrix calibration data — plus HumanEval/HumanEval+
33
- (pass@1, greedy, thinking on, 6k token budget; remaining quants in progress).
34
 
35
  | file | size | PPL | mean KLD | top-1 % | HumanEval | HumanEval+ |
36
  |---|---|---|---|---|---|---|
@@ -143,7 +154,7 @@ support automatically (`thinking = 1`), separates `reasoning_content` from
143
  working; rendering is byte-identical for native invocations.
144
 
145
  ```bash
146
- llama-server -m BLS-Mini-Code-1.0-Q5_K_M.gguf --jinja
147
  ```
148
 
149
  - thinking on (default): response arrives as `reasoning_content` + `content`
@@ -153,7 +164,7 @@ llama-server -m BLS-Mini-Code-1.0-Q5_K_M.gguf --jinja
153
 
154
  ## imatrix
155
 
156
- `BLS-Mini-Code-1.0.imatrix` (included) was computed on the **bf16** model over
157
  the v3 + code + chat mix described above (326x512-token chunks), reaching full
158
  coverage of all 128 experts in every layer.
159
 
@@ -164,5 +175,6 @@ coverage of all 128 experts in every layer.
164
  26/27, mean |dlogprob| 0.012 - the only disagreement a 0.013 near-tie.
165
  - Tool calling, parallel calls, multi-turn with reasoning passback, and a live
166
  agentic tool-execution loop verified end to end via `llama-server`.
167
- - 500k context advertised by the model; KV cache at long context stays small
 
168
  thanks to iSWA (only 13 of 49 layers are global; ~13.6 GB KV at 500k).
 
1
  ---
2
+ base_model: CohereLabs/North-Mini-Code-1.0
3
+ license: apache-2.0
4
  pipeline_tag: text-generation
5
  library_name: gguf
6
  tags:
 
8
  - moe
9
  - code
10
  - reasoning
11
+ - agent
12
  - gguf
13
  quantized_by: Arki05
14
  ---
15
 
16
+ # North-Mini-Code-1.0 — GGUF
17
 
18
+ GGUF quantizations of [CohereLabs/North-Mini-Code-1.0](https://huggingface.co/CohereLabs/North-Mini-Code-1.0),
19
  a 30.5B-total / ~2.9B-active sparse MoE code model by Cohere (`cohere2moe`
20
  architecture: Command-R7B-style hybrid SWA/full attention with NoPE on global
21
  layers, parallel residual blocks, 128 fine-grained experts with sigmoid top-8
22
+ routing, reasoning-by-default chat format). Trained with SFT followed by RL
23
+ with verifiable rewards, aimed at agentic coding and terminal/tool-use work —
24
+ see the [release blog post](https://huggingface.co/blog/CohereLabs/introducing-north-mini-code).
25
+
26
+ > **Provenance:** these GGUFs were converted from the weights Cohere first
27
+ > published as `CohereLabs/BLS-Mini-Code-1.0` (the pre-release name; the repo
28
+ > was renamed in place for the public launch). All 49 safetensors shards and
29
+ > the tokenizer of the final `North-Mini-Code-1.0` release are SHA256-identical
30
+ > to that pre-release — same weights, new name. Only file names and the GGUF
31
+ > `general.name`/`general.basename` metadata were updated here; tensor data is
32
+ > untouched, so all measurements below remain valid.
33
 
34
  > **Status / requirements:** needs llama.cpp with `cohere2moe` support —
35
  > [PR #24260](https://github.com/ggml-org/llama.cpp/pull/24260) (not yet merged).
36
+ > Build that branch until it lands. Weights are released under **Apache 2.0**,
37
+ > and these files inherit that license.
 
38
 
39
  ## Quants
40
 
41
  All quality numbers are measured against the **bf16 model as ground truth**.
42
  The headline table uses **wikitext-2 (test)** — the only evaluation set that is
43
  fully held out from the imatrix calibration data — plus HumanEval/HumanEval+
44
+ (pass@1, greedy, thinking on, 6k token budget).
45
 
46
  | file | size | PPL | mean KLD | top-1 % | HumanEval | HumanEval+ |
47
  |---|---|---|---|---|---|---|
 
154
  working; rendering is byte-identical for native invocations.
155
 
156
  ```bash
157
+ llama-server -m North-Mini-Code-1.0-Q5_K_M.gguf --jinja
158
  ```
159
 
160
  - thinking on (default): response arrives as `reasoning_content` + `content`
 
164
 
165
  ## imatrix
166
 
167
+ `North-Mini-Code-1.0.imatrix` (included) was computed on the **bf16** model over
168
  the v3 + code + chat mix described above (326x512-token chunks), reaching full
169
  coverage of all 128 experts in every layer.
170
 
 
175
  26/27, mean |dlogprob| 0.012 - the only disagreement a 0.013 near-tie.
176
  - Tool calling, parallel calls, multi-turn with reasoning passback, and a live
177
  agentic tool-execution loop verified end to end via `llama-server`.
178
+ - The official model card states 256K input / 64K output context; the config's
179
+ `max_position_embeddings` is 500k. KV cache at long context stays small
180
  thanks to iSWA (only 13 of 49 layers are global; ~13.6 GB KV at 500k).
eval-corpora.tar.zst CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:52c1d605d8429819d8e4f56a5fabfccee8b799e29e0e1e61c88a13f30a49551e
3
- size 640688
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:03eb4c521fe39eddea66bc581123f9a239b98b877586a65d050e787e8068d1b3
3
+ size 638292