RealFalconsAI commited on
Commit
3e2212d
·
verified ·
1 Parent(s): 519dfcb

Upload 14 files

Browse files
.gitattributes ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ compact-int8/model_int8.safetensors filter=lfs diff=lfs merge=lfs -text
2
+ model.safetensors filter=lfs diff=lfs merge=lfs -text
compact-int8/README.md CHANGED
@@ -4,12 +4,12 @@ language: [en]
4
  library_name: transformers
5
  tags: [decision-model, system-one, calibrated-decisions, multiple-choice, zero-shot-classification, falcondec]
6
  ---
7
- # FalconDec v1.0.0 — Falcon Decision
8
 
9
  Single-pass, typed (`choice` / `noul` / `score`), calibrated closed-set decision model. Successor to
10
  [`Falconsai/proof_v2`](https://huggingface.co/Falconsai/proof_v2).
11
 
12
- * Built with FalconDec notebook V2 · backbone `jhu-clsp/ettin-encoder-150m` · mode `scratch` · preset `standard` · lineage: jhu-clsp/ettin-encoder-150m
13
  * Parameters: 159.7M · weights: fp16 305 MB, int8 153 MB
14
  * Layout: `[CLS] question [SEP] [MASK] opt1 … [MASK] optk [SEP] state [SEP]`; set-transformer option head;
15
  temperature per (question type × option-count bucket).
@@ -19,7 +19,7 @@ Single-pass, typed (`choice` / `noul` / `score`), calibrated closed-set decision
19
  ```python
20
  import importlib.util
21
  from huggingface_hub import snapshot_download
22
- path = snapshot_download("<repo>") # or a local FalconDec-v1.0.0 directory
23
  spec = importlib.util.spec_from_file_location("falcondec_modeling", f"{path}/falcondec_modeling.py")
24
  fdm = importlib.util.module_from_spec(spec); spec.loader.exec_module(fdm)
25
  model, tok = fdm.load_falcondec(path) # add "/compact-int8" for the int8 artefact
@@ -31,69 +31,69 @@ out = fdm.decide(model, tok, state={"body": "I was charged twice, refund me or I
31
 
32
  ## Test results
33
 
34
- Overall: micro acc **0.725**, task-macro acc **0.725**, ECE **0.025**,
35
- NLL 0.652, Brier 0.358, AURC 0.097.
36
 
37
  | task | domain | heldout | n | chance | acc | proof_v2 (card) | ece |
38
  |---|---|---|---|---|---|---|---|
39
- | counsel/critique_quality | agentic | False | 201 | 0.333 | 0.557 | | 0.094 |
40
- | counsel/step_has_error | agentic | False | 201 | 0.500 | 0.791 | | 0.093 |
41
- | hotpotqa/comparison_yes_no | agentic | False | 17 | 0.500 | 0.824 | | 0.185 |
42
- | hotpotqa/retrieve | agentic | False | 298 | 0.168 | 0.842 | | 0.078 |
43
- | sst5/score | classification | True | 300 | 0.200 | 0.370 | | 0.086 |
44
- | emotion/6way | classification | True | 300 | 0.167 | 0.480 | | 0.049 |
45
- | yelp/score | classification | False | 300 | 0.200 | 0.640 | | 0.069 |
46
- | ag_news/topic | classification | False | 300 | 0.250 | 0.847 | | 0.051 |
47
- | devign/vulnerability | code | False | 300 | 0.500 | 0.553 | 0.542 | 0.026 |
48
- | mbpp/bugspot | code | False | 156 | 0.413 | 0.731 | 0.475 | 0.086 |
49
- | humaneval/completion | code | True | 119 | 0.394 | 0.756 | 0.575 | 0.152 |
50
- | bigclonebench/clone | code | False | 300 | 0.500 | 0.863 | 0.383 | 0.098 |
51
- | codexglue/func_name | code | False | 275 | 0.263 | 0.938 | 0.901 | 0.021 |
52
- | codexglue/doc_to_code | code | False | 300 | 0.274 | 0.973 | 0.961 | 0.022 |
53
- | codexglue/code_to_doc | code | False | 300 | 0.256 | 0.977 | 0.969 | 0.016 |
54
- | mbpp/solution | code | False | 300 | 0.250 | 0.987 | 0.992 | 0.014 |
55
- | codexglue/lang_id | code | False | 300 | 0.235 | 1.000 | 0.997 | 0.004 |
56
- | prompt_injections/detect | guardrails | True | 116 | 0.500 | 0.647 | | 0.272 |
57
- | agentharm/refuse | guardrails | True | 416 | 0.500 | 0.654 | | 0.178 |
58
- | civil_comments/toxic | guardrails | False | 300 | 0.500 | 0.807 | | 0.051 |
59
- | jailbreak/detect | guardrails | False | 262 | 0.500 | 0.966 | | 0.018 |
60
- | banking77/intent_77 | intents | True | 300 | 0.013 | 0.470 | | 0.200 |
61
- | clinc150/intent | intents | False | 300 | 0.122 | 0.793 | 0.850 | 0.081 |
62
- | massive_en/intent | intents | False | 300 | 0.122 | 0.907 | | 0.046 |
63
- | banking77/intent | intents | True | 300 | 0.317 | 0.923 | 0.883 | 0.037 |
64
- | policy/table_extreme_transfer | policy | False | 300 | 0.240 | 0.290 | | 0.036 |
65
- | policy/invoice_overdue_transfer | policy | False | 300 | 0.500 | 0.407 | | 0.364 |
66
- | policy/count_threshold_transfer | policy | False | 300 | 0.179 | 0.523 | | 0.085 |
67
- | policy/invoice_total_transfer | policy | False | 300 | 0.500 | 0.523 | | 0.020 |
68
- | policy/refund_approval_transfer | policy | False | 300 | 0.333 | 0.710 | | 0.219 |
69
- | policy/sla_urgency_transfer | policy | False | 300 | 0.250 | 0.713 | | 0.214 |
70
- | policy/free_shipping_transfer | policy | False | 300 | 0.500 | 0.850 | | 0.067 |
71
- | policy/access_control_transfer | policy | False | 300 | 0.333 | 1.000 | | 0.001 |
72
- | policy/return_window_transfer | policy | False | 300 | 0.333 | 1.000 | | 0.019 |
73
- | aqua_rat/math | reasoning | False | 247 | 0.200 | 0.259 | | 0.071 |
74
- | hellaswag/continuation | reasoning | False | 300 | 0.250 | 0.383 | | 0.171 |
75
- | anli/nli | reasoning | False | 300 | 0.333 | 0.393 | | 0.185 |
76
- | mmlu/mcq | reasoning | True | 300 | 0.250 | 0.393 | | 0.094 |
77
- | arc_challenge/mcq | reasoning | True | 300 | 0.250 | 0.423 | 0.308 | 0.085 |
78
- | winogrande/blank | reasoning | False | 300 | 0.500 | 0.540 | | 0.089 |
79
- | openbookqa/mcq | reasoning | False | 300 | 0.250 | 0.550 | 0.292 | 0.068 |
80
- | arc_easy/mcq | reasoning | True | 300 | 0.250 | 0.553 | 0.425 | 0.057 |
81
- | commonsense_qa/mcq | reasoning | False | 296 | 0.200 | 0.611 | 0.442 | 0.058 |
82
- | gsm8k/math | reasoning | False | 300 | 0.250 | 0.637 | 0.275 | 0.050 |
83
- | boolq/yes_no | reasoning | False | 300 | 0.500 | 0.783 | 0.717 | 0.071 |
84
- | mnli/claim | reasoning | False | 300 | 0.333 | 0.807 | 0.492 | 0.072 |
85
- | snli/nli | reasoning | False | 594 | 0.333 | 0.837 | | 0.035 |
86
- | sciq/mcq | reasoning | False | 300 | 0.250 | 0.950 | 0.692 | 0.023 |
87
- | scitail/support | reasoning | False | 300 | 0.500 | 0.957 | | 0.030 |
88
- | snli/contradicts | reasoning | False | 300 | 0.333 | 0.967 | 0.892 | 0.035 |
89
- | snli/must_be_true | reasoning | False | 300 | 0.333 | 0.970 | 0.908 | 0.029 |
90
- | qasc/mcq | reasoning | False | 300 | 0.125 | 0.983 | | 0.007 |
91
- | bitext/category | support | False | 300 | 0.172 | 0.993 | | 0.008 |
92
- | bitext/route | support | False | 300 | 0.200 | 1.000 | 0.958 | 0.002 |
93
- | typed_decisions/invoice_processing | workflows | False | 500 | 0.350 | 0.628 | | 0.101 |
94
- | typed_decisions/agent_trace_observability | workflows | False | 500 | 0.300 | 0.696 | | 0.106 |
95
- | typed_decisions/security_incidents | workflows | False | 500 | 0.340 | 0.712 | | 0.140 |
96
- | typed_decisions/customer_service | workflows | False | 500 | 0.280 | 0.720 | | 0.099 |
97
 
98
 
99
  ## Training
@@ -101,7 +101,7 @@ NLL 0.652, Brier 0.358, AURC 0.097.
101
  Strictly proper objective (log + 0.5·spherical + 1.0·RPS for ordinal), soft teacher targets where available,
102
  task-balanced sampling (α = 0.5), NOTA augmentation (8%), per-epoch option reshuffling,
103
  AdamW + LLRD 0.9, EMA 0.999, 2 epoch(s), RLCD stage off/reverted.
104
- Training time 46.1 min on NVIDIA GeForce RTX 5090 Laptop GPU.
105
 
106
  ## Limitations
107
 
 
4
  library_name: transformers
5
  tags: [decision-model, system-one, calibrated-decisions, multiple-choice, zero-shot-classification, falcondec]
6
  ---
7
+ # LightDec v1.0.1 — Falcon Decision
8
 
9
  Single-pass, typed (`choice` / `noul` / `score`), calibrated closed-set decision model. Successor to
10
  [`Falconsai/proof_v2`](https://huggingface.co/Falconsai/proof_v2).
11
 
12
+ * Built with FalconDec notebook V2 · backbone `jhu-clsp/ettin-encoder-150m` · mode `finetune` · preset `full` · lineage: jhu-clsp/ettin-encoder-150m → FalconDec-v1.0.0
13
  * Parameters: 159.7M · weights: fp16 305 MB, int8 153 MB
14
  * Layout: `[CLS] question [SEP] [MASK] opt1 … [MASK] optk [SEP] state [SEP]`; set-transformer option head;
15
  temperature per (question type × option-count bucket).
 
19
  ```python
20
  import importlib.util
21
  from huggingface_hub import snapshot_download
22
+ path = snapshot_download("<repo>") # or a local FalconDec-v1.0.1 directory
23
  spec = importlib.util.spec_from_file_location("falcondec_modeling", f"{path}/falcondec_modeling.py")
24
  fdm = importlib.util.module_from_spec(spec); spec.loader.exec_module(fdm)
25
  model, tok = fdm.load_falcondec(path) # add "/compact-int8" for the int8 artefact
 
31
 
32
  ## Test results
33
 
34
+ Overall: micro acc **0.781**, task-macro acc **0.769**, ECE **0.020**,
35
+ NLL 0.529, Brier 0.290, AURC 0.064.
36
 
37
  | task | domain | heldout | n | chance | acc | proof_v2 (card) | ece |
38
  |---|---|---|---|---|---|---|---|
39
+ | counsel/critique_quality | agentic | False | 201 | 0.333 | 0.592 | | 0.227 |
40
+ | counsel/step_has_error | agentic | False | 201 | 0.500 | 0.801 | | 0.169 |
41
+ | hotpotqa/comparison_yes_no | agentic | False | 26 | 0.500 | 0.846 | | 0.151 |
42
+ | hotpotqa/retrieve | agentic | False | 497 | 0.167 | 0.873 | | 0.083 |
43
+ | sst5/score | classification | True | 500 | 0.200 | 0.416 | | 0.046 |
44
+ | emotion/6way | classification | True | 500 | 0.167 | 0.502 | | 0.082 |
45
+ | yelp/score | classification | False | 500 | 0.200 | 0.654 | | 0.069 |
46
+ | ag_news/topic | classification | False | 500 | 0.250 | 0.888 | | 0.040 |
47
+ | devign/vulnerability | code | False | 500 | 0.500 | 0.596 | 0.542 | 0.050 |
48
+ | humaneval/completion | code | True | 119 | 0.394 | 0.672 | 0.575 | 0.107 |
49
+ | mbpp/bugspot | code | False | 256 | 0.406 | 0.730 | 0.475 | 0.102 |
50
+ | bigclonebench/clone | code | False | 500 | 0.500 | 0.952 | 0.383 | 0.016 |
51
+ | codexglue/func_name | code | False | 467 | 0.255 | 0.953 | 0.901 | 0.023 |
52
+ | codexglue/doc_to_code | code | False | 504 | 0.279 | 0.976 | 0.961 | 0.010 |
53
+ | mbpp/solution | code | False | 500 | 0.250 | 0.982 | 0.992 | 0.015 |
54
+ | codexglue/code_to_doc | code | False | 504 | 0.262 | 0.984 | 0.969 | 0.012 |
55
+ | codexglue/lang_id | code | False | 504 | 0.235 | 1.000 | 0.997 | 0.001 |
56
+ | prompt_injections/detect | guardrails | True | 116 | 0.500 | 0.578 | | 0.286 |
57
+ | agentharm/refuse | guardrails | True | 416 | 0.500 | 0.743 | | 0.067 |
58
+ | civil_comments/toxic | guardrails | False | 500 | 0.500 | 0.906 | | 0.035 |
59
+ | jailbreak/detect | guardrails | False | 262 | 0.500 | 0.954 | | 0.023 |
60
+ | banking77/intent_77 | intents | True | 300 | 0.013 | 0.533 | | 0.147 |
61
+ | massive_en/intent | intents | False | 500 | 0.130 | 0.906 | | 0.022 |
62
+ | banking77/intent | intents | True | 500 | 0.321 | 0.930 | 0.883 | 0.040 |
63
+ | clinc150/intent | intents | False | 500 | 0.124 | 0.942 | 0.850 | 0.036 |
64
+ | policy/invoice_total_transfer | policy | False | 500 | 0.500 | 0.486 | | 0.032 |
65
+ | policy/invoice_overdue_transfer | policy | False | 500 | 0.500 | 0.600 | | 0.201 |
66
+ | policy/sla_urgency_transfer | policy | False | 500 | 0.250 | 0.698 | | 0.212 |
67
+ | policy/free_shipping_transfer | policy | False | 500 | 0.500 | 0.766 | | 0.101 |
68
+ | policy/refund_approval_transfer | policy | False | 500 | 0.333 | 0.780 | | 0.186 |
69
+ | policy/count_threshold_transfer | policy | False | 500 | 0.178 | 0.846 | | 0.109 |
70
+ | policy/table_extreme_transfer | policy | False | 500 | 0.237 | 0.912 | | 0.014 |
71
+ | policy/return_window_transfer | policy | False | 500 | 0.333 | 0.998 | | 0.020 |
72
+ | policy/access_control_transfer | policy | False | 500 | 0.333 | 1.000 | | 0.000 |
73
+ | aqua_rat/math | reasoning | False | 247 | 0.200 | 0.275 | | 0.042 |
74
+ | mmlu/mcq | reasoning | True | 500 | 0.250 | 0.356 | | 0.109 |
75
+ | anli/nli | reasoning | False | 498 | 0.333 | 0.416 | | 0.184 |
76
+ | arc_challenge/mcq | reasoning | True | 500 | 0.250 | 0.460 | 0.308 | 0.088 |
77
+ | hellaswag/continuation | reasoning | False | 500 | 0.250 | 0.580 | | 0.073 |
78
+ | arc_easy/mcq | reasoning | True | 500 | 0.250 | 0.594 | 0.425 | 0.038 |
79
+ | openbookqa/mcq | reasoning | False | 500 | 0.250 | 0.610 | 0.292 | 0.092 |
80
+ | winogrande/blank | reasoning | False | 500 | 0.500 | 0.612 | | 0.111 |
81
+ | commonsense_qa/mcq | reasoning | False | 493 | 0.200 | 0.639 | 0.442 | 0.044 |
82
+ | gsm8k/math | reasoning | False | 500 | 0.250 | 0.662 | 0.275 | 0.061 |
83
+ | boolq/yes_no | reasoning | False | 500 | 0.500 | 0.816 | 0.717 | 0.027 |
84
+ | mnli/claim | reasoning | False | 500 | 0.333 | 0.838 | 0.492 | 0.071 |
85
+ | snli/nli | reasoning | False | 988 | 0.333 | 0.880 | | 0.050 |
86
+ | sciq/mcq | reasoning | False | 498 | 0.250 | 0.954 | 0.692 | 0.031 |
87
+ | scitail/support | reasoning | False | 500 | 0.500 | 0.964 | | 0.020 |
88
+ | snli/must_be_true | reasoning | False | 500 | 0.333 | 0.980 | 0.908 | 0.029 |
89
+ | snli/contradicts | reasoning | False | 500 | 0.333 | 0.986 | 0.892 | 0.014 |
90
+ | qasc/mcq | reasoning | False | 500 | 0.125 | 0.990 | | 0.006 |
91
+ | bitext/category | support | False | 500 | 0.161 | 1.000 | | 0.002 |
92
+ | bitext/route | support | False | 500 | 0.190 | 1.000 | 0.958 | 0.001 |
93
+ | typed_decisions/agent_trace_observability | workflows | False | 500 | 0.300 | 0.712 | | 0.120 |
94
+ | typed_decisions/security_incidents | workflows | False | 500 | 0.340 | 0.730 | | 0.144 |
95
+ | typed_decisions/customer_service | workflows | False | 500 | 0.280 | 0.740 | | 0.085 |
96
+ | typed_decisions/invoice_processing | workflows | False | 500 | 0.350 | 0.838 | | 0.122 |
97
 
98
 
99
  ## Training
 
101
  Strictly proper objective (log + 0.5·spherical + 1.0·RPS for ordinal), soft teacher targets where available,
102
  task-balanced sampling (α = 0.5), NOTA augmentation (8%), per-epoch option reshuffling,
103
  AdamW + LLRD 0.9, EMA 0.999, 2 epoch(s), RLCD stage off/reverted.
104
+ Training time 172.2 min on NVIDIA GeForce RTX 5090 Laptop GPU.
105
 
106
  ## Limitations
107
 
compact-int8/falcondec_config.json CHANGED
@@ -7,13 +7,14 @@
7
  "max_opts_single_pass": 96,
8
  "interact_layers": 2,
9
  "interact_heads": 8,
10
- "name": "FalconDec",
11
  "backbone": "jhu-clsp/ettin-encoder-150m",
12
  "backbone_init": "pretrained",
13
  "lineage": [
14
- "jhu-clsp/ettin-encoder-150m"
 
15
  ],
16
- "parent_version": null,
17
  "special": {
18
  "cls": 50281,
19
  "sep": 50282,
@@ -27,27 +28,27 @@
27
  "score": 2
28
  },
29
  "n_buckets": 4,
30
- "version": "1.0.0",
31
  "notebook_version": "V2",
32
- "created": "2026-09-26T04:35:05.339042Z",
33
  "temperature": [
34
  [
35
- 1.7067416906356812,
36
- 1.3524072170257568,
37
- 1.4660307168960571,
38
- 1.3491928577423096
39
  ],
40
  [
41
- 1.4017767906188965,
42
- 1.4017767906188965,
43
- 1.4017767906188965,
44
- 1.4017767906188965
45
  ],
46
  [
47
- 1.4534460306167603,
48
- 1.4534460306167603,
49
- 1.4534460306167603,
50
- 1.4534460306167603
51
  ]
52
  ],
53
  "weights": "model_int8.safetensors"
 
7
  "max_opts_single_pass": 96,
8
  "interact_layers": 2,
9
  "interact_heads": 8,
10
+ "name": "LightDec",
11
  "backbone": "jhu-clsp/ettin-encoder-150m",
12
  "backbone_init": "pretrained",
13
  "lineage": [
14
+ "jhu-clsp/ettin-encoder-150m",
15
+ "FalconDec-v1.0.0"
16
  ],
17
+ "parent_version": "1.0.0",
18
  "special": {
19
  "cls": 50281,
20
  "sep": 50282,
 
28
  "score": 2
29
  },
30
  "n_buckets": 4,
31
+ "version": "1.0.1",
32
  "notebook_version": "V2",
33
+ "created": "2026-09-26T11:28:22.623723Z",
34
  "temperature": [
35
  [
36
+ 1.3723914623260498,
37
+ 1.3805755376815796,
38
+ 1.5033340454101562,
39
+ 1.320592999458313
40
  ],
41
  [
42
+ 1.3903899192810059,
43
+ 1.3903899192810059,
44
+ 1.3903899192810059,
45
+ 1.3903899192810059
46
  ],
47
  [
48
+ 1.4738539457321167,
49
+ 1.4738539457321167,
50
+ 1.4738539457321167,
51
+ 1.4738539457321167
52
  ]
53
  ],
54
  "weights": "model_int8.safetensors"
compact-int8/model_int8.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:80796383c3b131a6540b01642a8cb929f53e8d84b4c0f3f159248e56c11a7624
3
  size 160530522
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:174e111a329070aaef5597db916a4045aef6de8ce4398c2f73a5975fbe9d700b
3
  size 160530522
compact-int8/tokenizer/tokenizer_config.json CHANGED
@@ -2,9 +2,10 @@
2
  "backend": "tokenizers",
3
  "clean_up_tokenization_spaces": true,
4
  "cls_token": "[CLS]",
5
- "is_local": false,
6
  "local_files_only": false,
7
  "mask_token": "[MASK]",
 
8
  "model_input_names": [
9
  "input_ids",
10
  "attention_mask"
@@ -12,6 +13,9 @@
12
  "model_max_length": 8192,
13
  "pad_token": "[PAD]",
14
  "sep_token": "[SEP]",
 
15
  "tokenizer_class": "TokenizersBackend",
 
 
16
  "unk_token": "[UNK]"
17
  }
 
2
  "backend": "tokenizers",
3
  "clean_up_tokenization_spaces": true,
4
  "cls_token": "[CLS]",
5
+ "is_local": true,
6
  "local_files_only": false,
7
  "mask_token": "[MASK]",
8
+ "max_length": 256,
9
  "model_input_names": [
10
  "input_ids",
11
  "attention_mask"
 
13
  "model_max_length": 8192,
14
  "pad_token": "[PAD]",
15
  "sep_token": "[SEP]",
16
+ "stride": 0,
17
  "tokenizer_class": "TokenizersBackend",
18
+ "truncation_side": "right",
19
+ "truncation_strategy": "longest_first",
20
  "unk_token": "[UNK]"
21
  }
encoder/config.json ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "ModernBertForMaskedLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 50281,
8
+ "causal_mask": false,
9
+ "classifier_activation": "gelu",
10
+ "classifier_bias": false,
11
+ "classifier_dropout": 0.0,
12
+ "classifier_pooling": "mean",
13
+ "cls_token_id": 50281,
14
+ "decoder_bias": true,
15
+ "deterministic_flash_attn": false,
16
+ "dtype": "float32",
17
+ "embedding_dropout": 0.0,
18
+ "eos_token_id": 50282,
19
+ "global_attn_every_n_layers": 3,
20
+ "gradient_checkpointing": false,
21
+ "hidden_activation": "gelu",
22
+ "hidden_size": 768,
23
+ "initializer_cutoff_factor": 2.0,
24
+ "initializer_range": 0.02,
25
+ "intermediate_size": 1152,
26
+ "is_causal": false,
27
+ "layer_norm_eps": 1e-05,
28
+ "layer_types": [
29
+ "full_attention",
30
+ "sliding_attention",
31
+ "sliding_attention",
32
+ "full_attention",
33
+ "sliding_attention",
34
+ "sliding_attention",
35
+ "full_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "full_attention",
39
+ "sliding_attention",
40
+ "sliding_attention",
41
+ "full_attention",
42
+ "sliding_attention",
43
+ "sliding_attention",
44
+ "full_attention",
45
+ "sliding_attention",
46
+ "sliding_attention",
47
+ "full_attention",
48
+ "sliding_attention",
49
+ "sliding_attention",
50
+ "full_attention"
51
+ ],
52
+ "local_attention": 128,
53
+ "max_position_embeddings": 7999,
54
+ "mlp_bias": false,
55
+ "mlp_dropout": 0.0,
56
+ "model_type": "modernbert",
57
+ "norm_bias": false,
58
+ "norm_eps": 1e-05,
59
+ "num_attention_heads": 12,
60
+ "num_hidden_layers": 22,
61
+ "pad_token_id": 50283,
62
+ "position_embedding_type": "sans_pos",
63
+ "rope_parameters": {
64
+ "full_attention": {
65
+ "rope_theta": 160000.0,
66
+ "rope_type": "default"
67
+ },
68
+ "sliding_attention": {
69
+ "rope_theta": 160000.0,
70
+ "rope_type": "default"
71
+ }
72
+ },
73
+ "sep_token_id": 50282,
74
+ "sparse_pred_ignore_index": -100,
75
+ "sparse_prediction": false,
76
+ "tie_word_embeddings": true,
77
+ "transformers_version": "5.17.0",
78
+ "vocab_size": 50368
79
+ }
falcondec_config.json ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "max_len": 512,
3
+ "long_max_len": 2048,
4
+ "long_opts_threshold": 24,
5
+ "head_max_len": 192,
6
+ "max_tok_per_opt": 24,
7
+ "max_opts_single_pass": 96,
8
+ "interact_layers": 2,
9
+ "interact_heads": 8,
10
+ "name": "LightDec",
11
+ "backbone": "jhu-clsp/ettin-encoder-150m",
12
+ "backbone_init": "pretrained",
13
+ "lineage": [
14
+ "jhu-clsp/ettin-encoder-150m",
15
+ "FalconDec-v1.0.0"
16
+ ],
17
+ "parent_version": "1.0.0",
18
+ "special": {
19
+ "cls": 50281,
20
+ "sep": 50282,
21
+ "mask": 50284,
22
+ "pad": 50283
23
+ },
24
+ "defer_threshold": 0.7,
25
+ "qtypes": {
26
+ "choice": 0,
27
+ "noul": 1,
28
+ "score": 2
29
+ },
30
+ "n_buckets": 4,
31
+ "version": "1.0.1",
32
+ "notebook_version": "V2",
33
+ "created": "2026-09-26T11:28:22.623723Z",
34
+ "temperature": [
35
+ [
36
+ 1.3723914623260498,
37
+ 1.3805755376815796,
38
+ 1.5033340454101562,
39
+ 1.320592999458313
40
+ ],
41
+ [
42
+ 1.3903899192810059,
43
+ 1.3903899192810059,
44
+ 1.3903899192810059,
45
+ 1.3903899192810059
46
+ ],
47
+ [
48
+ 1.4738539457321167,
49
+ 1.4738539457321167,
50
+ 1.4738539457321167,
51
+ 1.4738539457321167
52
+ ]
53
+ ],
54
+ "weights": "model.safetensors"
55
+ }
falcondec_modeling.py ADDED
@@ -0,0 +1,358 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """FalconDec — Falcon Decision model: single-pass, typed, calibrated closed-set decisions.
3
+
4
+ Layout : [CLS] question [SEP] [MASK] opt_1 ... [MASK] opt_k [SEP] state [SEP]
5
+ Head : marker vectors + CLS context + question-type embedding
6
+ -> permutation-equivariant set transformer (options attend to each other; no positions)
7
+ -> MLP -> one logit per option -> softmax over this question's options
8
+ Calib. : temperature per (question type, option-count bucket), stored in the checkpoint
9
+ """
10
+ from __future__ import annotations
11
+
12
+ import contextlib
13
+ import json
14
+ import shutil
15
+ from pathlib import Path
16
+
17
+ import numpy as np
18
+ import torch
19
+ import torch.nn as nn
20
+
21
+ QTYPES = {"choice": 0, "noul": 1, "score": 2}
22
+ N_BUCKETS = 4
23
+ CONFIG_FILE = "falcondec_config.json"
24
+ FP16_FILE = "model.safetensors"
25
+ INT8_FILE = "model_int8.safetensors"
26
+ MODEL_KEYS = ("input_ids", "attention_mask", "marker_pos", "marker_mask", "qtype")
27
+
28
+
29
+ def n_bucket(n: int) -> int:
30
+ return 0 if n <= 2 else 1 if n <= 5 else 2 if n <= 12 else 3
31
+
32
+
33
+ def special_ids(tok) -> dict:
34
+ sp = {"cls": tok.cls_token_id, "sep": tok.sep_token_id, "mask": tok.mask_token_id, "pad": tok.pad_token_id}
35
+ if sp["cls"] is None:
36
+ sp["cls"] = tok.bos_token_id
37
+ if sp["sep"] is None:
38
+ sp["sep"] = tok.eos_token_id
39
+ if sp["pad"] is None:
40
+ sp["pad"] = sp["sep"]
41
+ if sp["mask"] is None:
42
+ raise ValueError("FalconDec needs a tokenizer with a mask token (used as the option marker).")
43
+ return {k: int(v) for k, v in sp.items()}
44
+
45
+
46
+ def _as_list(x):
47
+ return x.tolist() if hasattr(x, "tolist") else list(x)
48
+
49
+
50
+ def assemble(q_ids, opt_ids, s_ids, sp, max_len=512, head_max_len=192, max_tok_per_opt=24,
51
+ long_max_len=2048, long_opts_threshold=24):
52
+ """Build one input sequence. Returns (ids, marker_positions)."""
53
+ n = len(opt_ids)
54
+ eff = max_len if n <= long_opts_threshold else max(max_len, long_max_len)
55
+ q = _as_list(q_ids)[:96]
56
+ need = len(q) + 2 + n * (max_tok_per_opt + 1)
57
+ head = min(eff - 64, max(head_max_len, need))
58
+ per = max(2, min(max_tok_per_opt, (head - len(q) - 2) // max(n, 1) - 1))
59
+ ids = [sp["cls"]] + q + [sp["sep"]]
60
+ markers = []
61
+ for o in opt_ids:
62
+ markers.append(len(ids))
63
+ ids.append(sp["mask"])
64
+ ids.extend(_as_list(o[:per]))
65
+ ids.append(sp["sep"])
66
+ room = eff - len(ids) - 1
67
+ if room > 0 and s_ids is not None and len(s_ids) > 0:
68
+ ids.extend(_as_list(s_ids[:room]))
69
+ ids.append(sp["sep"])
70
+ if len(ids) > eff:
71
+ ids = ids[:eff]
72
+ if markers and markers[-1] >= len(ids):
73
+ raise ValueError("Too many options for one pass; reduce options or raise long_max_len.")
74
+ return ids, markers
75
+
76
+
77
+ def collate_features(feats, pad_id, device=None):
78
+ """feats: list of (ids, markers, qtype_index) -> dict of padded tensors matching FalconDec.forward."""
79
+ B = len(feats)
80
+ T = max(len(f[0]) for f in feats)
81
+ K = max(len(f[1]) for f in feats)
82
+ ids = torch.full((B, T), pad_id, dtype=torch.long)
83
+ att = torch.zeros((B, T), dtype=torch.long)
84
+ mpos = torch.zeros((B, K), dtype=torch.long)
85
+ mmask = torch.zeros((B, K), dtype=torch.bool)
86
+ qt = torch.zeros(B, dtype=torch.long)
87
+ for i, (x, mk, q) in enumerate(feats):
88
+ ids[i, : len(x)] = torch.as_tensor(x, dtype=torch.long)
89
+ att[i, : len(x)] = 1
90
+ mpos[i, : len(mk)] = torch.as_tensor(mk, dtype=torch.long)
91
+ mmask[i, : len(mk)] = True
92
+ qt[i] = int(q)
93
+ out = dict(input_ids=ids, attention_mask=att, marker_pos=mpos, marker_mask=mmask, qtype=qt)
94
+ if device is not None:
95
+ out = {k: v.to(device, non_blocking=True) for k, v in out.items()}
96
+ return out
97
+
98
+
99
+ class OptionInteraction(nn.Module):
100
+ """Set transformer over the options of one question (no positional encoding => order-equivariant)."""
101
+
102
+ def __init__(self, d, n_layers=2, n_heads=8, dropout=0.1):
103
+ super().__init__()
104
+ layer = nn.TransformerEncoderLayer(d, n_heads, dim_feedforward=2 * d, dropout=dropout,
105
+ activation="gelu", batch_first=True, norm_first=True)
106
+ self.enc = nn.TransformerEncoder(layer, n_layers, enable_nested_tensor=False)
107
+
108
+ def forward(self, x, mask):
109
+ return self.enc(x, src_key_padding_mask=~mask)
110
+
111
+
112
+ class FalconDec(nn.Module):
113
+ def __init__(self, encoder, fcfg: dict):
114
+ super().__init__()
115
+ self.encoder = encoder
116
+ self.fcfg = dict(fcfg)
117
+ d = encoder.config.hidden_size
118
+ self.qtype_emb = nn.Embedding(len(QTYPES), d)
119
+ self.ctx_proj = nn.Linear(d, d)
120
+ self.opt_norm = nn.LayerNorm(d)
121
+ self.interact = OptionInteraction(d, int(self.fcfg.get("interact_layers", 2)),
122
+ int(self.fcfg.get("interact_heads", 8)))
123
+ self.scorer = nn.Sequential(nn.Linear(d, d), nn.GELU(), nn.Dropout(0.1), nn.Linear(d, 1))
124
+ self.register_buffer("temperature", torch.ones(len(QTYPES), N_BUCKETS))
125
+
126
+ @property
127
+ def device(self):
128
+ return next(self.parameters()).device
129
+
130
+ def num_parameters(self):
131
+ return sum(p.numel() for p in self.parameters())
132
+
133
+ def forward(self, input_ids, attention_mask, marker_pos, marker_mask, qtype):
134
+ h = self.encoder(input_ids=input_ids, attention_mask=attention_mask).last_hidden_state
135
+ B, K = marker_pos.shape
136
+ idx = marker_pos.unsqueeze(-1).expand(B, K, h.size(-1))
137
+ opt = h.gather(1, idx)
138
+ ctx = self.ctx_proj(h[:, 0]).unsqueeze(1)
139
+ x = self.opt_norm(opt + ctx + self.qtype_emb(qtype).unsqueeze(1))
140
+ x = self.interact(x, marker_mask)
141
+ logits = self.scorer(x).squeeze(-1).float()
142
+ return logits.masked_fill(~marker_mask, -1e4)
143
+
144
+
145
+ # ------------------------------------------------------------------ storage
146
+ def clean_state_dict(sd):
147
+ return {k.replace("_orig_mod.", ""): v.detach().cpu().contiguous() for k, v in sd.items()}
148
+
149
+
150
+ def quantize_int8(sd, min_numel=4096):
151
+ """Per-output-channel symmetric int8 for every matrix; fp16 for everything else."""
152
+ out = {}
153
+ for k, v in sd.items():
154
+ if v.is_floating_point() and v.ndim == 2 and v.numel() >= min_numel:
155
+ w = v.float()
156
+ s = (w.abs().amax(dim=1, keepdim=True) / 127.0).clamp_min(1e-12)
157
+ out[k] = torch.round(w / s).clamp_(-127, 127).to(torch.int8).contiguous()
158
+ out[k + "::scale"] = s.contiguous()
159
+ elif v.is_floating_point():
160
+ out[k] = v.half().contiguous()
161
+ else:
162
+ out[k] = v.contiguous()
163
+ return out
164
+
165
+
166
+ def dequantize_int8(sd):
167
+ out = {}
168
+ for k, v in sd.items():
169
+ if k.endswith("::scale"):
170
+ continue
171
+ s = sd.get(k + "::scale")
172
+ out[k] = (v.float() * s).half() if s is not None else v
173
+ return out
174
+
175
+
176
+ def save_falcondec(model, tok, out_dir, int8=False, extra_files=None):
177
+ from safetensors.torch import save_file
178
+ out = Path(out_dir)
179
+ out.mkdir(parents=True, exist_ok=True)
180
+ enc_cfg = model.encoder.config
181
+ if hasattr(enc_cfg, "reference_compile"):
182
+ enc_cfg.reference_compile = False
183
+ enc_cfg.save_pretrained(str(out / "encoder"))
184
+ tok.save_pretrained(str(out / "tokenizer"))
185
+ sd = clean_state_dict(model.state_dict())
186
+ fc = dict(model.fcfg)
187
+ fc["temperature"] = model.temperature.detach().float().cpu().tolist()
188
+ fc["weights"] = INT8_FILE if int8 else FP16_FILE
189
+ if int8:
190
+ save_file(quantize_int8(sd), str(out / INT8_FILE), metadata={"format": "falcondec-int8"})
191
+ else:
192
+ save_file({k: (v.half() if v.is_floating_point() else v) for k, v in sd.items()},
193
+ str(out / FP16_FILE), metadata={"format": "falcondec-fp16"})
194
+ (out / CONFIG_FILE).write_text(json.dumps(fc, indent=2, default=str), encoding="utf-8")
195
+ try:
196
+ shutil.copy(__file__, out / "falcondec_modeling.py")
197
+ except Exception:
198
+ pass
199
+ for name, content in (extra_files or {}).items():
200
+ (out / name).write_text(content, encoding="utf-8")
201
+ return out
202
+
203
+
204
+ def load_falcondec(path, device=None, dtype=None, attn_implementation="sdpa"):
205
+ """Load a FalconDec directory (fp16 or int8) or Hub repo. Returns (model, tokenizer)."""
206
+ from safetensors.torch import load_file
207
+ from transformers import AutoConfig, AutoModel, AutoTokenizer
208
+ p = Path(path)
209
+ if not p.exists():
210
+ from huggingface_hub import snapshot_download
211
+ p = Path(snapshot_download(str(path)))
212
+ fc = json.loads((p / CONFIG_FILE).read_text(encoding="utf-8"))
213
+ ecfg = AutoConfig.from_pretrained(str(p / "encoder"))
214
+ if hasattr(ecfg, "reference_compile"):
215
+ ecfg.reference_compile = False
216
+ try:
217
+ enc = AutoModel.from_config(ecfg, attn_implementation=attn_implementation)
218
+ except Exception:
219
+ enc = AutoModel.from_config(ecfg)
220
+ model = FalconDec(enc, fc)
221
+ wf = p / fc.get("weights", FP16_FILE)
222
+ if not wf.exists():
223
+ wf = p / (INT8_FILE if (p / INT8_FILE).exists() else FP16_FILE)
224
+ sd = load_file(str(wf))
225
+ if wf.name == INT8_FILE:
226
+ sd = dequantize_int8(sd)
227
+ missing, unexpected = model.load_state_dict(sd, strict=False)
228
+ if missing or unexpected:
229
+ print(f"[FalconDec] load warning: missing={list(missing)[:5]} unexpected={list(unexpected)[:5]}")
230
+ tok = AutoTokenizer.from_pretrained(str(p / "tokenizer"))
231
+ dev = torch.device(device) if device is not None else torch.device("cuda" if torch.cuda.is_available() else "cpu")
232
+ model.to(dev)
233
+ if dtype is not None:
234
+ model.to(dtype)
235
+ model.eval()
236
+ return model, tok
237
+
238
+
239
+ # ------------------------------------------------------------------ inference
240
+ def _amp(model):
241
+ if model.device.type == "cuda" and next(model.parameters()).dtype == torch.float32:
242
+ dt = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
243
+ return torch.autocast("cuda", dtype=dt)
244
+ return contextlib.nullcontext()
245
+
246
+
247
+ def _state_text(state):
248
+ if state is None:
249
+ return ""
250
+ return state if isinstance(state, str) else json.dumps(state, ensure_ascii=False)
251
+
252
+
253
+ @torch.no_grad()
254
+ def score_items(model, tok, items, batch_size=32):
255
+ """items: [{"state", "question", "options", "type"?, "option_tokens"?, "seq_len"?}] -> list of prob arrays."""
256
+ fc = model.fcfg
257
+ M = int(fc.get("max_opts_single_pass", 96))
258
+ out = [None] * len(items)
259
+ small = [i for i, it in enumerate(items) if len(it["options"]) <= M]
260
+ for i, it in enumerate(items):
261
+ if len(it["options"]) > M:
262
+ out[i] = _score_large(model, tok, it, batch_size)
263
+ if not small:
264
+ return out
265
+ sp = fc["special"]
266
+ prepared = {}
267
+ for i in small:
268
+ it = items[i]
269
+ if len(it["options"]) < 2:
270
+ raise ValueError("Every question needs at least two options.")
271
+ budget = int(it.get("option_tokens", fc["max_tok_per_opt"]))
272
+ q_ids = tok(it.get("question", ""), add_special_tokens=False, truncation=True, max_length=96)["input_ids"]
273
+ o_ids = tok([str(o) for o in it["options"]], add_special_tokens=False, truncation=True,
274
+ max_length=budget)["input_ids"]
275
+ s = _state_text(it.get("state"))
276
+ s_ids = tok(s, add_special_tokens=False, truncation=True,
277
+ max_length=int(fc["long_max_len"]))["input_ids"] if s else []
278
+ ids, mk = assemble(q_ids, o_ids, s_ids, sp, int(it.get("seq_len", fc["max_len"])), fc["head_max_len"],
279
+ budget, fc["long_max_len"], fc["long_opts_threshold"])
280
+ prepared[i] = (ids, mk, QTYPES[it.get("type", "choice")])
281
+ order = sorted(small, key=lambda i: len(prepared[i][0]))
282
+ model.eval()
283
+ for b0 in range(0, len(order), batch_size):
284
+ idxs = order[b0: b0 + batch_size]
285
+ batch = collate_features([prepared[i] for i in idxs], sp["pad"], model.device)
286
+ with _amp(model):
287
+ logits = model(**batch)
288
+ for j, i in enumerate(idxs):
289
+ n = len(prepared[i][1])
290
+ T = model.temperature[prepared[i][2], n_bucket(n)].float().clamp_min(1e-3)
291
+ out[i] = torch.softmax(logits[j, :n].float() / T, -1).cpu().numpy()
292
+ return out
293
+
294
+
295
+ def _score_large(model, tok, it, batch_size):
296
+ """Tournament for > max_opts_single_pass options: keep the best of each chunk, then one final pass."""
297
+ M = int(model.fcfg.get("max_opts_single_pass", 96))
298
+ opts = list(it["options"])
299
+ cand = list(range(len(opts)))
300
+ while len(cand) > M:
301
+ groups = [cand[i: i + M] for i in range(0, len(cand), M)]
302
+ keep = max(1, M // len(groups))
303
+ ps = score_items(model, tok, [dict(it, options=[opts[c] for c in g]) for g in groups], batch_size)
304
+ cand = [g[t] for g, p in zip(groups, ps) for t in np.argsort(-p)[:keep]]
305
+ p = score_items(model, tok, [dict(it, options=[opts[c] for c in cand])], batch_size)[0]
306
+ full = np.zeros(len(opts), dtype=np.float32)
307
+ full[cand] = p
308
+ return full
309
+
310
+
311
+ def _normalize_question(q):
312
+ qtype = q.get("type", "choice")
313
+ text = q.get("question") or q.get("instructions") or ""
314
+ if qtype == "noul":
315
+ lab = q.get("labels") or {}
316
+ return qtype, text, [True, False], [str(lab.get("true", "Yes")), str(lab.get("false", "No"))]
317
+ crit = q.get("criteria", q.get("options"))
318
+ if qtype == "score":
319
+ opts = [str(c) for c in crit]
320
+ return qtype, text, list(range(len(opts))), opts
321
+ if isinstance(crit, dict):
322
+ return qtype, text, list(crit), [f"{k}: {v}" if v else str(k) for k, v in crit.items()]
323
+ return qtype, text, list(crit), [str(c) for c in crit]
324
+
325
+
326
+ @torch.no_grad()
327
+ def decide(model, tok, state, questions, defer_threshold=None, batch_size=32):
328
+ """Answer typed questions about one state.
329
+
330
+ questions: list of dicts, or a Jev/Laya-style dict {key: question}. Each question:
331
+ {"type": "choice"|"noul"|"score", "question"/"instructions": str,
332
+ "options": [...] or "criteria": {key: description} / [levels], "option_tokens"?: int}
333
+ """
334
+ if isinstance(questions, dict):
335
+ questions = [dict(q, key=k) for k, q in questions.items()]
336
+ items, metas = [], []
337
+ for q in questions:
338
+ qtype, text, keys, opts = _normalize_question(q)
339
+ item = {"state": state, "question": text, "options": opts, "type": qtype}
340
+ for extra in ("option_tokens", "seq_len"):
341
+ if extra in q:
342
+ item[extra] = q[extra]
343
+ items.append(item)
344
+ metas.append((q, qtype, keys, opts))
345
+ probs = score_items(model, tok, items, batch_size)
346
+ thr = model.fcfg.get("defer_threshold", 0.0) if defer_threshold is None else defer_threshold
347
+ results = []
348
+ for (q, qtype, keys, opts), p in zip(metas, probs):
349
+ i = int(np.argmax(p))
350
+ r = {"key": q.get("key"), "type": qtype, "question": items[len(results)]["question"],
351
+ "choice": keys[i], "choice_text": opts[i], "confidence": float(p[i]),
352
+ "probs": {str(k): float(v) for k, v in zip(keys, p)}, "defer": float(p[i]) < thr}
353
+ if qtype == "noul":
354
+ r["p_true"] = float(p[0])
355
+ if qtype == "score":
356
+ r["expected_level"] = float(np.dot(p, np.arange(len(p))))
357
+ results.append(r)
358
+ return {"results": results, "answers": {r["key"]: r for r in results if r["key"] is not None}}
falcondec_report.json ADDED
@@ -0,0 +1,901 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "version": "1.0.1",
3
+ "notebook_version": "V2",
4
+ "mode": "finetune",
5
+ "parent_version": "1.0.0",
6
+ "lineage": [
7
+ "jhu-clsp/ettin-encoder-150m",
8
+ "FalconDec-v1.0.0"
9
+ ],
10
+ "config": {
11
+ "MODE": "finetune",
12
+ "FINETUNE_FROM": "falcondec_runs/FalconDec-v1.0.0",
13
+ "MODEL_NAME": "LightDec",
14
+ "BACKBONE": "jhu-clsp/ettin-encoder-150m",
15
+ "BACKBONE_INIT": "pretrained",
16
+ "PRESET": "full",
17
+ "CUSTOM_DATA_JSONL": "",
18
+ "CUSTOM_WEIGHT": 3.0,
19
+ "INCLUDE_BUILDERS": [],
20
+ "EXCLUDE_BUILDERS": [
21
+ "mind2web"
22
+ ],
23
+ "MAX_LEN": 512,
24
+ "LONG_MAX_LEN": 2048,
25
+ "LONG_OPTS_THRESHOLD": 24,
26
+ "HEAD_MAX_LEN": 192,
27
+ "MAX_TOK_PER_OPT": 24,
28
+ "MAX_OPTS_SINGLE_PASS": 96,
29
+ "INTERACT_LAYERS": 2,
30
+ "INTERACT_HEADS": 8,
31
+ "BATCH_SIZE": 32,
32
+ "TOKENS_PER_BATCH": 16384,
33
+ "GRAD_ACCUM": 1,
34
+ "LR_ENCODER": 4e-05,
35
+ "LR_HEAD": 0.0003,
36
+ "LLRD": 0.9,
37
+ "WEIGHT_DECAY": 0.01,
38
+ "WARMUP_FRAC": 0.06,
39
+ "EPOCHS": 0,
40
+ "EMA_DECAY": 0.999,
41
+ "GRAD_CLIP": 1.0,
42
+ "FINETUNE_LR_SCALE": 0.5,
43
+ "SPHERICAL_W": 0.5,
44
+ "BRIER_W": 0.0,
45
+ "RPS_W": 1.0,
46
+ "TASK_SAMPLING_ALPHA": 0.5,
47
+ "NOTA_PROB": 0.08,
48
+ "RLCD_EPOCHS": 0.0,
49
+ "RLCD_SAMPLES": 4,
50
+ "RLCD_SIGMA": 0.3,
51
+ "RLCD_CE_ANCHOR": 0.5,
52
+ "DEVICE": "cuda",
53
+ "SEED": 42,
54
+ "NUM_WORKERS": 0,
55
+ "USE_COMPILE": false,
56
+ "ATTN_IMPL": "auto",
57
+ "OUTPUT_ROOT": "./falcondec_runs",
58
+ "EXPORT_INT8": true,
59
+ "SIZE_TARGET_MB": 250,
60
+ "DEFER_THRESHOLD": 0.7,
61
+ "EVAL_BASELINE_PROOF_V2": true,
62
+ "BASELINE_MAX_PER_TASK": 120,
63
+ "RUN_CPU_BENCH": true,
64
+ "PUSH_TO_HUB": false,
65
+ "HUB_REPO": ""
66
+ },
67
+ "epochs": 2,
68
+ "best_epoch": 1,
69
+ "rlcd_kept": false,
70
+ "history": [
71
+ {
72
+ "epoch": 1,
73
+ "train_loss": 0.47463753718974483,
74
+ "train_acc": 0.8307763076996187,
75
+ "val_macro": 0.8253118682436162,
76
+ "val_micro": 0.834065625876414,
77
+ "val_nll": 0.4055366800702441
78
+ },
79
+ {
80
+ "epoch": 2,
81
+ "train_loss": 0.36265113689414885,
82
+ "train_acc": 0.8738680252304237,
83
+ "val_macro": 0.8252539304231531,
84
+ "val_micro": 0.8348602411891184,
85
+ "val_nll": 0.4680047751477799
86
+ }
87
+ ],
88
+ "train_minutes": 172.17712485790253,
89
+ "temperatures": [
90
+ [
91
+ 1.3723914623260498,
92
+ 1.3805755376815796,
93
+ 1.5033340454101562,
94
+ 1.320592999458313
95
+ ],
96
+ [
97
+ 1.3903899192810059,
98
+ 1.3903899192810059,
99
+ 1.3903899192810059,
100
+ 1.3903899192810059
101
+ ],
102
+ [
103
+ 1.4738539457321167,
104
+ 1.4738539457321167,
105
+ 1.4738539457321167,
106
+ 1.4738539457321167
107
+ ]
108
+ ],
109
+ "data_counts": {
110
+ "train": 946028,
111
+ "validation": 21394,
112
+ "test": 26597
113
+ },
114
+ "skipped_builders": [],
115
+ "test_overall": {
116
+ "n": 26597.0,
117
+ "micro_acc": 0.7810655337068091,
118
+ "macro_acc": 0.7694407230754166,
119
+ "nll": 0.5288964978022598,
120
+ "brier": 0.2897912847111225,
121
+ "ece": 0.01969622960708085,
122
+ "aurc": 0.06403486126686972,
123
+ "score_mae": 0.4928175845338412,
124
+ "coverage@0.7": 0.6595104711057638,
125
+ "acc_on_covered@0.7": 0.9167094236360527
126
+ },
127
+ "test_per_domain": {
128
+ "agentic": 0.7781070271610528,
129
+ "classification": 0.615,
130
+ "code": 0.8717717677968562,
131
+ "guardrails": 0.7951432854293641,
132
+ "intents": 0.8278333333333333,
133
+ "policy": 0.7873333333333333,
134
+ "reasoning": 0.7006267469170804,
135
+ "support": 1.0,
136
+ "workflows": 0.755
137
+ },
138
+ "test_per_task": [
139
+ {
140
+ "task": "ag_news/topic",
141
+ "n": 500,
142
+ "acc": 0.888,
143
+ "chance": 0.25,
144
+ "nll": 0.28043824258210953,
145
+ "brier": 0.154232526432949,
146
+ "heldout": false,
147
+ "domain": "classification",
148
+ "ece": 0.03968544936180117,
149
+ "proof_v2 (card)": NaN,
150
+ "\u0394 vs proof_v2": NaN
151
+ },
152
+ {
153
+ "task": "agentharm/refuse",
154
+ "n": 416,
155
+ "acc": 0.7427884615384616,
156
+ "chance": 0.5,
157
+ "nll": 0.5331332034928402,
158
+ "brier": 0.35988097897984717,
159
+ "heldout": true,
160
+ "domain": "guardrails",
161
+ "ece": 0.06706968007179408,
162
+ "proof_v2 (card)": NaN,
163
+ "\u0394 vs proof_v2": NaN
164
+ },
165
+ {
166
+ "task": "anli/nli",
167
+ "n": 498,
168
+ "acc": 0.41566265060240964,
169
+ "chance": 0.33333333333333326,
170
+ "nll": 1.1498472186171984,
171
+ "brier": 0.7016487136039561,
172
+ "heldout": false,
173
+ "domain": "reasoning",
174
+ "ece": 0.18407419437624845,
175
+ "proof_v2 (card)": NaN,
176
+ "\u0394 vs proof_v2": NaN
177
+ },
178
+ {
179
+ "task": "aqua_rat/math",
180
+ "n": 247,
181
+ "acc": 0.27530364372469635,
182
+ "chance": 0.2,
183
+ "nll": 1.6108178135746645,
184
+ "brier": 0.7991190068538744,
185
+ "heldout": false,
186
+ "domain": "reasoning",
187
+ "ece": 0.04157990880823329,
188
+ "proof_v2 (card)": NaN,
189
+ "\u0394 vs proof_v2": NaN
190
+ },
191
+ {
192
+ "task": "arc_challenge/mcq",
193
+ "n": 500,
194
+ "acc": 0.46,
195
+ "chance": 0.25006666666666666,
196
+ "nll": 1.2087727874126286,
197
+ "brier": 0.653150732152712,
198
+ "heldout": true,
199
+ "domain": "reasoning",
200
+ "ece": 0.08824656188488007,
201
+ "proof_v2 (card)": 0.308,
202
+ "\u0394 vs proof_v2": 0.15200000000000002
203
+ },
204
+ {
205
+ "task": "arc_easy/mcq",
206
+ "n": 500,
207
+ "acc": 0.594,
208
+ "chance": 0.24980000000000002,
209
+ "nll": 0.9366558539052494,
210
+ "brier": 0.5142523023860749,
211
+ "heldout": true,
212
+ "domain": "reasoning",
213
+ "ece": 0.0377397992014885,
214
+ "proof_v2 (card)": 0.425,
215
+ "\u0394 vs proof_v2": 0.16899999999999998
216
+ },
217
+ {
218
+ "task": "banking77/intent",
219
+ "n": 500,
220
+ "acc": 0.93,
221
+ "chance": 0.3214666666666666,
222
+ "nll": 0.21486007268911725,
223
+ "brier": 0.11047492184269238,
224
+ "heldout": true,
225
+ "domain": "intents",
226
+ "ece": 0.04031616294384006,
227
+ "proof_v2 (card)": 0.883,
228
+ "\u0394 vs proof_v2": 0.04700000000000004
229
+ },
230
+ {
231
+ "task": "banking77/intent_77",
232
+ "n": 300,
233
+ "acc": 0.5333333333333333,
234
+ "chance": 0.012987012987012986,
235
+ "nll": 2.063840282613334,
236
+ "brier": 0.6459762259581862,
237
+ "heldout": true,
238
+ "domain": "intents",
239
+ "ece": 0.14717036068439482,
240
+ "proof_v2 (card)": NaN,
241
+ "\u0394 vs proof_v2": NaN
242
+ },
243
+ {
244
+ "task": "bigclonebench/clone",
245
+ "n": 500,
246
+ "acc": 0.952,
247
+ "chance": 0.5,
248
+ "nll": 0.14842138955372502,
249
+ "brier": 0.07646761075289837,
250
+ "heldout": false,
251
+ "domain": "code",
252
+ "ece": 0.015586107730865495,
253
+ "proof_v2 (card)": 0.383,
254
+ "\u0394 vs proof_v2": 0.569
255
+ },
256
+ {
257
+ "task": "bitext/category",
258
+ "n": 500,
259
+ "acc": 1.0,
260
+ "chance": 0.16128124098124094,
261
+ "nll": 0.0018036427325441764,
262
+ "brier": 0.000553129204988908,
263
+ "heldout": false,
264
+ "domain": "support",
265
+ "ece": 0.0015759592056274496,
266
+ "proof_v2 (card)": NaN,
267
+ "\u0394 vs proof_v2": NaN
268
+ },
269
+ {
270
+ "task": "bitext/route",
271
+ "n": 500,
272
+ "acc": 1.0,
273
+ "chance": 0.19016313131313134,
274
+ "nll": 0.0011070070102664432,
275
+ "brier": 7.369536871875006e-05,
276
+ "heldout": false,
277
+ "domain": "support",
278
+ "ece": 0.0010853933095931502,
279
+ "proof_v2 (card)": 0.958,
280
+ "\u0394 vs proof_v2": 0.04200000000000004
281
+ },
282
+ {
283
+ "task": "boolq/yes_no",
284
+ "n": 500,
285
+ "acc": 0.816,
286
+ "chance": 0.5,
287
+ "nll": 0.4159143784940243,
288
+ "brier": 0.26228454396614875,
289
+ "heldout": false,
290
+ "domain": "reasoning",
291
+ "ece": 0.02717309403419496,
292
+ "proof_v2 (card)": 0.717,
293
+ "\u0394 vs proof_v2": 0.09899999999999998
294
+ },
295
+ {
296
+ "task": "civil_comments/toxic",
297
+ "n": 500,
298
+ "acc": 0.906,
299
+ "chance": 0.5,
300
+ "nll": 0.2400463573904126,
301
+ "brier": 0.13785060332637392,
302
+ "heldout": false,
303
+ "domain": "guardrails",
304
+ "ece": 0.035369446873664855,
305
+ "proof_v2 (card)": NaN,
306
+ "\u0394 vs proof_v2": NaN
307
+ },
308
+ {
309
+ "task": "clinc150/intent",
310
+ "n": 500,
311
+ "acc": 0.942,
312
+ "chance": 0.12404474984829318,
313
+ "nll": 0.16339450041492637,
314
+ "brier": 0.08445263661020003,
315
+ "heldout": false,
316
+ "domain": "intents",
317
+ "ece": 0.036179060459136984,
318
+ "proof_v2 (card)": 0.85,
319
+ "\u0394 vs proof_v2": 0.09199999999999997
320
+ },
321
+ {
322
+ "task": "codexglue/code_to_doc",
323
+ "n": 504,
324
+ "acc": 0.9841269841269841,
325
+ "chance": 0.26233465608465617,
326
+ "nll": 0.04367021227094591,
327
+ "brier": 0.02423435580505018,
328
+ "heldout": false,
329
+ "domain": "code",
330
+ "ece": 0.011909238107147682,
331
+ "proof_v2 (card)": 0.969,
332
+ "\u0394 vs proof_v2": 0.0151269841269841
333
+ },
334
+ {
335
+ "task": "codexglue/doc_to_code",
336
+ "n": 504,
337
+ "acc": 0.9761904761904762,
338
+ "chance": 0.2786044973544974,
339
+ "nll": 0.05729491385996503,
340
+ "brier": 0.030591094765289834,
341
+ "heldout": false,
342
+ "domain": "code",
343
+ "ece": 0.01021853999959103,
344
+ "proof_v2 (card)": 0.961,
345
+ "\u0394 vs proof_v2": 0.015190476190476199
346
+ },
347
+ {
348
+ "task": "codexglue/func_name",
349
+ "n": 467,
350
+ "acc": 0.9528907922912205,
351
+ "chance": 0.25499643112062814,
352
+ "nll": 0.17043053883545994,
353
+ "brier": 0.08725536960997705,
354
+ "heldout": false,
355
+ "domain": "code",
356
+ "ece": 0.023257157468183163,
357
+ "proof_v2 (card)": 0.901,
358
+ "\u0394 vs proof_v2": 0.05189079229122051
359
+ },
360
+ {
361
+ "task": "codexglue/lang_id",
362
+ "n": 504,
363
+ "acc": 1.0,
364
+ "chance": 0.2354166666666667,
365
+ "nll": 0.0005584346231148719,
366
+ "brier": 3.0963766792822474e-06,
367
+ "heldout": false,
368
+ "domain": "code",
369
+ "ece": 0.0005574606004214999,
370
+ "proof_v2 (card)": 0.997,
371
+ "\u0394 vs proof_v2": 0.0030000000000000027
372
+ },
373
+ {
374
+ "task": "commonsense_qa/mcq",
375
+ "n": 493,
376
+ "acc": 0.6389452332657201,
377
+ "chance": 0.19999999999999998,
378
+ "nll": 0.904400206106629,
379
+ "brier": 0.47107552904904443,
380
+ "heldout": false,
381
+ "domain": "reasoning",
382
+ "ece": 0.04405802657589711,
383
+ "proof_v2 (card)": 0.442,
384
+ "\u0394 vs proof_v2": 0.19694523326572005
385
+ },
386
+ {
387
+ "task": "counsel/critique_quality",
388
+ "n": 201,
389
+ "acc": 0.5920398009950248,
390
+ "chance": 0.3333333333333333,
391
+ "nll": 1.3054839486733487,
392
+ "brier": 0.6194979652829024,
393
+ "heldout": false,
394
+ "domain": "agentic",
395
+ "ece": 0.22704330650135063,
396
+ "proof_v2 (card)": NaN,
397
+ "\u0394 vs proof_v2": NaN
398
+ },
399
+ {
400
+ "task": "counsel/step_has_error",
401
+ "n": 201,
402
+ "acc": 0.8009950248756219,
403
+ "chance": 0.5,
404
+ "nll": 0.8227317674037697,
405
+ "brier": 0.3473574496250388,
406
+ "heldout": false,
407
+ "domain": "agentic",
408
+ "ece": 0.16924872979595879,
409
+ "proof_v2 (card)": NaN,
410
+ "\u0394 vs proof_v2": NaN
411
+ },
412
+ {
413
+ "task": "devign/vulnerability",
414
+ "n": 500,
415
+ "acc": 0.596,
416
+ "chance": 0.5,
417
+ "nll": 0.6676419580578804,
418
+ "brier": 0.4753553111644763,
419
+ "heldout": false,
420
+ "domain": "code",
421
+ "ece": 0.05009425449371341,
422
+ "proof_v2 (card)": 0.542,
423
+ "\u0394 vs proof_v2": 0.05399999999999994
424
+ },
425
+ {
426
+ "task": "emotion/6way",
427
+ "n": 500,
428
+ "acc": 0.502,
429
+ "chance": 0.16666666666666663,
430
+ "nll": 1.3070907976925372,
431
+ "brier": 0.6468766489597455,
432
+ "heldout": true,
433
+ "domain": "classification",
434
+ "ece": 0.08215402710437775,
435
+ "proof_v2 (card)": NaN,
436
+ "\u0394 vs proof_v2": NaN
437
+ },
438
+ {
439
+ "task": "gsm8k/math",
440
+ "n": 500,
441
+ "acc": 0.662,
442
+ "chance": 0.25,
443
+ "nll": 0.7718028125888668,
444
+ "brier": 0.4427847443219045,
445
+ "heldout": false,
446
+ "domain": "reasoning",
447
+ "ece": 0.06086966800689696,
448
+ "proof_v2 (card)": 0.275,
449
+ "\u0394 vs proof_v2": 0.387
450
+ },
451
+ {
452
+ "task": "hellaswag/continuation",
453
+ "n": 500,
454
+ "acc": 0.58,
455
+ "chance": 0.25,
456
+ "nll": 0.9953804229423404,
457
+ "brier": 0.5371978334623486,
458
+ "heldout": false,
459
+ "domain": "reasoning",
460
+ "ece": 0.07333891946077346,
461
+ "proof_v2 (card)": NaN,
462
+ "\u0394 vs proof_v2": NaN
463
+ },
464
+ {
465
+ "task": "hotpotqa/comparison_yes_no",
466
+ "n": 26,
467
+ "acc": 0.8461538461538461,
468
+ "chance": 0.5,
469
+ "nll": 0.43042818936877525,
470
+ "brier": 0.2727966735853732,
471
+ "heldout": false,
472
+ "domain": "agentic",
473
+ "ece": 0.15112700829139125,
474
+ "proof_v2 (card)": NaN,
475
+ "\u0394 vs proof_v2": NaN
476
+ },
477
+ {
478
+ "task": "hotpotqa/retrieve",
479
+ "n": 497,
480
+ "acc": 0.8732394366197183,
481
+ "chance": 0.16741720162243304,
482
+ "nll": 0.4343729186731075,
483
+ "brier": 0.20390003436957638,
484
+ "heldout": false,
485
+ "domain": "agentic",
486
+ "ece": 0.08316291382974779,
487
+ "proof_v2 (card)": NaN,
488
+ "\u0394 vs proof_v2": NaN
489
+ },
490
+ {
491
+ "task": "humaneval/completion",
492
+ "n": 119,
493
+ "acc": 0.6722689075630253,
494
+ "chance": 0.39355742296918783,
495
+ "nll": 0.7120104984726616,
496
+ "brier": 0.44585727429903876,
497
+ "heldout": true,
498
+ "domain": "code",
499
+ "ece": 0.10659772783768277,
500
+ "proof_v2 (card)": 0.575,
501
+ "\u0394 vs proof_v2": 0.09726890756302531
502
+ },
503
+ {
504
+ "task": "jailbreak/detect",
505
+ "n": 262,
506
+ "acc": 0.9541984732824428,
507
+ "chance": 0.5,
508
+ "nll": 0.10692274935157883,
509
+ "brier": 0.06256811361157136,
510
+ "heldout": false,
511
+ "domain": "guardrails",
512
+ "ece": 0.0232666851455019,
513
+ "proof_v2 (card)": NaN,
514
+ "\u0394 vs proof_v2": NaN
515
+ },
516
+ {
517
+ "task": "massive_en/intent",
518
+ "n": 500,
519
+ "acc": 0.906,
520
+ "chance": 0.12958727789085603,
521
+ "nll": 0.26250843075065133,
522
+ "brier": 0.12679194471370903,
523
+ "heldout": false,
524
+ "domain": "intents",
525
+ "ece": 0.022001883208751682,
526
+ "proof_v2 (card)": NaN,
527
+ "\u0394 vs proof_v2": NaN
528
+ },
529
+ {
530
+ "task": "mbpp/bugspot",
531
+ "n": 256,
532
+ "acc": 0.73046875,
533
+ "chance": 0.40559895833333326,
534
+ "nll": 0.5201096873639273,
535
+ "brier": 0.3294220878204956,
536
+ "heldout": false,
537
+ "domain": "code",
538
+ "ece": 0.1015084749087691,
539
+ "proof_v2 (card)": 0.475,
540
+ "\u0394 vs proof_v2": 0.25546875
541
+ },
542
+ {
543
+ "task": "mbpp/solution",
544
+ "n": 500,
545
+ "acc": 0.982,
546
+ "chance": 0.25,
547
+ "nll": 0.06299449819558049,
548
+ "brier": 0.03283427225700136,
549
+ "heldout": false,
550
+ "domain": "code",
551
+ "ece": 0.014863664209842696,
552
+ "proof_v2 (card)": 0.992,
553
+ "\u0394 vs proof_v2": -0.010000000000000009
554
+ },
555
+ {
556
+ "task": "mmlu/mcq",
557
+ "n": 500,
558
+ "acc": 0.356,
559
+ "chance": 0.25,
560
+ "nll": 1.370702707901597,
561
+ "brier": 0.738678903692766,
562
+ "heldout": true,
563
+ "domain": "reasoning",
564
+ "ece": 0.10921649205684662,
565
+ "proof_v2 (card)": NaN,
566
+ "\u0394 vs proof_v2": NaN
567
+ },
568
+ {
569
+ "task": "mnli/claim",
570
+ "n": 500,
571
+ "acc": 0.838,
572
+ "chance": 0.33333333333333326,
573
+ "nll": 0.41442326370440424,
574
+ "brier": 0.23035331540149503,
575
+ "heldout": false,
576
+ "domain": "reasoning",
577
+ "ece": 0.07134601444005968,
578
+ "proof_v2 (card)": 0.492,
579
+ "\u0394 vs proof_v2": 0.346
580
+ },
581
+ {
582
+ "task": "openbookqa/mcq",
583
+ "n": 500,
584
+ "acc": 0.61,
585
+ "chance": 0.25,
586
+ "nll": 1.024268753977958,
587
+ "brier": 0.5355113333378367,
588
+ "heldout": false,
589
+ "domain": "reasoning",
590
+ "ece": 0.09224175179004669,
591
+ "proof_v2 (card)": 0.292,
592
+ "\u0394 vs proof_v2": 0.318
593
+ },
594
+ {
595
+ "task": "policy/access_control_transfer",
596
+ "n": 500,
597
+ "acc": 1.0,
598
+ "chance": 0.33333333333333326,
599
+ "nll": 0.0002645651292309594,
600
+ "brier": 1.1635768837216181e-06,
601
+ "heldout": false,
602
+ "domain": "policy",
603
+ "ece": 0.00026422095298772597,
604
+ "proof_v2 (card)": NaN,
605
+ "\u0394 vs proof_v2": NaN
606
+ },
607
+ {
608
+ "task": "policy/count_threshold_transfer",
609
+ "n": 500,
610
+ "acc": 0.846,
611
+ "chance": 0.17848809523809525,
612
+ "nll": 0.4453474920627887,
613
+ "brier": 0.239936219050859,
614
+ "heldout": false,
615
+ "domain": "policy",
616
+ "ece": 0.10926523756980895,
617
+ "proof_v2 (card)": NaN,
618
+ "\u0394 vs proof_v2": NaN
619
+ },
620
+ {
621
+ "task": "policy/free_shipping_transfer",
622
+ "n": 500,
623
+ "acc": 0.766,
624
+ "chance": 0.5,
625
+ "nll": 0.39174148760084065,
626
+ "brier": 0.2747179012451008,
627
+ "heldout": false,
628
+ "domain": "policy",
629
+ "ece": 0.10137953925132755,
630
+ "proof_v2 (card)": NaN,
631
+ "\u0394 vs proof_v2": NaN
632
+ },
633
+ {
634
+ "task": "policy/invoice_overdue_transfer",
635
+ "n": 500,
636
+ "acc": 0.6,
637
+ "chance": 0.5,
638
+ "nll": 0.8301094107693061,
639
+ "brier": 0.5458620193712137,
640
+ "heldout": false,
641
+ "domain": "policy",
642
+ "ece": 0.20100946557521818,
643
+ "proof_v2 (card)": NaN,
644
+ "\u0394 vs proof_v2": NaN
645
+ },
646
+ {
647
+ "task": "policy/invoice_total_transfer",
648
+ "n": 500,
649
+ "acc": 0.486,
650
+ "chance": 0.5,
651
+ "nll": 0.69399339812994,
652
+ "brier": 0.5008197939882997,
653
+ "heldout": false,
654
+ "domain": "policy",
655
+ "ece": 0.03210310280323031,
656
+ "proof_v2 (card)": NaN,
657
+ "\u0394 vs proof_v2": NaN
658
+ },
659
+ {
660
+ "task": "policy/refund_approval_transfer",
661
+ "n": 500,
662
+ "acc": 0.78,
663
+ "chance": 0.33333333333333326,
664
+ "nll": 0.8478962361293161,
665
+ "brier": 0.38071118861218445,
666
+ "heldout": false,
667
+ "domain": "policy",
668
+ "ece": 0.18587953794002532,
669
+ "proof_v2 (card)": NaN,
670
+ "\u0394 vs proof_v2": NaN
671
+ },
672
+ {
673
+ "task": "policy/return_window_transfer",
674
+ "n": 500,
675
+ "acc": 0.998,
676
+ "chance": 0.33333333333333326,
677
+ "nll": 0.02459742584443666,
678
+ "brier": 0.005749234483560448,
679
+ "heldout": false,
680
+ "domain": "policy",
681
+ "ece": 0.019777906179428068,
682
+ "proof_v2 (card)": NaN,
683
+ "\u0394 vs proof_v2": NaN
684
+ },
685
+ {
686
+ "task": "policy/sla_urgency_transfer",
687
+ "n": 500,
688
+ "acc": 0.698,
689
+ "chance": 0.25,
690
+ "nll": 0.9843494065143168,
691
+ "brier": 0.49182705060356263,
692
+ "heldout": false,
693
+ "domain": "policy",
694
+ "ece": 0.21234694278240196,
695
+ "proof_v2 (card)": NaN,
696
+ "\u0394 vs proof_v2": NaN
697
+ },
698
+ {
699
+ "task": "policy/table_extreme_transfer",
700
+ "n": 500,
701
+ "acc": 0.912,
702
+ "chance": 0.23716666666666666,
703
+ "nll": 0.22187845180515342,
704
+ "brier": 0.12271458379265106,
705
+ "heldout": false,
706
+ "domain": "policy",
707
+ "ece": 0.013898970842361484,
708
+ "proof_v2 (card)": NaN,
709
+ "\u0394 vs proof_v2": NaN
710
+ },
711
+ {
712
+ "task": "prompt_injections/detect",
713
+ "n": 116,
714
+ "acc": 0.5775862068965517,
715
+ "chance": 0.5,
716
+ "nll": 1.0671347150446622,
717
+ "brier": 0.628204618498277,
718
+ "heldout": true,
719
+ "domain": "guardrails",
720
+ "ece": 0.2859799378904803,
721
+ "proof_v2 (card)": NaN,
722
+ "\u0394 vs proof_v2": NaN
723
+ },
724
+ {
725
+ "task": "qasc/mcq",
726
+ "n": 500,
727
+ "acc": 0.99,
728
+ "chance": 0.125,
729
+ "nll": 0.0315895584258742,
730
+ "brier": 0.01700795004425532,
731
+ "heldout": false,
732
+ "domain": "reasoning",
733
+ "ece": 0.006339998722076429,
734
+ "proof_v2 (card)": NaN,
735
+ "\u0394 vs proof_v2": NaN
736
+ },
737
+ {
738
+ "task": "sciq/mcq",
739
+ "n": 498,
740
+ "acc": 0.9538152610441767,
741
+ "chance": 0.25,
742
+ "nll": 0.11897318925779031,
743
+ "brier": 0.06399753046801272,
744
+ "heldout": false,
745
+ "domain": "reasoning",
746
+ "ece": 0.03059945360244999,
747
+ "proof_v2 (card)": 0.692,
748
+ "\u0394 vs proof_v2": 0.26181526104417674
749
+ },
750
+ {
751
+ "task": "scitail/support",
752
+ "n": 500,
753
+ "acc": 0.964,
754
+ "chance": 0.5,
755
+ "nll": 0.10391226119734347,
756
+ "brier": 0.05616228693429717,
757
+ "heldout": false,
758
+ "domain": "reasoning",
759
+ "ece": 0.020327572941780014,
760
+ "proof_v2 (card)": NaN,
761
+ "\u0394 vs proof_v2": NaN
762
+ },
763
+ {
764
+ "task": "snli/contradicts",
765
+ "n": 500,
766
+ "acc": 0.986,
767
+ "chance": 0.33333333333333326,
768
+ "nll": 0.048629490838502536,
769
+ "brier": 0.024212671137030826,
770
+ "heldout": false,
771
+ "domain": "reasoning",
772
+ "ece": 0.01413622707128527,
773
+ "proof_v2 (card)": 0.892,
774
+ "\u0394 vs proof_v2": 0.09399999999999997
775
+ },
776
+ {
777
+ "task": "snli/must_be_true",
778
+ "n": 500,
779
+ "acc": 0.98,
780
+ "chance": 0.33333333333333326,
781
+ "nll": 0.060363937412155795,
782
+ "brier": 0.02947848703942969,
783
+ "heldout": false,
784
+ "domain": "reasoning",
785
+ "ece": 0.02905823123455048,
786
+ "proof_v2 (card)": 0.908,
787
+ "\u0394 vs proof_v2": 0.07199999999999995
788
+ },
789
+ {
790
+ "task": "snli/nli",
791
+ "n": 988,
792
+ "acc": 0.8795546558704453,
793
+ "chance": 0.33333333333333326,
794
+ "nll": 0.3359528171680472,
795
+ "brier": 0.18385021007121183,
796
+ "heldout": false,
797
+ "domain": "reasoning",
798
+ "ece": 0.04961462647688053,
799
+ "proof_v2 (card)": NaN,
800
+ "\u0394 vs proof_v2": NaN
801
+ },
802
+ {
803
+ "task": "sst5/score",
804
+ "n": 500,
805
+ "acc": 0.416,
806
+ "chance": 0.2,
807
+ "nll": 1.246531030535698,
808
+ "brier": 0.6734807272262489,
809
+ "heldout": true,
810
+ "domain": "classification",
811
+ "ece": 0.04567561271786689,
812
+ "proof_v2 (card)": NaN,
813
+ "\u0394 vs proof_v2": NaN
814
+ },
815
+ {
816
+ "task": "typed_decisions/agent_trace_observability",
817
+ "n": 500,
818
+ "acc": 0.712,
819
+ "chance": 0.3,
820
+ "nll": 0.7305186161920428,
821
+ "brier": 0.42510022971515116,
822
+ "heldout": false,
823
+ "domain": "workflows",
824
+ "ece": 0.11957757359743115,
825
+ "proof_v2 (card)": NaN,
826
+ "\u0394 vs proof_v2": NaN
827
+ },
828
+ {
829
+ "task": "typed_decisions/customer_service",
830
+ "n": 500,
831
+ "acc": 0.74,
832
+ "chance": 0.28,
833
+ "nll": 0.6714302012249828,
834
+ "brier": 0.3723593006177565,
835
+ "heldout": false,
836
+ "domain": "workflows",
837
+ "ece": 0.08548467075824738,
838
+ "proof_v2 (card)": NaN,
839
+ "\u0394 vs proof_v2": NaN
840
+ },
841
+ {
842
+ "task": "typed_decisions/invoice_processing",
843
+ "n": 500,
844
+ "acc": 0.838,
845
+ "chance": 0.35,
846
+ "nll": 0.5010932807754725,
847
+ "brier": 0.26905854083557984,
848
+ "heldout": false,
849
+ "domain": "workflows",
850
+ "ece": 0.12220623338222505,
851
+ "proof_v2 (card)": NaN,
852
+ "\u0394 vs proof_v2": NaN
853
+ },
854
+ {
855
+ "task": "typed_decisions/security_incidents",
856
+ "n": 500,
857
+ "acc": 0.73,
858
+ "chance": 0.34,
859
+ "nll": 0.6859722163826227,
860
+ "brier": 0.3929527039537006,
861
+ "heldout": false,
862
+ "domain": "workflows",
863
+ "ece": 0.143579844892025,
864
+ "proof_v2 (card)": NaN,
865
+ "\u0394 vs proof_v2": NaN
866
+ },
867
+ {
868
+ "task": "winogrande/blank",
869
+ "n": 500,
870
+ "acc": 0.612,
871
+ "chance": 0.5,
872
+ "nll": 0.7296091562472283,
873
+ "brier": 0.508893263835874,
874
+ "heldout": false,
875
+ "domain": "reasoning",
876
+ "ece": 0.11118709647655486,
877
+ "proof_v2 (card)": NaN,
878
+ "\u0394 vs proof_v2": NaN
879
+ },
880
+ {
881
+ "task": "yelp/score",
882
+ "n": 500,
883
+ "acc": 0.654,
884
+ "chance": 0.2,
885
+ "nll": 0.7885563414096832,
886
+ "brier": 0.4562026964016218,
887
+ "heldout": false,
888
+ "domain": "classification",
889
+ "ece": 0.06915212672948838,
890
+ "proof_v2 (card)": NaN,
891
+ "\u0394 vs proof_v2": NaN
892
+ }
893
+ ],
894
+ "head_to_head": null,
895
+ "env": {
896
+ "torch": "2.11.0+cu128",
897
+ "transformers": "5.17.0",
898
+ "python": "3.14.4",
899
+ "gpu": "NVIDIA GeForce RTX 5090 Laptop GPU"
900
+ }
901
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:80486c432d95a5b8b69f79a728bc8a3e42f12a86cb16abf2016b2a3076faa7fd
3
+ size 319325890
tokenizer/tokenizer_config.json CHANGED
@@ -2,9 +2,10 @@
2
  "backend": "tokenizers",
3
  "clean_up_tokenization_spaces": true,
4
  "cls_token": "[CLS]",
5
- "is_local": false,
6
  "local_files_only": false,
7
  "mask_token": "[MASK]",
 
8
  "model_input_names": [
9
  "input_ids",
10
  "attention_mask"
@@ -12,6 +13,9 @@
12
  "model_max_length": 8192,
13
  "pad_token": "[PAD]",
14
  "sep_token": "[SEP]",
 
15
  "tokenizer_class": "TokenizersBackend",
 
 
16
  "unk_token": "[UNK]"
17
  }
 
2
  "backend": "tokenizers",
3
  "clean_up_tokenization_spaces": true,
4
  "cls_token": "[CLS]",
5
+ "is_local": true,
6
  "local_files_only": false,
7
  "mask_token": "[MASK]",
8
+ "max_length": 256,
9
  "model_input_names": [
10
  "input_ids",
11
  "attention_mask"
 
13
  "model_max_length": 8192,
14
  "pad_token": "[PAD]",
15
  "sep_token": "[SEP]",
16
+ "stride": 0,
17
  "tokenizer_class": "TokenizersBackend",
18
+ "truncation_side": "right",
19
+ "truncation_strategy": "longest_first",
20
  "unk_token": "[UNK]"
21
  }