LightDec / manifest.json
mstatt's picture
Upload manifest.json with huggingface_hub
809c344 verified
Raw History Blame Contribute Delete
98.1 kB
{
"contract": "https://surgeon.falcons.ai/model-surgeon/manifest/v1",
"tool": "FALCONS.AI Model Surgeon V7.99",
"copyright": "\u00a9 2026 FALCONS.AI",
"exported": "2026-09-26T11:51:25Z",
"architecture": {
"family": "transformer_nlp",
"label": "NLP \u00b7 Small Language Model (SLM)",
"confidence": 0.98,
"score": 8.92,
"tags": [],
"runners_up": [
{
"family": "mlp",
"label": "MLP / Tabular",
"score": 2.5
},
{
"family": "ssm",
"label": "State Space Model (Mamba/S4)",
"score": 1.8
}
],
"total_params": 159654157
},
"intended_task": null,
"execution_plan_source": "name_heuristic",
"source_format": "safetensors",
"totals": {
"params": 159654157,
"bytes": 319308314,
"tensors": 168,
"modules": 226
},
"removed_nodes": [],
"reparented": [],
"grafted_components": [],
"structure_only_tensors_zero_filled": [],
"merged_tensors": [],
"lineage": {
"parents": [
{
"role": "primary",
"file": "Falconsai/LightDec/model.safetensors",
"format": "safetensors"
}
],
"merges": [],
"quantized": [],
"head_prunes": [],
"source": {
"kind": "hub",
"repo": "Falconsai/LightDec",
"license": "apache-2.0",
"file": "",
"config": null,
"card": {
"repo": "Falconsai/LightDec",
"revision": "main",
"text": "---\nlicense: apache-2.0\nlanguage:\n- en\nlibrary_name: transformers\nbase_model: jhu-clsp/ettin-encoder-150m\npipeline_tag: zero-shot-classification\ntags:\n- decision-model\n- system-one\n- falcondec\n- lightdec\n- calibrated-decisions\n- multiple-choice\n- intent-classification\n- customer-support\n- natural-language-inference\n- code\n- guardrails\n- agents\n- selective-prediction\n- falconsai\n---\n> Source model card: `Falconsai/LightDec` @ `main`, carried verbatim below. Its licence is the repository's. The Model Surgeon record follows it.\n\n\n\n# Falconsai/LightDec\n\n**[View in Model Surgeon](https://surgeon.falcons.ai/?hub=Falconsai/LightDec)**\n\n**A lightweight, single-pass, typed, calibrated decision model for agentic systems.** Give it a **state** (text, code or JSON), one or more **typed questions** (`choice`, `noul` yes/no, `score` ordinal) and a closed set of options. It returns a calibrated probability for every option, from one encoder pass per question.\n\nLightDec is the FalconDec architecture trained with **FalconDec notebook V2** on the **`standard`** data preset. That run adds agent-specific decisions (AgentTrek, Counsel, HotpotQA) to a balanced mix of 58 test tasks across 9 domains. This page is both the **model card** and the **developer guide**: how to load it, test it, and use it as a decision component in an agentic system.\n\n> **At a glance.** Test accuracy **0.725** (micro and task-macro) on 17,498 decisions from 58 tasks, with ECE **0.025**. At a 0.70 confidence threshold, LightDec answers **56%** of decisions at **89.6%** accuracy and defers the rest. Weights: **319 MB** fp16, **161 MB** int8. It is strongest on support routing, code understanding, intents, guardrails and agent-step checks, and weakest on multi-step arithmetic, date and table reasoning, and very wide label sets. Evaluate it on your own traffic (\u00a76.4) before acting on its answers.\n\n## Contents\n\n1. [Model summary](#1-model-summary)\n2. [Results](#2-results)\n3. [Intended and out-of-scope uses](#3-intended-and-out-of-scope-uses)\n4. [The Hub repository](#4-the-hub-repository)\n5. [Install, load and read a result](#5-install-load-and-read-a-result)\n6. [Testing the model](#6-testing-the-model)\n7. [Using it in an agentic system](#7-using-it-in-an-agentic-system)\n8. [Tuning for your domain](#8-tuning-for-your-domain)\n9. [Architecture](#9-architecture)\n10. [Training data](#10-training-data)\n11. [Training procedure and calibration](#11-training-procedure-and-calibration)\n12. [Operational notes](#12-operational-notes)\n13. [Bias, risks and limitations](#13-bias-risks-and-limitations)\n14. [Versioning and lineage](#14-versioning-and-lineage)\n15. [API reference](#15-api-reference)\n16. [Citation and references](#16-citation-and-references)\n\n---\n\n## 1. Model summary\n\n| | |\n|---|---|\n| **Model** | LightDec: the FalconDec architecture, notebook V2, `standard` preset. The checkpoint's own config reports `FalconDec` version `1.0.0` |\n| **Task** | Closed-set decisions: given a state, a question and 2\u2013N options, return a calibrated probability per option |\n| **Question types** | `choice` (pick one), `noul` (yes/no, returns P(true)), `score` (ordinal rubric, returns the expected level) |\n| **Backbone** | [`jhu-clsp/ettin-encoder-150m`](https://huggingface.co/jhu-clsp/ettin-encoder-150m) (ModernBERT-style encoder), fully fine-tuned |\n| **Decision head** | Option-marker scoring plus a permutation-equivariant set-transformer head (\u00a79) |\n| **Parameters** | \u2248160M (Hub reports 0.2B) |\n| **Weights** | fp16 **319 MB** (`model.safetensors`) \u00b7 per-channel int8 **161 MB** (`compact-int8/model_int8.safetensors`) |\n| **Context** | 512 tokens; automatically 2,048 for questions with more than 24 options; a tournament above 96 options |\n| **Calibration** | One temperature per (question type \u00d7 option-count bucket), stored in the checkpoint and applied automatically |\n| **Inference cost** | One encoder pass per question. The same architecture (`Falconsai/proof_v3`) measured 15.7 ms p50 for one question on a GPU; LightDec's own latency is not in its report (measure with \u00a76.3) |\n| **Output** | Probabilities, the chosen option, confidence, a `defer` flag, `p_true` (noul) and `expected_level` (score) |\n| **Language** | English, plus code in Python, Java, JavaScript, PHP, Ruby, Go and C |\n| **Custom code** | `falcondec_modeling.py` ships with the weights and holds the model and all inference logic. Load it with `importlib` (\u00a75); `AutoModel.from_pretrained` alone won't build the decision head |\n| **License** | Apache-2.0 for the weights and code. Check each training dataset's license before redistributing derived data |\n\nThe model has no generative component. It can only rank the options you give it, so it cannot produce text outside that set.\n\n---\n\n## 2. Results\n\nAll numbers come from this checkpoint's `falcondec_report.json`: one run, `standard` preset, 2 epochs, seed 42, `MODE=\"scratch\"`, notebook V2.\n\n### 2.1 At a glance\n\n| Metric | Value |\n|---|---|\n| Test decisions / tasks | 17,498 / 58 (10 of them held out) |\n| Test accuracy, micro | **0.725** |\n| Test accuracy, task-macro | **0.725** |\n| Held-out tasks, task-macro (10 tasks never trained on) | **0.567** |\n| Expected calibration error (ECE, 15 bins) | **0.025** |\n| Negative log-likelihood / Brier score | 0.652 / 0.358 |\n| Area under the risk\u2013coverage curve (AURC, lower is better) | **0.097** |\n| Ordinal (`score`) mean absolute error, in levels | 0.572 |\n| Coverage / accuracy at confidence \u2265 0.70 | **56.4% / 0.896** |\n| Weights | fp16 319 MB \u00b7 int8 161 MB |\n\n**Selective prediction is the headline.** Calibration is good (ECE 0.025), so the confidence score is a reliable gate. Acting only on decisions with confidence \u2265 0.70 covers 56% of traffic at 89.6% accuracy, against 72.5% accuracy when answering everything. That is the property an agent loop needs: answer the easy majority locally, and hand the rest to an LLM or a human.\n\n### 2.2 Per domain\n\n| Domain | Test decisions | Task-macro accuracy |\n|---|---|---|\n| support | 600 | **0.997** |\n| code | 2,350 | **0.864** |\n| intents | 1,200 | **0.773** |\n| guardrails | 1,094 | **0.768** |\n| agentic | 717 | **0.754** |\n| workflows | 2,000 | **0.689** |\n| policy | 2,700 | **0.669** |\n| reasoning | 5,637 | **0.667** |\n| classification | 1,200 | **0.584** |\n\n### 2.3 Per task\n\n*Held-out tasks were never used for training, calibration or model selection. \"proof_v2 (card)\" lists proof_v2's published score for the same source and task (different samples; indicative only).*\n\n| Domain | Task | n | Chance | **Accuracy** | ECE | proof_v2 (card) |\n|---|---|---|---|---|---|---|\n| support | `bitext/route` | 300 | 0.200 | **1.000** | 0.002 | 0.958 |\n| support | `bitext/category` | 300 | 0.172 | **0.993** | 0.008 | |\n| code | `codexglue/lang_id` | 300 | 0.235 | **1.000** | 0.004 | 0.997 |\n| code | `mbpp/solution` | 300 | 0.250 | **0.987** | 0.014 | 0.992 |\n| code | `codexglue/code_to_doc` | 300 | 0.256 | **0.977** | 0.016 | 0.969 |\n| code | `codexglue/doc_to_code` | 300 | 0.274 | **0.973** | 0.022 | 0.961 |\n| code | `codexglue/func_name` | 275 | 0.263 | **0.938** | 0.021 | 0.901 |\n| code | `bigclonebench/clone` | 300 | 0.500 | **0.863** | 0.098 | 0.383 |\n| code | `humaneval/completion` *(held out)* | 119 | 0.394 | **0.756** | 0.152 | 0.575 |\n| code | `mbpp/bugspot` | 156 | 0.413 | **0.731** | 0.086 | 0.475 |\n| code | `devign/vulnerability` | 300 | 0.500 | **0.553** | 0.026 | 0.542 |\n| intents | `banking77/intent` *(held out)* | 300 | 0.317 | **0.923** | 0.037 | 0.883 |\n| intents | `massive_en/intent` | 300 | 0.122 | **0.907** | 0.046 | |\n| intents | `clinc150/intent` | 300 | 0.122 | **0.793** | 0.081 | 0.850 |\n| intents | `banking77/intent_77` *(held out)* | 300 | 0.013 | **0.470** | 0.200 | |\n| guardrails | `jailbreak/detect` | 262 | 0.500 | **0.966** | 0.018 | |\n| guardrails | `civil_comments/toxic` | 300 | 0.500 | **0.807** | 0.051 | |\n| guardrails | `agentharm/refuse` *(held out)* | 416 | 0.500 | **0.654** | 0.178 | |\n| guardrails | `prompt_injections/detect` *(held out)* | 116 | 0.500 | **0.647** | 0.272 | |\n| agentic | `hotpotqa/retrieve` | 298 | 0.168 | **0.842** | 0.078 | |\n| agentic | `hotpotqa/comparison_yes_no` | 17 | 0.500 | **0.824** | 0.185 | |\n| agentic | `counsel/step_has_error` | 201 | 0.500 | **0.791** | 0.093 | |\n| agentic | `counsel/critique_quality` | 201 | 0.333 | **0.557** | 0.094 | |\n| workflows | `typed_decisions/customer_service` | 500 | 0.280 | **0.720** | 0.099 | |\n| workflows | `typed_decisions/security_incidents` | 500 | 0.340 | **0.712** | 0.140 | |\n| workflows | `typed_decisions/agent_trace_observability` | 500 | 0.300 | **0.696** | 0.106 | |\n| workflows | `typed_decisions/invoice_processing` | 500 | 0.350 | **0.628** | 0.101 | |\n| policy | `policy/access_control_transfer` | 300 | 0.333 | **1.000** | 0.001 | |\n| policy | `policy/return_window_transfer` | 300 | 0.333 | **1.000** | 0.019 | |\n| policy | `policy/free_shipping_transfer` | 300 | 0.500 | **0.850** | 0.067 | |\n| policy | `policy/sla_urgency_transfer` | 300 | 0.250 | **0.713** | 0.214 | |\n| policy | `policy/refund_approval_transfer` | 300 | 0.333 | **0.710** | 0.219 | |\n| policy | `policy/count_threshold_transfer` | 300 | 0.179 | **0.523** | 0.085 | |\n| policy | `policy/invoice_total_transfer` | 300 | 0.500 | **0.523** | 0.020 | |\n| policy | `policy/invoice_overdue_transfer` | 300 | 0.500 | **0.407** | 0.364 | |\n| policy | `policy/table_extreme_transfer` | 300 | 0.240 | **0.290** | 0.036 | |\n| reasoning | `qasc/mcq` | 300 | 0.125 | **0.983** | 0.007 | |\n| reasoning | `snli/must_be_true` | 300 | 0.333 | **0.970** | 0.029 | 0.908 |\n| reasoning | `snli/contradicts` | 300 | 0.333 | **0.967** | 0.035 | 0.892 |\n| reasoning | `scitail/support` | 300 | 0.500 | **0.957** | 0.030 | |\n| reasoning | `sciq/mcq` | 300 | 0.250 | **0.950** | 0.023 | 0.692 |\n| reasoning | `snli/nli` | 594 | 0.333 | **0.837** | 0.035 | |\n| reasoning | `mnli/claim` | 300 | 0.333 | **0.807** | 0.072 | 0.492 |\n| reasoning | `boolq/yes_no` | 300 | 0.500 | **0.783** | 0.071 | 0.717 |\n| reasoning | `gsm8k/math` | 300 | 0.250 | **0.637** | 0.050 | 0.275 |\n| reasoning | `commonsense_qa/mcq` | 296 | 0.200 | **0.611** | 0.058 | 0.442 |\n| reasoning | `arc_easy/mcq` *(held out)* | 300 | 0.250 | **0.553** | 0.057 | 0.425 |\n| reasoning | `openbookqa/mcq` | 300 | 0.250 | **0.550** | 0.068 | 0.292 |\n| reasoning | `winogrande/blank` | 300 | 0.500 | **0.540** | 0.089 | |\n| reasoning | `arc_challenge/mcq` *(held out)* | 300 | 0.250 | **0.423** | 0.085 | 0.308 |\n| reasoning | `anli/nli` | 300 | 0.333 | **0.393** | 0.185 | |\n| reasoning | `mmlu/mcq` *(held out)* | 300 | 0.250 | **0.393** | 0.094 | |\n| reasoning | `hellaswag/continuation` | 300 | 0.250 | **0.383** | 0.171 | |\n| reasoning | `aqua_rat/math` | 247 | 0.200 | **0.259** | 0.071 | |\n| classification | `ag_news/topic` | 300 | 0.250 | **0.847** | 0.051 | |\n| classification | `yelp/score` | 300 | 0.200 | **0.640** | 0.069 | |\n| classification | `emotion/6way` *(held out)* | 300 | 0.167 | **0.480** | 0.049 | |\n| classification | `sst5/score` *(held out)* | 300 | 0.200 | **0.370** | 0.086 | |\n\n### 2.4 Comparison with proof_v2 (indicative)\n\nOn the 22 tasks that both this report and proof_v2's model card cover, LightDec's task-macro accuracy is **0.807 vs 0.679**, and it scores higher on **20 of 22**. The largest gains are on the tasks proof_v2 reported as weak:\n\n| Task | proof_v2 (card) | LightDec |\n|---|---|---|\n| BigCloneBench clone detection | 0.383 | **0.863** |\n| GSM8K (4-option numeric) | 0.275 | **0.637** |\n| MultiNLI claim | 0.492 | **0.807** |\n| OpenBookQA | 0.292 | **0.550** |\n| MBPP bug spotting | 0.475 | **0.731** |\n| HumanEval completion *(held out)* | 0.575 | **0.756** |\n| ARC-Challenge *(held out)* | 0.308 | **0.423** |\n\nLightDec is lower on CLINC150 (0.793 vs 0.850) and MBPP task\u2192solution (0.987 vs 0.992).\n\nThese are **different test samples and, for some tasks, different question formats**. For example, LightDec's MultiNLI task is three-way NLI, and its bug-spotting mutants are verified to fail the unit tests. The like-for-like head-to-head, which runs proof_v2 on identical decisions (notebook cell 23), **did not run** for this checkpoint (`head_to_head: null`).\n\n### 2.5 Comparison with Laya and TypeSafe Jev (indicative)\n\n| Benchmark | LightDec | Laya | TypeSafe Jev 1.13.0 |\n|---|---|---|---|\n| typed-decisions test (2,000 decisions, 4 workflows) | 0.689 | 0.766 (fine-tuned) \u00b7 0.362 (zero-shot) | 0.727 |\n| AG News (4 labels) | 0.847 | 0.950 | 0.910 |\n| DAIR Emotion, 6 labels *(held out for LightDec)* | 0.480 | 0.595 | 0.480 |\n| Banking77, all 77 labels in one question *(held out)* | 0.470 | 0.425 | 0.870 (72 labels) |\n| SST-5 (ordinal) *(held out)* | 0.370 | 0.372 | \u2014 |\n\nLaya's numbers are from its own benchmark report; Jev's are third-party published. LightDec was trained on the typed-decisions training split, like the fine-tuned Laya checkpoint. **On typed-decisions LightDec trails both** (the teacher-agreement ceiling is 0.735 and the majority-class baseline 0.461). It matches Jev on Emotion, edges Laya on all-77 Banking77, and trails both on AG News. LightDec is 2.6\u00d7 smaller than Laya's 421M English checkpoint.\n\n---\n\n## 3. Intended and out-of-scope uses\n\n### Intended\n\n| Use | Measured evidence |\n|---|---|\n| **Support and ticket routing** | Bitext route 1.000, Bitext category 0.993; Banking77 (held out, 2\u20135 options) 0.923; MASSIVE 0.907 |\n| **Code understanding against a menu** | Language ID 1.000; code\u2194description 0.973\u20130.977; task\u2192solution 0.987; function naming 0.938 |\n| **Guardrails** | Jailbreak detection 0.966; toxicity 0.807. Held out: prompt-injection 0.647, AgentHarm refusal 0.654, so recalibrate and test on your own traffic |\n| **Agent loops** | Retrieval routing (HotpotQA) 0.842; \"does this agent step contain an error?\" (Counsel) 0.791 |\n| **Statement verification** | SNLI must-be-true / contradicts 0.970 / 0.967; SciTail 0.957; SciQ 0.950 |\n| **Selective automation** | 89.6% accuracy on the 56% of decisions with confidence \u2265 0.70 |\n\n### Out of scope\n\n- **Multi-step arithmetic and quantitative reasoning.** AQuA 0.259 (chance 0.20); counting and summing policies 0.52; comparing values in a table 0.290. GSM8K reaches 0.637 only because it is posed as 4-option multiple choice with near-miss distractors. Route real math to an LLM or code.\n- **Date reasoning in unfamiliar formats.** The invoice-overdue transfer test (ISO dates, whereas training used \"Month DD, YYYY\") scores **0.407, below chance, with ECE 0.364**. It is confidently wrong there. Normalise dates before asking, or compute them in code.\n- **Very wide label sets.** All 77 Banking77 intents in one question score 0.470. Pre-filter to a shortlist of about 20 options (\u00a77.1).\n- **Hard commonsense and exam knowledge**: HellaSwag 0.383, ANLI 0.393, MMLU 0.393.\n- **Code security and correctness gating**: Devign 0.553 is near chance. Don't use it to approve code.\n- **Open-ended questions**: it always picks one of your options. Add \"None of the above\" when appropriate; it was trained with that option.\n- **Non-English text**, and **high-stakes decisions without human oversight**.\n\n---\n\n## 4. The Hub repository\n\n| File | Contents |\n|---|---|\n| `model.safetensors` | fp16 weights (319 MB): encoder, decision head and the temperature buffer |\n| `falcondec_config.json` | Layout (sequence lengths, option budgets), special-token ids, temperatures, defer threshold, version, lineage |\n| `encoder/` | Backbone configuration (the encoder is rebuilt from this, then the weights are loaded) |\n| `tokenizer/` | Tokenizer files |\n| `falcondec_modeling.py` | `FalconDec`, `load_falcondec`, `decide`, `score_items`, `save_falcondec` and the int8 codec |\n| `falcondec_report.json` | Training configuration, data counts, history, temperatures and all test results |\n| `compact-int8/` | The same model with per-channel int8 weights (`model_int8.safetensors`, 161 MB); a complete, self-contained directory with its own config, tokenizer and modeling file |\n| `README.md` | This card |\n\n**Pin a revision in production.** The repo can change, so pass a commit hash when you load.\n\n---\n\n## 5. Install, load and read a result\n\n```bash\npip install torch \"transformers>=4.48\" safetensors huggingface_hub numpy\n```\n\n### 5.1 Load\n\nSave this helper as `lightdec.py` next to your code. Every example below uses it.\n\n```python\n# lightdec.py\nimport importlib.util, json, shutil\nfrom pathlib import Path\nfrom huggingface_hub import snapshot_download\n\n\ndef load_lightdec(repo=\"Falconsai/LightDec\", revision=None, variant=\"fp16\", device=None, dtype=None):\n \"\"\"Returns (fdm, model, tokenizer). variant: \"fp16\" (319 MB) or \"int8\" (161 MB, dequantised on load).\"\"\"\n path = Path(repo) if Path(repo).exists() else Path(snapshot_download(repo, revision=revision))\n if variant == \"int8\":\n path = path / \"compact-int8\"\n fc = json.loads((path / \"falcondec_config.json\").read_text(encoding=\"utf-8\"))\n expected = fc.get(\"weights\", \"model.safetensors\")\n if not (path / expected).exists(): # e.g. a renamed weight file in a processed copy\n cands = sorted(path.glob(\"*.safetensors\"))\n if not cands:\n raise FileNotFoundError(f\"no .safetensors weights in {path}\")\n local = Path(\"lightdec_local\") / variant\n shutil.copytree(path, local, dirs_exist_ok=True)\n shutil.copy(cands[0], local / expected)\n path = local\n spec = importlib.util.spec_from_file_location(\"falcondec_modeling\", str(path / \"falcondec_modeling.py\"))\n fdm = importlib.util.module_from_spec(spec)\n spec.loader.exec_module(fdm)\n model, tok = fdm.load_falcondec(str(path), device=device, dtype=dtype) # cuda if available, else cpu\n return fdm, model, tok\n```\n\n```python\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec() # or load_lightdec(revision=\"<commit>\", variant=\"int8\")\nprint(model.fcfg[\"name\"], model.fcfg[\"version\"], round(model.num_parameters() / 1e6, 1), \"M params\")\n```\n\nIf loading prints `[FalconDec] load warning: missing=\u2026 unexpected=\u2026`, the weights didn't match the architecture. Treat that as a failed load (\u00a76.1 checks for it).\n\n### 5.2 Decide\n\n`decide()` takes one state and any number of typed questions: either a list, or a Jev/Laya-style dict keyed by name.\n\n```python\nstate = {\"from\": \"user@acme.com\", \"subject\": \"Duplicate charge on invoice #4411\",\n \"body\": \"We were billed twice for March. Please refund the duplicate today or we will cancel our plan.\"}\n\nout = fdm.decide(model, tok, state, {\n \"department\": {\"type\": \"choice\", \"instructions\": \"Which department should handle this request?\",\n \"criteria\": {\"billing\": \"invoices, payments, refunds\", \"technical\": \"bugs, outages\",\n \"sales\": \"pricing, contracts\", \"other\": \"everything else\"}},\n \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent is this request?\",\n \"criteria\": [\"not urgent\", \"soon\", \"critical deadline or blocking issue\"]},\n \"churn_risk\": {\"type\": \"noul\", \"instructions\": \"Does the user threaten to cancel or leave?\"},\n})\na = out[\"answers\"]\nprint(a[\"department\"][\"choice\"], round(a[\"department\"][\"confidence\"], 3), a[\"department\"][\"defer\"])\nprint(\"urgency level\", round(a[\"urgency\"][\"expected_level\"], 2), \"of\", 2)\nprint(\"P(churn)\", round(a[\"churn_risk\"][\"p_true\"], 3))\n```\n\nPlain options work too: `{\"question\": \"Which team?\", \"options\": [\"Accounts\", \"Billing\", \"Shipping\"]}`.\n\n### 5.3 Reading a result\n\nEach item in `out[\"results\"]` (and `out[\"answers\"][key]`) contains:\n\n| Field | Meaning |\n|---|---|\n| `key` | The question's name (dict input) or `None` |\n| `type` | `choice`, `noul` or `score` |\n| `choice` | The chosen key: the criteria key, the option text, `True`/`False` for `noul`, or the level index for `score` |\n| `choice_text` | The option text the model saw |\n| `confidence` | Calibrated probability of `choice` |\n| `probs` | The full distribution, keyed by `str(key)` |\n| `defer` | `True` when `confidence` is below the defer threshold (default 0.70, stored in the config): don't act on it (\u00a77.3) |\n| `p_true` | `noul` only: calibrated P(yes) |\n| `expected_level` | `score` only: probability-weighted level (0 \u2026 k\u22121); better than the argmax for ordinal rubrics (test MAE 0.57 levels) |\n\nHow to read them:\n\n- **Low confidence, spread probabilities**: the state doesn't support any option clearly. Defer, or add \"None of the above\".\n- **`noul` near 0.5**: genuinely ambiguous. Ask for more information rather than guessing.\n- **`score`**: use `expected_level` for thresholds (\"escalate if \u2265 1.5\") rather than `choice`.\n\n---\n\n## 6. Testing the model\n\nTests 6.1\u20136.3 need no labelled data, so run them in CI whenever you change the revision. Test 6.4 is the one that tells you whether to ship.\n\n### 6.1 Integrity and determinism\n\nSave as `check_lightdec.py` and run `python check_lightdec.py [revision]`.\n\n```python\nimport contextlib, io, sys\nimport numpy as np\nfrom lightdec import load_lightdec\n\nrev = sys.argv[1] if len(sys.argv) > 1 else None\nlog = io.StringIO()\nwith contextlib.redirect_stdout(log):\n fdm, model, tok = load_lightdec(revision=rev)\nassert \"load warning\" not in log.getvalue(), log.getvalue()\n\nfc = model.fcfg\nassert fc[\"name\"] == \"FalconDec\", fc[\"name\"] # LightDec checkpoints use the FalconDec architecture name\nT = model.temperature.float().cpu().numpy()\nassert T.shape == (3, 4) and (T > 0).all(), T\nq = {\"team\": {\"question\": \"Which team?\", \"options\": [\"recover password\", \"shipping\", \"invoicing\"]}}\na = fdm.decide(model, tok, \"I forgot my password and can't sign in.\", q)[\"answers\"][\"team\"]\nb = fdm.decide(model, tok, \"I forgot my password and can't sign in.\", q)[\"answers\"][\"team\"]\nassert all(abs(a[\"probs\"][k] - b[\"probs\"][k]) < 1e-4 for k in a[\"probs\"]), \"non-deterministic\"\nassert abs(sum(a[\"probs\"].values()) - 1) < 1e-3\nprint(f\"OK LightDec (FalconDec v{fc['version']}, notebook {fc.get('notebook_version')}) choice={a['choice']} \"\n f\"conf={a['confidence']:.3f} defer_threshold={fc.get('defer_threshold')}\")\n```\n\nThe stored temperatures should read approximately `[[1.707, 1.352, 1.466, 1.349], [1.402 \u00d74], [1.453 \u00d74]]` (\u00a711.2).\n\n### 6.2 Behavioural tests (pytest)\n\nSave as `tests/test_lightdec.py` and run `pytest -q`. Set `LIGHTDEC_REVISION` to test a pinned commit.\n\n```python\nimport os, random\nimport pytest\nfrom lightdec import load_lightdec\n\n\n@pytest.fixture(scope=\"session\")\ndef fd():\n return load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"))\n\n\ndef ask(fd, state, question, options, **kw):\n fdm, model, tok = fd\n return fdm.decide(model, tok, state, [dict(question=question, options=options, **kw)])[\"results\"][0]\n\n\ndef test_probabilities_are_valid(fd):\n r = ask(fd, \"The build failed on main.\", \"What next?\", [\"Revert\", \"Ignore\", \"Retry\"])\n assert all(0 <= p <= 1 for p in r[\"probs\"].values()) and abs(sum(r[\"probs\"].values()) - 1) < 1e-3\n\n\ndef test_typed_outputs(fd):\n fdm, model, tok = fd\n out = fdm.decide(model, tok, \"I was charged twice. Refund me or I'm leaving.\", {\n \"refund\": {\"type\": \"noul\", \"instructions\": \"Does the user ask for a refund?\"},\n \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent?\", \"criteria\": [\"low\", \"medium\", \"high\"]}})[\"answers\"]\n assert 0 <= out[\"refund\"][\"p_true\"] <= 1 and out[\"refund\"][\"choice\"] in (True, False)\n assert 0 <= out[\"urgency\"][\"expected_level\"] <= 2\n\n\ndef test_support_routing(fd):\n r = ask(fd, \"I forgot my password and the reset email never arrived.\", \"Which team should handle this?\",\n [\"recover password\", \"billing and payment\", \"delivery information\"])\n assert r[\"choice\"] == \"recover password\"\n\n\ndef test_fanout_matches_single_questions(fd):\n # Batching changes padding; under bf16 that moves probabilities slightly, never the substance.\n fdm, model, tok = fd\n state = \"I forgot my password and can't sign in.\"\n qs = [{\"question\": \"Team?\", \"options\": [\"recover password\", \"shipping\", \"invoicing\"]},\n {\"question\": \"Urgent?\", \"options\": [\"Yes\", \"No\"]}]\n together = fdm.decide(model, tok, state, qs)[\"results\"]\n for q, t in zip(qs, together):\n alone = fdm.decide(model, tok, state, [q])[\"results\"][0]\n assert all(abs(alone[\"probs\"][k] - t[\"probs\"][k]) < 2e-2 for k in alone[\"probs\"])\n\n\ndef test_option_order_is_mostly_irrelevant(fd):\n # The head is order-equivariant, but the encoder sees positions; training reshuffled options every epoch.\n state, q = \"Where is my parcel? It's three days late.\", \"What should support do?\"\n opts = [\"Give the delivery status\", \"Start a refund\", \"Book an appointment\"]\n base = ask(fd, state, q, opts)[\"choice\"]\n same = sum(ask(fd, state, q, random.Random(s).sample(opts, len(opts)))[\"choice\"] == base for s in range(5))\n assert same >= 4\n\n\ndef test_many_options_use_the_tournament(fd):\n opts = [f\"topic number {i}\" for i in range(119)] + [\"reset my password\"]\n r = ask(fd, \"I can't log in, I need to reset my password.\", \"What does the user want?\", opts)\n assert len(r[\"probs\"]) == 120 and abs(sum(r[\"probs\"].values()) - 1) < 1e-3\n\n\ndef test_int8_agrees_with_fp16(fd):\n fdm8, m8, tok8 = load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"), variant=\"int8\")\n fdm, model, tok = fd\n items = [dict(state=s, question=\"Which team?\", options=[\"billing\", \"shipping\", \"accounts\", \"technical\"])\n for s in [\"I was double charged\", \"Where is my parcel?\", \"Change my email\", \"The app crashes on start\",\n \"Refund the duplicate payment\", \"Package never arrived\", \"Reset my login\", \"Error 500 on checkout\"]]\n a = [p.argmax() for p in fdm.score_items(model, tok, items)]\n b = [p.argmax() for p in fdm8.score_items(m8, tok8, items)]\n assert sum(x == y for x, y in zip(a, b)) >= len(items) - 1\n```\n\n### 6.3 Latency\n\n```python\nimport time, numpy as np, torch\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nq = [{\"question\": \"Route?\", \"options\": [\"billing and payment\", \"shipping\", \"recover password\"]}]\nfor _ in range(5):\n fdm.decide(model, tok, \"I was charged twice.\", q)\nt = []\nfor _ in range(100):\n if torch.cuda.is_available(): torch.cuda.synchronize()\n t0 = time.perf_counter(); fdm.decide(model, tok, \"I was charged twice.\", q)\n if torch.cuda.is_available(): torch.cuda.synchronize()\n t.append((time.perf_counter() - t0) * 1000)\nprint(f\"p50 {np.percentile(t, 50):.1f} ms p95 {np.percentile(t, 95):.1f} ms on {model.device}\")\n```\n\nFor CPU serving, load the int8 variant with `device=\"cpu\", dtype=torch.float32`, and optionally apply `torch.ao.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)` for int8 matrix multiplies.\n\n### 6.4 Accuracy on your own labelled data\n\nWrite 50\u2013500 decisions that look like your real traffic, one JSON object per line. `expected` may be a letter, a 0-based index or the option text; `type` is optional.\n\n```json\n{\"id\": \"t1\", \"tag\": \"support\", \"state\": \"\u2026\", \"question\": \"\u2026\", \"options\": [\"\u2026\", \"\u2026\"], \"expected\": \"B\", \"type\": \"choice\"}\n```\n\n```python\nimport json, string\nimport numpy as np\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nrows = [json.loads(l) for l in open(\"my_eval.jsonl\", encoding=\"utf-8\") if l.strip()]\n\ndef idx(v, opts):\n if isinstance(v, int): return v\n v = str(v).strip()\n if len(v) == 1 and v.upper() in string.ascii_uppercase[:len(opts)]: return string.ascii_uppercase.index(v.upper())\n return [o.lower() for o in opts].index(v.lower())\n\nitems = [dict(state=r[\"state\"], question=r.get(\"question\", \"\"), options=r[\"options\"], type=r.get(\"type\", \"choice\"))\n for r in rows]\nprobs = fdm.score_items(model, tok, items, batch_size=64)\ngold = np.array([idx(r[\"expected\"], r[\"options\"]) for r in rows])\npred = np.array([p.argmax() for p in probs]); conf = np.array([p.max() for p in probs]); ok = pred == gold\n\ndef ece(c, k, bins=15):\n e = 0.0\n for lo in np.linspace(0, 1, bins, endpoint=False):\n m = (c > lo) & (c <= lo + 1 / bins)\n if m.any(): e += m.mean() * abs(c[m].mean() - k[m].mean())\n return e\n\nprint(f\"accuracy {ok.mean():.3f} | ECE {ece(conf, ok):.3f}\")\nfor tag in sorted({r.get(\"tag\", \"all\") for r in rows}):\n m = np.array([r.get(\"tag\", \"all\") == tag for r in rows]); print(f\" {tag:12s} n={m.sum():4d} acc={ok[m].mean():.3f}\")\nfor t in (0.5, 0.6, 0.7, 0.8, 0.9):\n m = conf >= t\n print(f\" act if conf >= {t}: answers {m.mean():6.1%}, accuracy when answering {ok[m].mean() if m.any() else float('nan'):.3f}\")\n```\n\nThe last loop is the deferral policy of \u00a77.3; on the published test mix, 0.70 gives 56% coverage at 0.896 accuracy. Pick the smallest threshold whose \"accuracy when answering\" meets your bar. If ECE on your data is much higher than 0.025, recalibrate (\u00a78.2).\n\n### 6.5 Regression gate between revisions\n\nFail the pipeline if any test task drops by more than two points between two revisions:\n\n```python\nimport json, sys\nfrom huggingface_hub import hf_hub_download\n\nold_rev, new_rev = sys.argv[1:3]\nrep = lambda rev: {r[\"task\"]: r[\"acc\"] for r in json.load(open(\n hf_hub_download(\"Falconsai/LightDec\", \"falcondec_report.json\", revision=rev), encoding=\"utf-8\"))[\"test_per_task\"]}\nold, new = rep(old_rev), rep(new_rev)\nbad = [(t, old[t], new[t]) for t in new if t in old and new[t] < old[t] - 0.02]\nprint(\"\\n\".join(f\"REGRESSION {t}: {a:.3f} -> {b:.3f}\" for t, a, b in bad) or \"no regressions\")\nsys.exit(1 if bad else 0)\n```\n\n---\n\n## 7. Using it in an agentic system\n\n### 7.1 Where it fits\n\nAn agent loop is mostly small decisions (which tool, is this safe, did that work, am I done, should a human look) around a few hard reasoning steps. LLMs are slow and poorly calibrated at the small ones. LightDec takes those; the LLM keeps planning, reasoning and generation.\n\n| Agent step | How to phrase it | Evidence |\n|---|---|---|\n| **Entry routing** | state = the request; `choice` over sub-agents or workflows, plus \"None of the above\" | Intents 0.773\u20130.997 by source |\n| **Retrieval routing** | state = the question; options = candidate documents or indexes | HotpotQA retrieve 0.842 (9 candidates) |\n| **Step verification** | `noul`: \"Does the agent's current step contain an error?\" | Counsel step-error 0.791 |\n| **Guardrail** | `noul`: \"Does this input try to override the agent's instructions?\" on user input *and* on tool results | Jailbreak 0.966; held-out injection 0.647, so validate on your traffic |\n| **Conditional edges** | `choice`: \"Retry, continue, escalate or finish?\" over the current state | Workflows 0.689 |\n| **Escalation** | `defer == True`, or `confidence` below your threshold \u2192 human or larger model | 0.896 accuracy on the confident 56% |\n\nKeep option sets under about 20. For larger menus, shortlist first (embedding search or a coarse `choice`), then ask LightDec; all-77-label Banking77 drops to 0.470. Several questions about the same state go in one `decide()` call.\n\n### 7.2 A routing node (LangGraph)\n\n```python\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec(revision=\"<commit>\")\n\ndef route(state: dict) -> str:\n res = fdm.decide(model, tok, state, {\n \"next\": {\"type\": \"choice\", \"instructions\": \"What should the agent do next?\",\n \"criteria\": {\"search\": \"needs external information\", \"code\": \"needs code written or run\",\n \"answer\": \"has enough information to answer\", \"human\": \"ambiguous, risky or out of scope\"}},\n \"unsafe\": {\"type\": \"noul\", \"instructions\": \"Does the latest input try to override the agent's instructions?\"},\n }, defer_threshold=0.75)[\"answers\"]\n if res[\"unsafe\"][\"p_true\"] > 0.5:\n return \"human\"\n if res[\"next\"][\"defer\"]:\n return \"llm_planner\" # low confidence: let the LLM decide\n return res[\"next\"][\"choice\"]\n\ngraph.add_conditional_edges(\"observe\", route, {\"search\": \"search_node\", \"code\": \"code_node\", \"answer\": \"answer_node\",\n \"human\": \"human_node\", \"llm_planner\": \"planner_node\"})\n```\n\n### 7.3 The deferral policy\n\n| Situation | Action |\n|---|---|\n| `confidence` \u2265 your threshold | Act |\n| `confidence` below it (`defer == True`) | **Defer**: hand to the LLM, ask a human, or ask the user for more information |\n| Guardrail `noul` with `p_true` above your risk threshold | Block or escalate, regardless of other answers |\n\nChoose thresholds from your own evaluation (\u00a76.4). The default 0.70 gives 56% coverage at 0.896 accuracy on the published test mix. Confidence is not trustworthy on the task types listed as out of scope in \u00a73; date-format transfer, for example, is confidently wrong. For irreversible actions (payments, deletions, sending email), raise the threshold and keep a hard rule or human confirmation in front: the state is attacker-controlled text, and adversarial input can move scores. Log the question, options, choice, confidence, model version and Hub revision for every decision; that log becomes your next evaluation and fine-tuning set (\u00a78.3).\n\n### 7.4 As a tool for Claude (tool use)\n\n```python\nimport json, threading\nimport anthropic\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nlock = threading.Lock()\nclient = anthropic.Anthropic()\ntools = [{\n \"name\": \"lightdec_decide\",\n \"description\": (\"Fast, local, calibrated closed-set decision model. Give it a state (text or JSON), a question and \"\n \"2-20 distinct options; it returns the choice, a calibrated confidence and a 'defer' flag. \"\n \"If 'defer' is true, don't rely on the answer. Not for arithmetic, dates or multi-step reasoning.\"),\n \"input_schema\": {\"type\": \"object\", \"properties\": {\n \"state\": {\"type\": \"string\", \"description\": \"The message, document excerpt or JSON state.\"},\n \"question\": {\"type\": \"string\"},\n \"options\": {\"type\": \"array\", \"items\": {\"type\": \"string\"}, \"minItems\": 2},\n \"type\": {\"type\": \"string\", \"enum\": [\"choice\", \"score\"], \"description\": \"score = options are ordered levels\"}},\n \"required\": [\"state\", \"question\", \"options\"]},\n}]\n\ndef run_tool(inp):\n with lock:\n r = fdm.decide(model, tok, inp[\"state\"], [{\"question\": inp[\"question\"], \"options\": inp[\"options\"],\n \"type\": inp.get(\"type\", \"choice\")}])[\"results\"][0]\n return {k: r[k] for k in (\"choice\", \"confidence\", \"defer\", \"probs\") if k in r}\n\nmessages = [{\"role\": \"user\", \"content\": \"Triage: 'I forgot my password and the reset email never arrived.' \"\n \"Teams: Accounts, Billing, Shipping.\"}]\nwhile True:\n resp = client.messages.create(model=\"claude-sonnet-5\", max_tokens=1024, tools=tools, messages=messages)\n if resp.stop_reason != \"tool_use\":\n print(\"\".join(b.text for b in resp.content if b.type == \"text\"))\n break\n messages.append({\"role\": \"assistant\", \"content\": resp.content})\n results = []\n for block in resp.content:\n if block.type == \"tool_use\" and block.name == \"lightdec_decide\":\n try:\n results.append({\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": json.dumps(run_tool(block.input))})\n except Exception as exc:\n results.append({\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": str(exc), \"is_error\": True})\n messages.append({\"role\": \"user\", \"content\": results})\n```\n\n### 7.5 As an MCP server\n\n`pip install mcp`, then save `lightdec_mcp.py` next to `lightdec.py`:\n\n```python\nimport os, threading\nfrom mcp.server.fastmcp import FastMCP\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"),\n variant=os.environ.get(\"LIGHTDEC_VARIANT\", \"fp16\"))\nMIN_CONF = float(os.environ.get(\"LIGHTDEC_MIN_CONF\", \"0.7\"))\nlock = threading.Lock()\nmcp = FastMCP(\"lightdec\")\n\n\n@mcp.tool()\ndef decide(state: str, question: str, options: list[str], type: str = \"choice\") -> dict:\n \"\"\"Choose one of 2-20 distinct options for a question about a state. type=\"score\" means ordered levels.\n Returns the choice, a calibrated confidence and 'defer' (true = not reliable enough to act on).\"\"\"\n with lock:\n r = fdm.decide(model, tok, state, [{\"question\": question, \"options\": options, \"type\": type}],\n defer_threshold=MIN_CONF)[\"results\"][0]\n return {k: r[k] for k in (\"choice\", \"confidence\", \"defer\", \"probs\", \"expected_level\") if k in r}\n\n\n@mcp.tool()\ndef decide_many(state: str, questions: dict) -> dict:\n \"\"\"Several typed questions about one state: {name: {\"type\": \"choice\"|\"noul\"|\"score\",\n \"instructions\": str, \"criteria\": {key: description} | [levels]}}.\"\"\"\n with lock:\n return fdm.decide(model, tok, state, questions, defer_threshold=MIN_CONF)[\"answers\"]\n\n\nif __name__ == \"__main__\":\n mcp.run()\n```\n\nRegister it in Claude Desktop's `claude_desktop_config.json`:\n\n```json\n{\n \"mcpServers\": {\n \"lightdec\": {\n \"command\": \"C:\\\\path\\\\to\\\\python.exe\",\n \"args\": [\"C:\\\\path\\\\to\\\\lightdec_mcp.py\"],\n \"env\": { \"LIGHTDEC_REVISION\": \"<commit>\", \"LIGHTDEC_VARIANT\": \"int8\" }\n }\n }\n}\n```\n\n### 7.6 As an HTTP microservice\n\n`pip install fastapi uvicorn`, then save `serve_lightdec.py`:\n\n```python\nimport threading\nfrom fastapi import FastAPI, HTTPException\nfrom pydantic import BaseModel\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nlock = threading.Lock()\napp = FastAPI(title=\"LightDec\")\n\n\nclass Decide(BaseModel):\n state: str | dict\n questions: dict | list\n defer_threshold: float = 0.7\n\n\n@app.get(\"/health\")\ndef health():\n return {\"ok\": True, \"model\": \"LightDec\", \"version\": model.fcfg.get(\"version\")}\n\n\n@app.post(\"/v1/decide\")\ndef decide(req: Decide):\n try:\n with lock:\n return fdm.decide(model, tok, req.state, req.questions, defer_threshold=req.defer_threshold)\n except (ValueError, KeyError, TypeError) as exc:\n raise HTTPException(422, str(exc))\n```\n\nRun it with `uvicorn serve_lightdec:app --port 9904 --workers 1`. Each worker holds its own copy of the model; scale out with more processes.\n\n### 7.7 Without an LLM: a support intake step\n\n```python\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nQUEUES = {\"billing\": \"charges, invoices, payments, refunds\", \"returns\": \"returning or exchanging items\",\n \"delivery\": \"shipping status, late or missing parcels\", \"accounts\": \"login, password, profile\",\n \"human\": \"complaints or requests to speak to a person\"}\n\ndef intake(message: str) -> dict:\n a = fdm.decide(model, tok, message, {\n \"queue\": {\"type\": \"choice\", \"instructions\": \"Which team should handle this message?\", \"criteria\": QUEUES},\n \"urgent\": {\"type\": \"noul\", \"instructions\": \"Does this need a reply within the hour?\"},\n })[\"answers\"]\n if a[\"queue\"][\"defer\"]:\n return {\"action\": \"human_review\", \"suggestion\": a[\"queue\"][\"choice\"],\n \"reason\": f\"low confidence ({a['queue']['confidence']:.0%})\"}\n return {\"action\": \"enqueue\", \"queue\": a[\"queue\"][\"choice\"],\n \"priority\": \"high\" if a[\"urgent\"][\"p_true\"] >= 0.5 else \"normal\",\n \"evidence\": {\"confidence\": round(a[\"queue\"][\"confidence\"], 3), \"model\": \"LightDec\"}}\n```\n\n---\n\n## 8. Tuning for your domain\n\n### 8.1 Change the deferral threshold\n\nPass `defer_threshold=` to `decide()`, or set `model.fcfg[\"defer_threshold\"]`. This changes nothing in the model and is usually enough.\n\n### 8.2 Recalibrate on your data\n\nIf ECE on your traffic (\u00a76.4) is noticeably worse than 0.025, fit one extra temperature on top of the stored ones and save a recalibrated copy:\n\n```python\nimport numpy as np, torch\n# probs, gold: from \u00a76.4 (probabilities already include the stored temperatures)\nlogp = [np.log(np.clip(p, 1e-12, 1)) for p in probs]\ndef nll(s):\n return -np.mean([(lp / s)[g] - np.log(np.exp(lp / s).sum()) for lp, g in zip(logp, gold)])\ns = min(np.linspace(0.5, 3.0, 51), key=nll)\nprint(\"extra temperature\", s)\nwith torch.no_grad():\n model.temperature.mul_(float(s))\nfdm.save_falcondec(model, tok, \"lightdec_recalibrated\") # add int8=True for the compact variant\n```\n\nFit on one labelled set and measure on another.\n\n### 8.3 Fine-tune on your own decisions\n\nUse the FalconDec training notebook (V2):\n\n1. Export your logged and corrected decisions as JSONL: `{\"state\", \"question\", \"options\", \"answer\", \"type\"?, \"task\"?}`.\n2. In cell 2, set `MODE=\"finetune\"`, `FINETUNE_FROM=\"Falconsai/LightDec\"`, `CUSTOM_DATA_JSONL=\"your_file.jsonl\"`, and a preset.\n3. Run the notebook. The version bumps automatically (1.0.0 \u2192 1.0.1), and the report records the lineage.\n\nYour tasks are up-weighted (`CUSTOM_WEIGHT`), while the public tasks keep the model general. Adding decisions with ISO-format dates and numeric tables is the most direct fix for the transfer weaknesses in \u00a73. Check the regression gate (\u00a76.5) before publishing.\n\n### 8.4 Save and publish\n\n```python\nfdm.save_falcondec(model, tok, \"lightdec_out\") # fp16\nfdm.save_falcondec(model, tok, \"lightdec_out/compact-int8\", int8=True)\n\nfrom huggingface_hub import HfApi # needs a write token: huggingface-cli login\nHfApi().upload_folder(folder_path=\"lightdec_out\", repo_id=\"Falconsai/LightDec\", commit_message=\"LightDec v1.0.x\")\n```\n\n---\n\n## 9. Architecture\n\n```\n[CLS] question [SEP] [MASK] option\u2081 [MASK] option\u2082 \u2026 [MASK] option\u2096 [SEP] state [SEP]\n \u2502\n Ettin-150M encoder (22 layers, hidden 768)\n \u2502\n hidden state at each [MASK] + CLS context + question-type embedding\n \u2502\n set transformer: 2 layers, 8 heads, no positional encoding \u2192 options attend to each other, order-equivariant\n \u2502\n MLP \u2192 one logit per option \u2192 \u00f7 temperature[type, option-count bucket] \u2192 softmax\n```\n\n- **Option markers** (the approach Laya uses): every option is read at its own `[MASK]` token, so all options are scored in **one** pass.\n- **Question first, state last**: when the input is too long, the tail of the state is truncated, never the options.\n- **Adaptive option budget**: up to 24 tokens per option within a 192-token head budget that grows with the option count, and 2,048-token sequences above 24 options, so labels stay distinct. Above 96 options, `decide()` runs a tournament.\n- **Typed primitives**: `noul` is rendered as a neutral two-option Yes/No choice; `score` keeps its level order and reports an expected level.\n- **Calibration** lives in the model: a 3 \u00d7 4 temperature table (question type \u00d7 option-count bucket: \u22642, 3\u20135, 6\u201312, >12).\n- **int8 storage**: per-output-channel symmetric int8 for every weight matrix, fp16 elsewhere, dequantised on load.\n\n---\n\n## 10. Training data\n\n**155,747 training, 13,148 validation and 17,498 test decisions** from 58 tasks. No source failed to load (`skipped_builders` is empty). Every example is a **decision**: `state`, `question`, `options`, `answer`, plus type, task and domain.\n\n| Domain | Sources | Decisions built |\n|---|---|---|\n| Support | Bitext customer support | Intent routing and category routing; 30% of messages wrapped as JSON program state |\n| Intents | CLINC150, MASSIVE (en), Banking77 *(held out)* | Intents as runtime-defined options, 2\u201348 per question; Banking77 also as one 77-option question |\n| Code | CodeXGLUE code-to-text (6 languages), Devign, BigCloneBench, MBPP, HumanEval *(held out)* | Language ID, code\u2194description, function naming, vulnerability, clones, task\u2192solution; bug spotting against single-fault mutants **verified to fail the unit tests** |\n| Guardrails | Jailbreak classification, Civil Comments; deepset prompt-injections and AgentHarm *(held out)* | `noul` detection and refusal |\n| Agentic | Counsel (human meta-evaluations of agent-step critiques), HotpotQA, AgentTrek | Step-error and critique-quality, retrieval routing and comparison yes/no; AgentTrek next-action type and finish-now. AgentTrek loaded without errors, but none of its decisions appear in the test split, so its contribution isn't measured |\n| Workflows | LocalLLaMA/typed-decisions | `choice`/`noul`/`score` with the teacher's soft probabilities; test = the benchmark's 2,000-decision test split |\n| Policy | Synthetic, executable (in-notebook) | Return windows, approval tiers, AND/OR eligibility, overdue invoices, table look-ups, counting, SLA urgency, access control, invoice totals; the test split uses **transfer** wording, currencies and date formats |\n| Reasoning | ARC-Easy/Challenge and MMLU *(held out)*, OpenBookQA, SciQ, CommonsenseQA, QASC, HellaSwag, WinoGrande, MMLU auxiliary-train, BoolQ, GSM8K, AQuA-RAT, SNLI, MultiNLI, ANLI, SciTail | Multiple choice, yes/no, NLI, numeric answers with near-miss distractors |\n| Classification | AG News, Yelp (ordinal); DAIR Emotion and SST-5 *(held out)* | Topic; 5-level `score` sentiment |\n\n**Augmentation**: options are reshuffled every epoch (ordinal levels keep their order). In 8% of choice questions the gold answer is removed and \"None of the above\" becomes correct; in another 4% it is added as a distractor. **Balance**: tasks are sampled with p \u221d n^0.5 each epoch. **Leak guard**: training decisions whose (state, question) appears in validation or test were removed. Mind2Web was excluded (opt-in in the notebook). Held-out sources were never used for training, calibration or model selection.\n\nCheck each dataset's card for its license before redistributing derived data. AgentHarm is used only as a held-out evaluation, in line with its intended use.\n\n---\n\n## 11. Training procedure and calibration\n\n### 11.1 Setup\n\n| Setting | Value |\n|---|---|\n| Mode / preset | `scratch` from the pretrained backbone / `standard` |\n| Data caps | \u22644,000 train, \u2264300 validation, \u2264300 test decisions per source split |\n| Epochs | 2 (best: epoch 2) |\n| Objective | Strictly proper scoring rules: log score + 0.5 \u00d7 spherical score, + 1.0 \u00d7 ranked probability score for `score` questions; soft targets (50/50 with the hard label) where a teacher distribution exists. These are the RLCD rewards, optimised with exact gradients |\n| Optional RLCD stage | Off |\n| Optimiser | AdamW (\u03b2 0.9/0.98, weight decay 0.01), encoder LR 4e-5 with layer-wise decay 0.9, head LR 3e-4, 6% warm-up, cosine decay, gradient clipping 1.0 |\n| Batching | Token-budget batches (16,384 tokens, \u226432 decisions), length-bucketed |\n| Weights kept | EMA of the weights (decay 0.999), best validation task-macro accuracy |\n| Precision | bf16 autocast, TF32 matmuls; attention `auto` (FlashAttention-2 if installed, else SDPA) |\n| Sequence | 512 tokens; 2,048 above 24 options; question \u226496 tokens; \u226424 tokens per option (more for code options) |\n| Hardware / time | NVIDIA GeForce RTX 5090 Laptop GPU \u00b7 **46 minutes** |\n| Software | Python 3.14.4 \u00b7 PyTorch 2.11.0+cu128 \u00b7 transformers 5.17.0 |\n| Seed | 42 |\n\n| Epoch | Train loss | Train acc | Val macro | Val micro | Val NLL |\n|---|---|---|---|---|---|\n| 1 | 0.868 | 0.654 | 0.731 | 0.738 | 0.565 |\n| 2 | 0.540 | 0.802 | **0.763** | 0.775 | 0.520 |\n\nValidation was still improving at epoch 2, so a longer schedule (the `full` preset) is likely to help.\n\n### 11.2 Fitted temperatures\n\n| Question type | k \u2264 2 | k 3\u20135 | k 6\u201312 | k > 12 |\n|---|---|---|---|---|\n| choice | 1.707 | 1.352 | 1.466 | 1.349 |\n| noul | 1.402 | 1.402 | 1.402 | 1.402 |\n| score | 1.453 | 1.453 | 1.453 | 1.453 |\n\nAll temperatures are above 1, so the raw model was over-confident, as Laya's checkpoints are. `noul` and `score` use a single per-type temperature because their questions fall into one option-count bucket (2 options for `noul`, mostly 3\u20135 levels for `score`).\n\n---\n\n## 12. Operational notes\n\n- **Concurrency.** `decide()` isn't internally locked. Serialise calls with a lock per process (as in \u00a77.4\u20137.6) and scale out with processes.\n- **Hardware.** On a GPU, expect tens of milliseconds per call; measure yours with \u00a76.3. On CPU, use the int8 variant (loaded as fp32) and optionally dynamic int8 quantisation.\n- **Input length.** 512 tokens by default. The question and options come first, so an over-long state loses its *end*. Put the decisive information early, or summarise.\n- **Determinism.** Repeated identical calls give identical results on the same hardware and library versions. Under bf16 on GPU, batching different questions together changes padding and can move probabilities slightly (about 1e-2); near-ties can flip. Decisions with a clear margin don't change.\n- **Traceability.** Log the model version, the Hub revision and the variant (fp16/int8) with every decision.\n- **Offline use.** After the first download, set `HF_HUB_OFFLINE=1`, or save a local copy and load it by path.\n\n---\n\n## 13. Bias, risks and limitations\n\n- **Quantitative and date reasoning.** Arithmetic (AQuA 0.259), table comparisons (0.290) and counting or summing (0.523) are near chance. Overdue-invoice checks with ISO dates score 0.407 with ECE 0.364, meaning confidently wrong.\n- **Wide option sets.** Accuracy falls to 0.470 with all 77 Banking77 intents in one question; shortlist first.\n- **Generalisation gap.** Held-out tasks average 0.567 against 0.725 overall. Expect lower accuracy on traffic unlike the training mix, and measure it (\u00a76.4).\n- **Guardrail transfer.** In-distribution jailbreak detection is 0.966, but held-out prompt-injection (0.647) and AgentHarm refusal (0.654) are much lower, with ECE 0.272 and 0.178. Don't rely on it as the only safety layer.\n- **No completed head-to-head with proof_v2.** The comparisons in \u00a72.4\u20132.5 use published numbers on different samples.\n- **Closed world.** The model always picks one of your options. Add \"None of the above\" when appropriate.\n- **Distribution-dependent calibration.** ECE was measured on this test mix; re-check it on your traffic and recalibrate if needed (\u00a78.2).\n- **Adversarial input.** The state is untrusted text. The model can't be instructed like an LLM, but crafted input can shift its scores. Don't make it the only safeguard before irreversible actions.\n- **Data provenance.** Training data is English, largely crowd-sourced, templated, synthetic or scraped from public code and web tasks. Biases in these sources and in the backbone's pre-training can carry into decisions. Automated routing can systematically misroute users whose phrasing differs from the training data (dialects, non-native speakers, assistive phrasing); monitor misroutes by group where possible.\n- **Oversight.** Not for high-stakes decisions without human review.\n\n---\n\n## 14. Versioning and lineage\n\n| | proof_V_1 | proof_v2 | proof_v3 | **LightDec** |\n|---|---|---|---|---|\n| Model | falconsproof v1 | falconsproof v2.0.0 | FalconDec v1.0.0 (notebook V1) | **FalconDec v1.0.0 (notebook V2)** |\n| Backbone | DistilBERT, 128 tokens | ModernBERT-base, 384 tokens | Ettin-150M, 512 / 2,048 | Ettin-150M, 512 / 2,048 |\n| Preset / epochs | \u2014 | small / 1 per stage | small / 1 | **standard / 2** |\n| Training decisions | \u2014 | 18,795 | \u2264500 per task | **155,747** |\n| Agentic data | \u2014 | \u2014 | \u2014 | AgentTrek, Counsel, HotpotQA |\n| Test accuracy | \u2014 | 0.813 (23-task, code-heavy mix) | 0.518 micro / 0.535 macro | **0.725 / 0.725** (58-task mix) |\n| ECE | \u2014 | 0.008 | 0.014 | 0.025 |\n| Weights | \u2014 | \u2248596 MB fp32 | 319 MB fp16 | 319 MB fp16 \u00b7 161 MB int8 |\n\nLightDec is a fresh `scratch` run from the pretrained Ettin backbone (lineage: `jhu-clsp/ettin-encoder-150m`). It doesn't inherit proof_v3's or proof_v2's weights. Fine-tuning LightDec with the notebook bumps the patch version (1.0.0 \u2192 1.0.1).\n\n**Changelog.** LightDec 1.0.0: first release. Notebook V2, `standard` preset, 2 epochs, seed 42.\n\n---\n\n## 15. API reference\n\nAll functions live in `falcondec_modeling.py`.\n\n| Function | Description |\n|---|---|\n| `load_falcondec(path, device=None, dtype=None, attn_implementation=\"sdpa\")` | Loads a FalconDec directory (fp16 or int8) or Hub repo id. Returns `(model, tokenizer)`; cuda if available |\n| `decide(model, tok, state, questions, defer_threshold=None, batch_size=32)` | Typed questions about one state. `questions` is a list or a `{key: question}` dict. Returns `{\"results\": [...], \"answers\": {key: result}}` |\n| `score_items(model, tok, items, batch_size=32)` | Batch scoring. `items = [{\"state\", \"question\", \"options\", \"type\"?, \"option_tokens\"?, \"seq_len\"?}]`. Returns a calibrated probability array per item |\n| `save_falcondec(model, tok, out_dir, int8=False, extra_files=None)` | Writes a self-contained directory (copies the modeling file) |\n| `quantize_int8(state_dict)` / `dequantize_int8(state_dict)` | The per-channel int8 codec used for `compact-int8/` |\n| `assemble(...)`, `collate_features(...)` | Low-level sequence building and batching |\n\n**Question fields**: `type` (`choice` / `noul` / `score`, default `choice`); `question` or `instructions`; `options` (list) or `criteria` (dict `{key: description}` for choice, list of levels for score); `labels` (`{\"true\": \u2026, \"false\": \u2026}` wording for noul); `option_tokens` and `seq_len` (optional per-question budgets).\n\n**Model attributes**: `model.fcfg` (the live config: layout, `special`, `defer_threshold`, `version`, `lineage`); `model.temperature` (3 \u00d7 4 tensor); `model.num_parameters()`; `model.device`.\n\n---\n\n## 16. Citation and references\n\n```bibtex\n@misc{falconsai_lightdec_2026,\n title = {LightDec: a lightweight, single-pass, typed, calibrated decision model for agentic systems},\n author = {{Falconsai}},\n year = {2026},\n howpublished = {\\url{https://huggingface.co/Falconsai/LightDec}},\n note = {FalconDec architecture, Ettin-150M backbone; successor to Falconsai/proof_v3}\n}\n```\n\n**Methods.** Warner et al. (2024), *ModernBERT*. Weller et al. (2025), *Ettin* encoders. Gneiting and Raftery (2007), *Strictly Proper Scoring Rules*. Guo et al. (2017), *On Calibration of Modern Neural Networks*. Geifman and El-Yaniv (2017), *Selective Classification*. Zaheer et al. (2017), *Deep Sets*; Lee et al. (2019), *Set Transformer*. Williams (1992), REINFORCE; Shao et al. (2024), GRPO. Hinton, Vinyals and Dean (2015), *Distillation*. Related decision models: Laya (Convai Innovations), TypeSafe Jev, Together Tev1.\n\n**Data.** ARC, OpenBookQA, SciQ, CommonsenseQA, QASC, HellaSwag, WinoGrande, MMLU, BoolQ, GSM8K, AQuA-RAT, SNLI, MultiNLI, ANLI, SciTail, CLINC150, MASSIVE, Bitext, Banking77, AG News, Yelp, DAIR Emotion, SST-5, jailbreak-classification, Civil Comments, deepset prompt-injections, AgentHarm, CodeXGLUE, MBPP, HumanEval, LocalLLaMA/typed-decisions, AgentTrek, Counsel, HotpotQA.\n\nReport issues, misroutes or evaluation results through the Community tab of this repository.\n\n---\n\nThis card is generated from the surgical record itself; the package's\n`lineage.intoto.jsonl` is the signed source of truth (verify it free at\nthe Surgeon's public verifier or with the bundled `verify_attestation.py`).\n\n## Architecture\n- Identification: **NLP \u00b7 Small Language Model (SLM)** (98% confidence)\n- Source format: `safetensors` \u00b7 Intended task: not declared\n- `config.json`: synthesized from the anatomy (no source config.json); model_type omitted \u2014 no architecture name in the source (QA-F-126)\n- Source license: apache-2.0\n- Lineage chain: 1 surgery (no prior attestation reachable) \u00b7 Falconsai/LightDec\n- Post-surgery totals: 159,654,157 parameters \u00b7\n 168 tensors\n- Compute estimate: 15.02439 GFLOPs (comparison\n metric, not a measurement)\n\n## Provenance & operations\n- Parents: Falconsai/LightDec/model.safetensors\n- Operations performed: load\u00d71\n- Weight merges recorded: 0\n- Quantized tensors (F32\u2192F16): 0\n\n## Surgery Log (ordered)\n1. **load** \u2014 hub:Falconsai/LightDec/model.safetensors (319.3 MB, safetensors)\n\n## Validation\n- Tissue imaging: not run\n- Structural integrity is testable offline via the packaged\n `load_and_test.py`.\n\n## Compliance note\nThe signed attestation + this card together document model composition,\nmodification history, and validation evidence \u2014 the record structure\ntechnical-documentation obligations (e.g. EU AI Act Annex IV) ask for.\nThis is evidence, not legal advice.\n\n---\n*Operated with Model Surgeon \u2014 verify this package at https://surgeon.falcons.ai/verify*\n*\u00a9 2026 FALCONS.AI \u2014 Model Surgeon record format. The model weights remain their owner's.*\n\n",
"blocks": [
{
"lang": "bash",
"code": "pip install torch \"transformers>=4.48\" safetensors huggingface_hub numpy"
},
{
"lang": "python",
"code": "# lightdec.py\nimport importlib.util, json, shutil\nfrom pathlib import Path\nfrom huggingface_hub import snapshot_download\n\n\ndef load_lightdec(repo=\"Falconsai/LightDec\", revision=None, variant=\"fp16\", device=None, dtype=None):\n \"\"\"Returns (fdm, model, tokenizer). variant: \"fp16\" (319 MB) or \"int8\" (161 MB, dequantised on load).\"\"\"\n path = Path(repo) if Path(repo).exists() else Path(snapshot_download(repo, revision=revision))\n if variant == \"int8\":\n path = path / \"compact-int8\"\n fc = json.loads((path / \"falcondec_config.json\").read_text(encoding=\"utf-8\"))\n expected = fc.get(\"weights\", \"model.safetensors\")\n if not (path / expected).exists(): # e.g. a renamed weight file in a processed copy\n cands = sorted(path.glob(\"*.safetensors\"))\n if not cands:\n raise FileNotFoundError(f\"no .safetensors weights in {path}\")\n local = Path(\"lightdec_local\") / variant\n shutil.copytree(path, local, dirs_exist_ok=True)\n shutil.copy(cands[0], local / expected)\n path = local\n spec = importlib.util.spec_from_file_location(\"falcondec_modeling\", str(path / \"falcondec_modeling.py\"))\n fdm = importlib.util.module_from_spec(spec)\n spec.loader.exec_module(fdm)\n model, tok = fdm.load_falcondec(str(path), device=device, dtype=dtype) # cuda if available, else cpu\n return fdm, model, tok"
},
{
"lang": "python",
"code": "from lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec() # or load_lightdec(revision=\"<commit>\", variant=\"int8\")\nprint(model.fcfg[\"name\"], model.fcfg[\"version\"], round(model.num_parameters() / 1e6, 1), \"M params\")"
},
{
"lang": "python",
"code": "state = {\"from\": \"user@acme.com\", \"subject\": \"Duplicate charge on invoice #4411\",\n \"body\": \"We were billed twice for March. Please refund the duplicate today or we will cancel our plan.\"}\n\nout = fdm.decide(model, tok, state, {\n \"department\": {\"type\": \"choice\", \"instructions\": \"Which department should handle this request?\",\n \"criteria\": {\"billing\": \"invoices, payments, refunds\", \"technical\": \"bugs, outages\",\n \"sales\": \"pricing, contracts\", \"other\": \"everything else\"}},\n \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent is this request?\",\n \"criteria\": [\"not urgent\", \"soon\", \"critical deadline or blocking issue\"]},\n \"churn_risk\": {\"type\": \"noul\", \"instructions\": \"Does the user threaten to cancel or leave?\"},\n})\na = out[\"answers\"]\nprint(a[\"department\"][\"choice\"], round(a[\"department\"][\"confidence\"], 3), a[\"department\"][\"defer\"])\nprint(\"urgency level\", round(a[\"urgency\"][\"expected_level\"], 2), \"of\", 2)\nprint(\"P(churn)\", round(a[\"churn_risk\"][\"p_true\"], 3))"
},
{
"lang": "python",
"code": "import contextlib, io, sys\nimport numpy as np\nfrom lightdec import load_lightdec\n\nrev = sys.argv[1] if len(sys.argv) > 1 else None\nlog = io.StringIO()\nwith contextlib.redirect_stdout(log):\n fdm, model, tok = load_lightdec(revision=rev)\nassert \"load warning\" not in log.getvalue(), log.getvalue()\n\nfc = model.fcfg\nassert fc[\"name\"] == \"FalconDec\", fc[\"name\"] # LightDec checkpoints use the FalconDec architecture name\nT = model.temperature.float().cpu().numpy()\nassert T.shape == (3, 4) and (T > 0).all(), T\nq = {\"team\": {\"question\": \"Which team?\", \"options\": [\"recover password\", \"shipping\", \"invoicing\"]}}\na = fdm.decide(model, tok, \"I forgot my password and can't sign in.\", q)[\"answers\"][\"team\"]\nb = fdm.decide(model, tok, \"I forgot my password and can't sign in.\", q)[\"answers\"][\"team\"]\nassert all(abs(a[\"probs\"][k] - b[\"probs\"][k]) < 1e-4 for k in a[\"probs\"]), \"non-deterministic\"\nassert abs(sum(a[\"probs\"].values()) - 1) < 1e-3\nprint(f\"OK LightDec (FalconDec v{fc['version']}, notebook {fc.get('notebook_version')}) choice={a['choice']} \"\n f\"conf={a['confidence']:.3f} defer_threshold={fc.get('defer_threshold')}\")"
},
{
"lang": "python",
"code": "import os, random\nimport pytest\nfrom lightdec import load_lightdec\n\n\n@pytest.fixture(scope=\"session\")\ndef fd():\n return load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"))\n\n\ndef ask(fd, state, question, options, **kw):\n fdm, model, tok = fd\n return fdm.decide(model, tok, state, [dict(question=question, options=options, **kw)])[\"results\"][0]\n\n\ndef test_probabilities_are_valid(fd):\n r = ask(fd, \"The build failed on main.\", \"What next?\", [\"Revert\", \"Ignore\", \"Retry\"])\n assert all(0 <= p <= 1 for p in r[\"probs\"].values()) and abs(sum(r[\"probs\"].values()) - 1) < 1e-3\n\n\ndef test_typed_outputs(fd):\n fdm, model, tok = fd\n out = fdm.decide(model, tok, \"I was charged twice. Refund me or I'm leaving.\", {\n \"refund\": {\"type\": \"noul\", \"instructions\": \"Does the user ask for a refund?\"},\n \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent?\", \"criteria\": [\"low\", \"medium\", \"high\"]}})[\"answers\"]\n assert 0 <= out[\"refund\"][\"p_true\"] <= 1 and out[\"refund\"][\"choice\"] in (True, False)\n assert 0 <= out[\"urgency\"][\"expected_level\"] <= 2\n\n\ndef test_support_routing(fd):\n r = ask(fd, \"I forgot my password and the reset email never arrived.\", \"Which team should handle this?\",\n [\"recover password\", \"billing and payment\", \"delivery information\"])\n assert r[\"choice\"] == \"recover password\"\n\n\ndef test_fanout_matches_single_questions(fd):\n # Batching changes padding; under bf16 that moves probabilities slightly, never the substance.\n fdm, model, tok = fd\n state = \"I forgot my password and can't sign in.\"\n qs = [{\"question\": \"Team?\", \"options\": [\"recover password\", \"shipping\", \"invoicing\"]},\n {\"question\": \"Urgent?\", \"options\": [\"Yes\", \"No\"]}]\n together = fdm.decide(model, tok, state, qs)[\"results\"]\n for q, t in zip(qs, together):\n alone = fdm.decide(model, tok, state, [q])[\"results\"][0]\n assert all(abs(alone[\"probs\"][k] - t[\"probs\"][k]) < 2e-2 for k in alone[\"probs\"])\n\n\ndef test_option_order_is_mostly_irrelevant(fd):\n # The head is order-equivariant, but the encoder sees positions; training reshuffled options every epoch.\n state, q = \"Where is my parcel? It's three days late.\", \"What should support do?\"\n opts = [\"Give the delivery status\", \"Start a refund\", \"Book an appointment\"]\n base = ask(fd, state, q, opts)[\"choice\"]\n same = sum(ask(fd, state, q, random.Random(s).sample(opts, len(opts)))[\"choice\"] == base for s in range(5))\n assert same >= 4\n\n\ndef test_many_options_use_the_tournament(fd):\n opts = [f\"topic number {i}\" for i in range(119)] + [\"reset my password\"]\n r = ask(fd, \"I can't log in, I need to reset my password.\", \"What does the user want?\", opts)\n assert len(r[\"probs\"]) == 120 and abs(sum(r[\"probs\"].values()) - 1) < 1e-3\n\n\ndef test_int8_agrees_with_fp16(fd):\n fdm8, m8, tok8 = load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"), variant=\"int8\")\n fdm, model, tok = fd\n items = [dict(state=s, question=\"Which team?\", options=[\"billing\", \"shipping\", \"accounts\", \"technical\"])\n for s in [\"I was double charged\", \"Where is my parcel?\", \"Change my email\", \"The app crashes on start\",\n \"Refund the duplicate payment\", \"Package never arrived\", \"Reset my login\", \"Error 500 on checkout\"]]\n a = [p.argmax() for p in fdm.score_items(model, tok, items)]\n b = [p.argmax() for p in fdm8.score_items(m8, tok8, items)]\n assert sum(x == y for x, y in zip(a, b)) >= len(items) - 1"
},
{
"lang": "python",
"code": "import time, numpy as np, torch\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nq = [{\"question\": \"Route?\", \"options\": [\"billing and payment\", \"shipping\", \"recover password\"]}]\nfor _ in range(5):\n fdm.decide(model, tok, \"I was charged twice.\", q)\nt = []\nfor _ in range(100):\n if torch.cuda.is_available(): torch.cuda.synchronize()\n t0 = time.perf_counter(); fdm.decide(model, tok, \"I was charged twice.\", q)\n if torch.cuda.is_available(): torch.cuda.synchronize()\n t.append((time.perf_counter() - t0) * 1000)\nprint(f\"p50 {np.percentile(t, 50):.1f} ms p95 {np.percentile(t, 95):.1f} ms on {model.device}\")"
},
{
"lang": "json",
"code": "{\"id\": \"t1\", \"tag\": \"support\", \"state\": \"\u2026\", \"question\": \"\u2026\", \"options\": [\"\u2026\", \"\u2026\"], \"expected\": \"B\", \"type\": \"choice\"}"
},
{
"lang": "python",
"code": "import json, string\nimport numpy as np\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nrows = [json.loads(l) for l in open(\"my_eval.jsonl\", encoding=\"utf-8\") if l.strip()]\n\ndef idx(v, opts):\n if isinstance(v, int): return v\n v = str(v).strip()\n if len(v) == 1 and v.upper() in string.ascii_uppercase[:len(opts)]: return string.ascii_uppercase.index(v.upper())\n return [o.lower() for o in opts].index(v.lower())\n\nitems = [dict(state=r[\"state\"], question=r.get(\"question\", \"\"), options=r[\"options\"], type=r.get(\"type\", \"choice\"))\n for r in rows]\nprobs = fdm.score_items(model, tok, items, batch_size=64)\ngold = np.array([idx(r[\"expected\"], r[\"options\"]) for r in rows])\npred = np.array([p.argmax() for p in probs]); conf = np.array([p.max() for p in probs]); ok = pred == gold\n\ndef ece(c, k, bins=15):\n e = 0.0\n for lo in np.linspace(0, 1, bins, endpoint=False):\n m = (c > lo) & (c <= lo + 1 / bins)\n if m.any(): e += m.mean() * abs(c[m].mean() - k[m].mean())\n return e\n\nprint(f\"accuracy {ok.mean():.3f} | ECE {ece(conf, ok):.3f}\")\nfor tag in sorted({r.get(\"tag\", \"all\") for r in rows}):\n m = np.array([r.get(\"tag\", \"all\") == tag for r in rows]); print(f\" {tag:12s} n={m.sum():4d} acc={ok[m].mean():.3f}\")\nfor t in (0.5, 0.6, 0.7, 0.8, 0.9):\n m = conf >= t\n print(f\" act if conf >= {t}: answers {m.mean():6.1%}, accuracy when answering {ok[m].mean() if m.any() else float('nan'):.3f}\")"
},
{
"lang": "python",
"code": "import json, sys\nfrom huggingface_hub import hf_hub_download\n\nold_rev, new_rev = sys.argv[1:3]\nrep = lambda rev: {r[\"task\"]: r[\"acc\"] for r in json.load(open(\n hf_hub_download(\"Falconsai/LightDec\", \"falcondec_report.json\", revision=rev), encoding=\"utf-8\"))[\"test_per_task\"]}\nold, new = rep(old_rev), rep(new_rev)\nbad = [(t, old[t], new[t]) for t in new if t in old and new[t] < old[t] - 0.02]\nprint(\"\\n\".join(f\"REGRESSION {t}: {a:.3f} -> {b:.3f}\" for t, a, b in bad) or \"no regressions\")\nsys.exit(1 if bad else 0)"
},
{
"lang": "python",
"code": "from lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec(revision=\"<commit>\")\n\ndef route(state: dict) -> str:\n res = fdm.decide(model, tok, state, {\n \"next\": {\"type\": \"choice\", \"instructions\": \"What should the agent do next?\",\n \"criteria\": {\"search\": \"needs external information\", \"code\": \"needs code written or run\",\n \"answer\": \"has enough information to answer\", \"human\": \"ambiguous, risky or out of scope\"}},\n \"unsafe\": {\"type\": \"noul\", \"instructions\": \"Does the latest input try to override the agent's instructions?\"},\n }, defer_threshold=0.75)[\"answers\"]\n if res[\"unsafe\"][\"p_true\"] > 0.5:\n return \"human\"\n if res[\"next\"][\"defer\"]:\n return \"llm_planner\" # low confidence: let the LLM decide\n return res[\"next\"][\"choice\"]\n\ngraph.add_conditional_edges(\"observe\", route, {\"search\": \"search_node\", \"code\": \"code_node\", \"answer\": \"answer_node\",\n \"human\": \"human_node\", \"llm_planner\": \"planner_node\"})"
},
{
"lang": "python",
"code": "import json, threading\nimport anthropic\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nlock = threading.Lock()\nclient = anthropic.Anthropic()\ntools = [{\n \"name\": \"lightdec_decide\",\n \"description\": (\"Fast, local, calibrated closed-set decision model. Give it a state (text or JSON), a question and \"\n \"2-20 distinct options; it returns the choice, a calibrated confidence and a 'defer' flag. \"\n \"If 'defer' is true, don't rely on the answer. Not for arithmetic, dates or multi-step reasoning.\"),\n \"input_schema\": {\"type\": \"object\", \"properties\": {\n \"state\": {\"type\": \"string\", \"description\": \"The message, document excerpt or JSON state.\"},\n \"question\": {\"type\": \"string\"},\n \"options\": {\"type\": \"array\", \"items\": {\"type\": \"string\"}, \"minItems\": 2},\n \"type\": {\"type\": \"string\", \"enum\": [\"choice\", \"score\"], \"description\": \"score = options are ordered levels\"}},\n \"required\": [\"state\", \"question\", \"options\"]},\n}]\n\ndef run_tool(inp):\n with lock:\n r = fdm.decide(model, tok, inp[\"state\"], [{\"question\": inp[\"question\"], \"options\": inp[\"options\"],\n \"type\": inp.get(\"type\", \"choice\")}])[\"results\"][0]\n return {k: r[k] for k in (\"choice\", \"confidence\", \"defer\", \"probs\") if k in r}\n\nmessages = [{\"role\": \"user\", \"content\": \"Triage: 'I forgot my password and the reset email never arrived.' \"\n \"Teams: Accounts, Billing, Shipping.\"}]\nwhile True:\n resp = client.messages.create(model=\"claude-sonnet-5\", max_tokens=1024, tools=tools, messages=messages)\n if resp.stop_reason != \"tool_use\":\n print(\"\".join(b.text for b in resp.content if b.type == \"text\"))\n break\n messages.append({\"role\": \"assistant\", \"content\": resp.content})\n results = []\n for block in resp.content:\n if block.type == \"tool_use\" and block.name == \"lightdec_decide\":\n try:\n results.append({\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": json.dumps(run_tool(block.input))})\n except Exception as exc:\n results.append({\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": str(exc), \"is_error\": True})\n messages.append({\"role\": \"user\", \"content\": results})"
},
{
"lang": "python",
"code": "import os, threading\nfrom mcp.server.fastmcp import FastMCP\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec(revision=os.environ.get(\"LIGHTDEC_REVISION\"),\n variant=os.environ.get(\"LIGHTDEC_VARIANT\", \"fp16\"))\nMIN_CONF = float(os.environ.get(\"LIGHTDEC_MIN_CONF\", \"0.7\"))\nlock = threading.Lock()\nmcp = FastMCP(\"lightdec\")\n\n\n@mcp.tool()\ndef decide(state: str, question: str, options: list[str], type: str = \"choice\") -> dict:\n \"\"\"Choose one of 2-20 distinct options for a question about a state. type=\"score\" means ordered levels.\n Returns the choice, a calibrated confidence and 'defer' (true = not reliable enough to act on).\"\"\"\n with lock:\n r = fdm.decide(model, tok, state, [{\"question\": question, \"options\": options, \"type\": type}],\n defer_threshold=MIN_CONF)[\"results\"][0]\n return {k: r[k] for k in (\"choice\", \"confidence\", \"defer\", \"probs\", \"expected_level\") if k in r}\n\n\n@mcp.tool()\ndef decide_many(state: str, questions: dict) -> dict:\n \"\"\"Several typed questions about one state: {name: {\"type\": \"choice\"|\"noul\"|\"score\",\n \"instructions\": str, \"criteria\": {key: description} | [levels]}}.\"\"\"\n with lock:\n return fdm.decide(model, tok, state, questions, defer_threshold=MIN_CONF)[\"answers\"]\n\n\nif __name__ == \"__main__\":\n mcp.run()"
},
{
"lang": "json",
"code": "{\n \"mcpServers\": {\n \"lightdec\": {\n \"command\": \"C:\\\\path\\\\to\\\\python.exe\",\n \"args\": [\"C:\\\\path\\\\to\\\\lightdec_mcp.py\"],\n \"env\": { \"LIGHTDEC_REVISION\": \"<commit>\", \"LIGHTDEC_VARIANT\": \"int8\" }\n }\n }\n}"
},
{
"lang": "python",
"code": "import threading\nfrom fastapi import FastAPI, HTTPException\nfrom pydantic import BaseModel\nfrom lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nlock = threading.Lock()\napp = FastAPI(title=\"LightDec\")\n\n\nclass Decide(BaseModel):\n state: str | dict\n questions: dict | list\n defer_threshold: float = 0.7\n\n\n@app.get(\"/health\")\ndef health():\n return {\"ok\": True, \"model\": \"LightDec\", \"version\": model.fcfg.get(\"version\")}\n\n\n@app.post(\"/v1/decide\")\ndef decide(req: Decide):\n try:\n with lock:\n return fdm.decide(model, tok, req.state, req.questions, defer_threshold=req.defer_threshold)\n except (ValueError, KeyError, TypeError) as exc:\n raise HTTPException(422, str(exc))"
},
{
"lang": "python",
"code": "from lightdec import load_lightdec\n\nfdm, model, tok = load_lightdec()\nQUEUES = {\"billing\": \"charges, invoices, payments, refunds\", \"returns\": \"returning or exchanging items\",\n \"delivery\": \"shipping status, late or missing parcels\", \"accounts\": \"login, password, profile\",\n \"human\": \"complaints or requests to speak to a person\"}\n\ndef intake(message: str) -> dict:\n a = fdm.decide(model, tok, message, {\n \"queue\": {\"type\": \"choice\", \"instructions\": \"Which team should handle this message?\", \"criteria\": QUEUES},\n \"urgent\": {\"type\": \"noul\", \"instructions\": \"Does this need a reply within the hour?\"},\n })[\"answers\"]\n if a[\"queue\"][\"defer\"]:\n return {\"action\": \"human_review\", \"suggestion\": a[\"queue\"][\"choice\"],\n \"reason\": f\"low confidence ({a['queue']['confidence']:.0%})\"}\n return {\"action\": \"enqueue\", \"queue\": a[\"queue\"][\"choice\"],\n \"priority\": \"high\" if a[\"urgent\"][\"p_true\"] >= 0.5 else \"normal\",\n \"evidence\": {\"confidence\": round(a[\"queue\"][\"confidence\"], 3), \"model\": \"LightDec\"}}"
},
{
"lang": "python",
"code": "import numpy as np, torch\n# probs, gold: from \u00a76.4 (probabilities already include the stored temperatures)\nlogp = [np.log(np.clip(p, 1e-12, 1)) for p in probs]\ndef nll(s):\n return -np.mean([(lp / s)[g] - np.log(np.exp(lp / s).sum()) for lp, g in zip(logp, gold)])\ns = min(np.linspace(0.5, 3.0, 51), key=nll)\nprint(\"extra temperature\", s)\nwith torch.no_grad():\n model.temperature.mul_(float(s))\nfdm.save_falcondec(model, tok, \"lightdec_recalibrated\") # add int8=True for the compact variant"
},
{
"lang": "python",
"code": "fdm.save_falcondec(model, tok, \"lightdec_out\") # fp16\nfdm.save_falcondec(model, tok, \"lightdec_out/compact-int8\", int8=True)\n\nfrom huggingface_hub import HfApi # needs a write token: huggingface-cli login\nHfApi().upload_folder(folder_path=\"lightdec_out\", repo_id=\"Falconsai/LightDec\", commit_message=\"LightDec v1.0.x\")"
},
{
"lang": "",
"code": "[CLS] question [SEP] [MASK] option\u2081 [MASK] option\u2082 \u2026 [MASK] option\u2096 [SEP] state [SEP]\n \u2502\n Ettin-150M encoder (22 layers, hidden 768)\n \u2502\n hidden state at each [MASK] + CLS context + question-type embedding\n \u2502\n set transformer: 2 layers, 8 heads, no positional encoding \u2192 options attend to each other, order-equivariant\n \u2502\n MLP \u2192 one logit per option \u2192 \u00f7 temperature[type, option-count bucket] \u2192 softmax"
},
{
"lang": "bibtex",
"code": "@misc{falconsai_lightdec_2026,\n title = {LightDec: a lightweight, single-pass, typed, calibrated decision model for agentic systems},\n author = {{Falconsai}},\n year = {2026},\n howpublished = {\\url{https://huggingface.co/Falconsai/LightDec}},\n note = {FalconDec architecture, Ettin-150M backbone; successor to Falconsai/proof_v3}\n}"
}
]
}
},
"chain": {
"depth": 1,
"prior": null
},
"tool": "FALCONS.AI Model Surgeon V7.99",
"signature": "6491b1d562a5cd9a38a83934f708ed8226739c09fd65632ae667845f068486cc",
"copyright": "\u00a9 2026 FALCONS.AI"
},
"identification_probe": {
"input": "41 token ids (hash-mapped, vocab 50368)",
"input_shape": [
1,
41
],
"output_shape": [
1,
2304
],
"output_sample": [
-2.31253,
-1.21176,
0.92579,
1.48794,
-0.25607,
0.6288
],
"finite": true,
"degenerate": false,
"layers_executed": 7,
"layers_planned": 149,
"stats": [
{
"node": "encoder.embeddings.tok_embeddings",
"shape": [
1,
41,
768
],
"mean_abs": 0.07562,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "interact.enc.layers.0.norm1",
"shape": [
1,
41,
768
],
"mean_abs": 0.669011,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "encoder.layers.0.mlp_norm",
"shape": [
1,
41,
768
],
"mean_abs": 0.100415,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "interact.enc.layers.0.norm2",
"shape": [
1,
41,
768
],
"mean_abs": 0.649311,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "encoder.layers.0.attn.Wo",
"shape": [
1,
41,
768
],
"mean_abs": 0.777682,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "encoder.layers.0.attn.Wqkv",
"shape": [
1,
41,
2304
],
"mean_abs": 1.247892,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
},
{
"node": "encoder.layers.0.attn [attention]",
"shape": [
1,
41,
768
],
"mean_abs": 0.778897,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false,
"attention": {
"heads": 12,
"fused": true,
"q": "encoder.layers.0.attn.Wqkv",
"entropy": 1.6132,
"per_head": [
{
"head": 0,
"mass": 0.751513,
"share": 0.0804,
"entropy": 1.7443
},
{
"head": 1,
"mass": 0.96201,
"share": 0.1029,
"entropy": 1.5727
},
{
"head": 2,
"mass": 0.896842,
"share": 0.096,
"entropy": 2.4325
},
{
"head": 3,
"mass": 0.543106,
"share": 0.0581,
"entropy": 1.1946
},
{
"head": 4,
"mass": 0.864718,
"share": 0.0925,
"entropy": 1.5858
},
{
"head": 5,
"mass": 0.939124,
"share": 0.1005,
"entropy": 1.2323
},
{
"head": 6,
"mass": 0.396543,
"share": 0.0424,
"entropy": 1.1537
},
{
"head": 7,
"mass": 0.53258,
"share": 0.057,
"entropy": 2.0694
},
{
"head": 8,
"mass": 0.714835,
"share": 0.0765,
"entropy": 1.8161
},
{
"head": 9,
"mass": 0.981614,
"share": 0.105,
"entropy": 1.4514
},
{
"head": 10,
"mass": 0.879077,
"share": 0.0941,
"entropy": 1.7476
},
{
"head": 11,
"mass": 0.884806,
"share": 0.0947,
"entropy": 1.3584
}
]
}
},
{
"node": "encoder.layers.0.mlp.Wi",
"shape": [
1,
41,
2304
],
"mean_abs": 3.870604,
"sparsity": 0.0,
"dead_channels": 0,
"relu": false
}
],
"trace": [
{
"node": "encoder.embeddings.tok_embeddings",
"status": "ok"
},
{
"node": "interact.enc.layers.0.norm1",
"status": "ok"
},
{
"node": "encoder.layers.0.mlp_norm",
"status": "ok"
},
{
"node": "interact.enc.layers.0.norm2",
"status": "ok"
},
{
"node": "encoder.layers.0.attn.Wo",
"status": "ok"
},
{
"node": "encoder.layers.0.attn.Wqkv",
"status": "ok"
},
{
"node": "encoder.layers.0.attn [attention]",
"status": "ok"
},
{
"node": "encoder.layers.0.mlp.Wi",
"status": "ok"
},
{
"node": "encoder.layers.0.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.0.linear1",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.0.linear2",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.0.self_attn.out_proj",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.1.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.1.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.1.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.1.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.1.linear1",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.1.linear2",
"status": "skip: feature mismatch"
},
{
"node": "interact.enc.layers.1.self_attn.out_proj",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.2.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.2.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.2.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.2.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.3.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.3.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.3.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.3.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.4.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.4.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.4.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.4.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.5.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.5.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.5.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.5.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.6.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.6.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.6.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.6.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.7.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.7.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.7.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.7.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.8.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.8.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.8.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.8.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.9.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.9.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.9.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.9.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.10.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.10.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.10.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.10.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.11.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.11.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.11.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.11.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.12.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.12.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.12.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.12.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.13.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.13.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.13.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.13.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.14.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.14.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.14.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.14.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.15.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.15.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.15.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.15.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.16.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.16.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.16.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.16.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.17.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.17.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.17.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.17.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.18.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.18.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.18.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.18.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.19.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.19.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.19.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.19.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.20.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.20.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.20.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.20.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.21.attn.Wo",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.21.attn.Wqkv",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.21.mlp.Wi",
"status": "skip: feature mismatch"
},
{
"node": "encoder.layers.21.mlp.Wo",
"status": "skip: feature mismatch"
},
{
"node": "ctx_proj",
"status": "skip: feature mismatch"
},
{
"node": "qtype_emb",
"status": "skip: feature mismatch"
},
{
"node": "scorer.0",
"status": "skip: feature mismatch"
},
{
"node": "scorer.3",
"status": "skip: feature mismatch"
}
],
"verdict": "WEAK",
"identification_upgraded": false,
"architecture": {
"family": "transformer_nlp",
"label": "NLP \u00b7 Small Language Model (SLM)",
"confidence": 0.98,
"score": 8.92,
"tags": [],
"runners_up": [
{
"family": "mlp",
"label": "MLP / Tabular",
"score": 2.5
},
{
"family": "ssm",
"label": "State Space Model (Mamba/S4)",
"score": 1.8
}
],
"total_params": 159654157
}
},
"original_format_export": {
"file": null,
"note": "upload was already safetensors \u2014 model_edited.safetensors IS the original format"
},
"watermark": {
"text": "FALCONS.AI Model Surgeon V7.99 | session oNuktkC6 | 2026-09-26T11:51:25Z | NLP \u00b7 Small Language Model (SLM)",
"hmac_sha256": "43e443504348bac8e425ea7b47bb48ce1a7362fef0f6710571a6245633cbc1e5",
"tensor": "_falconsai_",
"decode": "((values - 0.5) * 255) -> uint8 -> utf-8"
}
}