Configuration Parsing Warning:In adapter_config.json: "peft.base_model_name_or_path" must be a string

D1A-E2B v0.6

D1A is a small open decision model. One document and a set of typed questions go in (yes/no, choice, score), and a calibrated probability for every option comes out, in one forward pass, with no generated text. It speaks the TypeSafe System One API (POST /v1/systemone), so the TypeSafe SDK and existing clients work against it.

v0.6 brings D1A-E2B up to the skills of D1A-E4B v0.6, at about half the size, for on-device use. E2B v0.2.1 had only general decisions. v0.6 adds, in one training run: pull-request labelling, hard, developer-tool and long-document decisions, Japanese, and agent routing. Against v0.2.1:

  • hard decisions 31.4% β†’ 62.1%;
  • developer tools 50.5% β†’ 65.4%;
  • long documents 72.7% β†’ 86.3%;
  • JGLUE 63.5% β†’ 76.9%;
  • agent-factory routing 64.2% β†’ 89.4%;
  • PR change type 40.0% β†’ 84.5% and severity 30.1% β†’ 71.2%.

On 487 pull requests newer than all of its training, it labels severity at 77.2% and change type at 87.5% (E4B v0.6: 79.1% and 90.6%). It is still behind E4B v0.6 on almost every suite, and it never predicted P0 or P1 on those newer PRs. See Results and Limits.

Release exception. v0.6 fails one line of the bar registered for it before training (jonpol01/d1a#229). The bar required hard-v1 and devtools-v1 to rise by at least 3 points (they rose 30.7 and 14.9), and no card suite to fall by more than 1. General decisions (decision-v7) fell by 1.4 [βˆ’2.8, +0.1], from 82.2% to 80.8% (30 fixes, 48 breaks). D1A's general release rule (scripts/decide.py) passes it: βˆ’1.4 is inside its βˆ’2.0 card-suite limit, and pooled over all 21 lines v0.6 βˆ’ v0.2.1 = +18.9 [+17.4, +20.4]. It ships by the owner's decision of 2026-10-11, so that both D1A sizes exist as MLX and Core ML builds and can be compared on device. Stay on v0.2.1-2epoch-calibrated if general decisions are all you need.

The quality gate's verdict was INCOMPLETE only because its harness could not score the photo, voice and video demos on v0.2.1's build (that tag has no media/). A separate HTTP check scored all 14 (below). scripts/decide.py also asks a person to read the demo answers that changed (V2). The lead read them on 2026-10-11 against E4B v0.6's answers. Of the 29 that changed, 21 move to E4B v0.6's answer or toward its score: routing 3, inbox 6, PR labeller 5, the evals ratings 4, bulk 1, rerank 1, and the Japanese payment gate (deny β†’ ask). 5 move away: the seventh inbox email is now "urgent" (English and Japanese), the ambiguous refund case is "replace" where E4B says "deny", and two blast-radius calls are broader. One matters for safety. The English tool-gate example is a $4,999 payments.create call, made during a task that only investigates a failed invoice. v0.2.1 said deny and E4B v0.6 says ask (0.95), but E2B v0.6 gives a three-way near-tie: allow 0.36, ask 0.35, deny 0.29. Read by its top answer alone, it would allow the call. See Limits.

What's inside

A LoRA adapter (rank 16 on the attention and MLP projections: q, k, v, o, gate, up, down) and a small pointer head (256-dimensional) on google/gemma-4-E2B at revision d29ff6b4.

Version Stage Training data Settings
v0.1 / v0.2 General decisions decision-v7 suite: ~12,500 records from public classification datasets plus generated policy and rule data 1 / 2 epochs, lr 5e-5; v0.2.1-2epoch-calibrated adds the calibration temperature (1.52)
v0.6 Every later D1A skill, in one run PR labels: 4,000 English (severity-stratified), 753 Japanese, 1,500 with a blast-radius label, and 479 hermes-agent PRs created 2026-10-02 to 10-05 shown twice (once in each half); hard-v1 4,000; devtools-v1 4,000; documents-v1 2,500; JGLUE 2,000; routing: agent factory 1,216, generic 1,000; decision-v7 replay 1,500: 23,427 records from v0.2.1, 1 epoch, lr 5e-5, states up to 5,120 tokens, 2,929 steps (1 h 43 m on one NVIDIA L40S)

E4B reached the same skills over several stages (v0.2 to v0.6). E2B gets them in one mix at the final proportions, so no stage has to replay an earlier one. The version number jumps from v0.2 to v0.6 to match the E4B release with the same skills. Before training, the mix was screened against every evaluation partition, the 487 newer PRs below included: none of their documents or PRs is in it.

Calibration: a single temperature, T = 1.45 (90% interval [1.38, 1.52]), fitted on pooled rows of the decision-v7 calibration split and the PR-labelling development set. Calibration error is 0.051 before and 0.024 after, out of fold. It is stored in d1a_config.json and head.pt, and there are no per-use-case temperatures.

Results

Pull requests newer than all training

487 hermes-agent pull requests created after every PR in v0.6's training, held back before it was trained. All models are MLX 8-bit at their served temperatures, with the serving context. Differences are paired over PRs, with 95% bootstrap intervals. E4B v0.6 is shown for reference:

Question Majority answer E2B v0.2.1 E2B v0.6 E4B v0.6 E2B v0.6 βˆ’ v0.2.1
Severity (P0–P4) 67.8% 33.5% 77.2% 79.1% +43.7 [+38.2, +49.3]
Change type 51.5% 26.9% 87.5% 90.6% +60.6 [+56.1, +64.9]
P0 recall (7 PRs) 0/7 0/7 5/7
P1 recall (23 PRs) 0/23 0/23 8/23
Security type recall (14 PRs) 0/14 12/14 9/14

Every suite, against E2B v0.2.1

Held-out data only. Both E2B models are MLX 8-bit (JohnP1/d1a-e2b-mlx-q8) at their served temperatures (v0.2.1 1.52, v0.6 1.45), scored with d1a.eval.benchmark through D1A's quality gate. Differences are paired over questions, with 95% bootstrap intervals. E4B v0.6's numbers (its own card) are for reference:

Set E2B v0.2.1 E2B v0.6 Ξ” [95% CI] Calibration error E4B v0.6
Hard decisions (hard-v1, 1,083 questions) 31.4% 62.1% +30.7 [+26.9, +34.4] 0.187 β†’ 0.056 68.8%
Developer-tool decisions (devtools-v1, 1,074) 50.5% 65.4% +14.9 [+12.1, +18.1] 0.157 β†’ 0.042 70.1%
Long documents (documents-v1, 920) 72.7% 86.3% +13.6 [+10.9, +16.4] 0.017 β†’ 0.036 89.5%
Held-out transfer-v4 (656, sources never trained on) 60.2% 63.6% +3.4 [+0.6, +6.1] 0.077 β†’ 0.077 73.0%*
General decisions (decision-v7, 1,264) 82.2% 80.8% βˆ’1.4 [βˆ’2.8, +0.1] 0.030 β†’ 0.042 83.9%
JGLUE development / test (1,500 each, Japanese) 63.5% / 62.8% 76.9% / 75.3% +13.3 / +12.5 0.049 β†’ 0.037 / 0.045 β†’ 0.032 81.7% / 80.5%
Agent factory development (640) 64.2% 89.4% +25.2 [+21.0, +29.0] 0.031 β†’ 0.091 93.9%
Model routing: generic (270) / hand-labelled (45) 62.2% / 66.7% 97.8% / 93.3% +35.6 / +26.7 0.076 β†’ 0.022 / 0.080 β†’ 0.064 98.5% / 97.8%
External: semif-v1 (144) / typesafe-v1 (102) 72.9% / 53.9% 76.4% / 53.9% +3.5 / +0.0 0.090 β†’ 0.103 / 0.106 β†’ 0.212 83.3%* / 74.5%
External: wanli-v1 (256) / wanli-v2 (1,002) 64.5% / 60.0% 63.3% / 61.1% βˆ’1.2 [βˆ’5.1, +2.3] / +1.1 0.153 β†’ 0.167 / 0.180 β†’ 0.181 70.3% / 67.1%
Dates (900) / unknowable facts (525) / assertions (1,020) 83.4% / 64.2% / 90.4% 80.9% / 68.6% / 93.9% βˆ’2.6 [βˆ’5.0, βˆ’0.1] / +4.4 / +3.5 0.043 β†’ 0.109 / 0.212 β†’ 0.111 / 0.070 β†’ 0.053 92.1% / 67.8% / 94.9%
PR labels, English test (953 PRs): change type / severity / blast radius (139) 40.0% / 30.1% / 34.5% 84.5% / 71.2% / 53.2% +44.5 / +41.1 / +18.7 0.077 β†’ 0.032 85.3% / 77.9% / 61.9%
PR labels, English development (946 PRs): change type / severity / blast radius (138) 40.7% / 32.9% / 31.2% 87.8% / 69.9% / 49.3% +47.1 / +37.0 / +18.1 0.053 β†’ 0.012 90.7% / 72.5% / 52.9%
PR labels, Japanese test (92 PRs, all questions) 35.2% 72.0% +36.8 [+28.5, +45.7] 0.070 β†’ 0.054 72.5%
Owner's labelling job, replayed (193 questions, review-bot labels): type / blast radius 40.0% / 45.9% 74.7% / 57.1% +22.8 [+14.9, +31.1] overall 90.5% / 66.3%
Recent PRs from the owner's repositories, labels checked by hand (109 questions): type / blast radius / severity 55.3% / 56.4% / 15.6% 76.3% / 69.2% / 65.6% +26.6 [+14.7, +38.9] overall 86.8% / 66.7% / 65.6%

* Not like for like: E4B v0.6's card scored transfer-v4 on 764 questions and semif-v1 on 252, different partitions from the 656 and 144 here.

Pooled over all 21 lines, v0.6 βˆ’ v0.2.1 = +18.9 [+17.4, +20.4]. Latency is unchanged against v0.2.1: v0.6 over v0.2.1 on short requests is 1.0013, within run-to-run noise (measured interleaved).

Calibration: error is lower than v0.2.1's on 12 of the 22 suites, most on hard and developer-tool decisions. It is higher on 10, most on typesafe-v1 (0.106 β†’ 0.212), dates (0.043 β†’ 0.109), agent factory (0.031 β†’ 0.091), long documents (0.017 β†’ 0.036) and decision-v7 (0.030 β†’ 0.042).

Photos, voice and video

With the media extra, d1a.serving.serve answers questions about a photo, a voice clip or a video with this same model through Gemma 4 E2B's own vision and audio encoders, unchanged by the adapter (POST /v1/systemone/media). No media in training (zero-shot). Zero-shot, measured on only the playground's 14 photo, voice and video demo requests. v0.6 answers all 14, and calls 2 of the 3 intact photos damaged (the locker and the mailbox):

  • 26 of 28 answers match E2B v0.2.1. The changed two: the damaged wet-parcel video is now called damaged (right), and the intact-mailbox photo is now called damaged too (wrong).
  • 25 of 28 match E4B v0.6. The intact-locker and intact-mailbox photos are called damaged, and the mailbox's place is "in a delivery locker". So it calls intact items damaged more often than E4B does.

The questions it was trained for

PR labelling works best with these questions, worded exactly so, over a document of the form title / author / stats / body / files (one - status path +added/-deleted line per file):

  • type (choice): "Primary change type from files and body, not the title prefix." Options: type/bug, type/docs, type/feature, type/perf, type/refactor, type/security, type/test.
  • blast (choice): "How far a mistake in this PR spreads in production." Options: review:blast-contained (one module), -moderate (one subsystem), -broad (shared helper or config), -massive (auth, permissions, or all paths).
  • sev (choice): "How serious the problem this PR addresses is β€” not the risk of merging the diff as-is." Options P0 to P4.

The PR-labeler recipe has the exact wording, the document builder and the training and scoring scripts. The mix builder builds and screens mixes like v0.6's.

Run it

pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a"
python -m d1a.serving.serve --run JohnP1/d1a-e2b@v0.6 --port 8009

or in-process: from d1a import D1A; D1A.load("JohnP1/d1a-e2b@v0.6").decide(state, questions). On Apple Silicon the 8-bit MLX build, JohnP1/d1a-e2b-mlx-q8 v0.6, is faster and smaller (4.4 GB with the photo, voice and video encoders). For general decisions only, @v0.2.1-2epoch-calibrated is 1.4 points better on decision-v7.

Limits

  • Never act on its raw top answer for tool gating. On the playground's $4,999 payment example it is a near-tie (allow 0.36, ask 0.35, deny 0.29), so the top answer allows the call. Gate tool calls through D1A's fail-safe thresholds instead (d1a.agents.presets.advise: deny at p β‰₯ 0.5, allow only at p β‰₯ 0.8, otherwise ask), which turn this case into "ask".
  • Not a replacement for E4B v0.6 as a PR labeller. On the 487 newer PRs it never predicted P0 or P1: P0/P1 recall is 0/30 (E4B v0.6: 13/30). On the frozen English test set it finds 4 of 47 P1 PRs (11 P1 calls) and never predicts P0. With severity offsets fitted on the development set it finds 26 of 47, but at a precision of 0.19. Use E4B for severity, or flag a PR when p(P0) + p(P1) is high instead of taking the top answer.
  • Behind E4B v0.6 on almost every suite: hard decisions 62.1% against 68.8%, developer tools 65.4% against 70.1%, transfer 63.6% against 73.0%, JGLUE about 5 points lower, typesafe-v1 53.9% against 74.5%, dates 80.9% against 92.1%.
  • General decisions are 1.4 points below v0.2.1 (the release exception above), and dates 2.6 points below: 83.4% β†’ 80.9%, βˆ’2.6 [βˆ’5.0, βˆ’0.1].
  • Calibration error rose on 10 of 22 suites. On typesafe-v1, TypeSafe's own suite for its hosted API, it went from 0.106 to 0.212. On dates it went from 0.043 to 0.109, and the share of answers given at 90% confidence or more fell from 0.36 to 0.10. It also rose on agent factory, long documents and decision-v7 (numbers above). Probabilities on those sets are less trustworthy than v0.2.1's.
  • Media: zero-shot, measured on 14 demo requests only. It calls 2 of 3 intact photos damaged, more often than E4B v0.6 does (above).
  • The playground's evals demo grades answers less sharply than v0.2.1 did. For the English "good answer", error goes from 0.088 to 0.153 and quality from 3.81 to 3.51 (out of 5).
  • On the owner's review-bot replay it calls security 5 times, where v0.2.1 called it once.
  • One temperature for every question, and no use-case temperatures.
  • The recent PR labels come from one large open-source project and a few days of it. If its conventions move again, so will these numbers.
  • Questions are in English. The documents can be English or Japanese.

Versions

Tag What it adds
v0.1, v0.1-1epoch general decisions (decision-v7, 1 epoch)
v0.2, v0.2-2epoch general decisions, 2 epochs
v0.2.1-2epoch-calibrated the same weights with a calibration temperature (1.52)
v0.6 every later D1A skill in one run: PR labels, hard, developer-tool and long-document decisions, Japanese, agent routing; a release exception (decision-v7 βˆ’1.4)

License and data

Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0). Japanese decision data derived from JGLUE by Yahoo Japan Corporation and Waseda University (CC BY-SA 4.0). Pull-request data from NousResearch/hermes-agent (MIT); the labelled PR dataset itself, including the recent PRs, is private.

Downloads last month
39
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for JohnP1/d1a-e2b

Adapter
(46)
this model
Quantizations
1 model