Post
92
๐ JackOD-9B-Coder โ a 9B merge built to FINISH agentic coding tasks. Four-way omnimerge_v2 over Qwen3.5-9B: Jack = Qwopus3.5-9B-Coder (0.30), O = Ornith-1.5-9B (0.15), D = DeltaCoder (0.55). MTP head kept.
๐ Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 โ merge / base / DeltaCoder / Qwopus / Ornith:
โก LiveCodeBench v6 (55 hard) โ 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
โ HumanEval โ 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
โ HumanEval+ โ 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
๐ค MultiPL-E โ 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
๐ IFEval โ 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
๐ฏ LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
๐ And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 ยท base 18/55 ยท Qwopus 8/55 ยท Ornith 2/55 ยท JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
๐ค tool-eval-bench hardmode, 5 seeds: JackOD 144.4 ยฑ4.7, second behind Ornith 145.6 ยฑ4.2, above base 142.0 โ all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst โ it stops too early. The merge does both.
๐ง Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 โ the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
๐งช Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
๐ฆ 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
๐ danielcherubini/Qwen3.5-DeltaCoder-9B ยท ornith-ai/Ornith-1.5-9B
๐ ManniX-ITA/JackOD-9B-Coder
๐ ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
๐ https://ollama.com/mannix/JackOD-9B-Coder
๐ Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 โ merge / base / DeltaCoder / Qwopus / Ornith:
โก LiveCodeBench v6 (55 hard) โ 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
โ HumanEval โ 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
โ HumanEval+ โ 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
๐ค MultiPL-E โ 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
๐ IFEval โ 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
๐ฏ LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
๐ And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 ยท base 18/55 ยท Qwopus 8/55 ยท Ornith 2/55 ยท JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
๐ค tool-eval-bench hardmode, 5 seeds: JackOD 144.4 ยฑ4.7, second behind Ornith 145.6 ยฑ4.2, above base 142.0 โ all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst โ it stops too early. The merge does both.
๐ง Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 โ the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
๐งช Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
๐ฆ 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
๐ danielcherubini/Qwen3.5-DeltaCoder-9B ยท ornith-ai/Ornith-1.5-9B
๐ ManniX-ITA/JackOD-9B-Coder
๐ ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
๐ https://ollama.com/mannix/JackOD-9B-Coder