Instructions to use andyzhang232/ajev-gemma4-26b-a4b-lora1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use andyzhang232/ajev-gemma4-26b-a4b-lora1 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
AJev · Gemma 4 26B-A4B LoRA
AJev is an open decision model in the style of Jev. You give it some context (text or JSON) and a few typed questions: yes / no, single choice (up to 255 options), or an ordered score. It returns a calibrated probability for every option. Each question takes a single forward pass; no text is generated.
This repository holds a LoRA adapter for google/gemma-4-26B-A4B-it. The per-type calibration temperatures are in ajev_lm_config.json. Inference code: github.com/cmzy/ajev-infer.
Results
Self-run on the full Jev Decision Index 0.2.1 suite with the official kit:
| Model | Base | Decision Index |
|---|---|---|
| Jev (TypeSafe, hosted service) | — | 57.91 |
| Surogate Rune 26B-A4B v3 | Gemma 4 26B-A4B | 57.44 |
| AJev 26B-A4B (this model) | Gemma 4 26B-A4B | 57.42 |
| AJev lora5 (previous version) | Gemma 4 12B | 52.22 |
By area: Knowledge 41.8, Language 62.5, Retrieval 63.5, Tools 73.2, Arts 43.7. These are our own full-suite results (all 150,317 scoreable requests answered, HLE included) and have not yet been reproduced by the maintainers.
Accuracy on held-out test sets:
| Test set | AJev 12B (lora5) | This model |
|---|---|---|
| JevBench public | 0.853 | 0.887 |
| Kev transfer v9 | 0.771 | 0.790 |
| eikos heldout | 0.931 | 0.937 |
| typed-decisions | 0.790 | 0.780 |
Median latency for sequential requests on one RTX PRO 6000 is 49 ms.
Usage
pip install "ajev-infer @ git+https://github.com/cmzy/ajev-infer"
from ajev.lm.predictor import LMPredictor
from ajev.schema import decisions_from_jev, jev_answer
p = LMPredictor("google/gemma-4-26B-A4B-it", adapter="andyzhang232/ajev-gemma4-26b-a4b-lora1")
ds = decisions_from_jev(
{"ticket": "I was charged twice for order #1182."},
{"topic": {"type": "choice", "instructions": "What is the ticket about?",
"criteria": {"billing": "charges, refunds", "delivery": "shipping", "account": "login"}},
"escalate": {"type": "noul", "instructions": "Should a human agent take this now?"}})
for d, probs in zip(ds, p.predict(ds)):
print(d.meta["question_id"], jev_answer(d, probs))
Jev-compatible server (POST /v1/systemone):
pip install "ajev-infer[vllm] @ git+https://github.com/cmzy/ajev-infer"
python -m ajev.serve_vllm --base-model google/gemma-4-26B-A4B-it \
--adapter andyzhang232/ajev-gemma4-26b-a4b-lora1 --port 8000
- Requires transformers ≥ 5.17.
- bf16 inference needs about 55 GB of GPU memory.
- Load it as a LoRA adapter, or merge it in memory only. Do not save a merged model and reload it.
Training
- Data: about 68k questions. Sources are public classification, NLI and decision datasets, procedurally generated business-rule questions, and the training splits of benchmarks related to the leaderboard (e.g. HoVer, VAST, POP909, ContractNLI, ACOS, BANKING77, CLINC150).
- Decontamination: every training question was checked against the full leaderboard suite, and any with text overlap was removed.
- Setup: LoRA r 32 / α 64, learning rate 3e-5, 1 epoch, one RTX PRO 6000.
Limitations
- Knowledge-reasoning benchmarks such as GPQA and ChessBench, and some Arts benchmarks, are still weak.
- Probabilities were calibrated on our own held-out data. If your data distribution is very different, recalibrate.
- Some training data carries non-commercial or attribution terms. Check them yourself before commercial use.
- Not affiliated with TypeSafe AI.
- Downloads last month
- 28