kev-0.8b-vision-intent
A fine-tune of Kev-0.8B (Qwen3.5-0.8B-Base, LoRA + pointer head) as
System 1 of the intent browser tool. The tool takes one stated
intent per call ("click the comments link of the top story"), builds a lexical shortlist from the page, and asks this
model one question:
- ACT: which of ~24 controls the intent means (or NONE), with a screenshot carrying each control's ref mark;
- RETRIEVE: which of ~16 page text segments answers the intent (or NONE);
- VERIFY: whether a claim holds, given the top evidence segments.
The answer is a calibrated probability distribution from one forward pass: no text is generated. The tool acts above a threshold and returns candidates below it.
A larger sibling, kev-4b-vision-intent, is more accurate and about 3.5× slower.
Files
kev.js/: the kev.js bundle (q8f32 ONNX decoder with the LoRA merged, Qwen3.5's stock vision tower, pointer head, tokenizer; 1 GB), tested with kev.js 0.6.0.kev.js/manifest.jsonnames every file.- Root: the PyTorch checkpoint in Kev's format (
adapter_model.safetensors+adapter_config.jsonforQwen/Qwen3.5-0.8B-Base,head.ptwith the pointer head, the serving temperature and the training arguments).
Use
import * as ort from "onnxruntime-web/webgpu";
import { loadKev } from "@ai-ecoverse/kev.js";
const kev = await loadKev("https://huggingface.co/ai-ecoverse/kev-0.8b-vision-intent/resolve/main/kev.js", { ort, variant: "q8f32" });
const res = await kev.systemOne({
state: "Intent: open the comments of the top story\nPage: Hacker News (https://news.ycombinator.com/)",
image, // the 1024 x 576 screenshot with the shortlist's ref marks, as RGBA pixels
questions: { action: { type: "choice", instructions: "Which control does the intent mean?",
criteria: { e11: 'link "comments"', e41: 'link "77 comments" (1st of 27)', NONE: "none of these controls is the one the intent means" } } },
});
In intent: --from https://huggingface.co/ai-ecoverse/kev-0.8b-vision-intent/resolve/main/kev.js.
Settings for intent
kev.js/manifest.json carries the question wording and thresholds the tool should use with this model:
"intent": { "question": "plain", "act": 0.7, "retrieve": 0.8, "verify": 0.9 }
question: "plain": the model was trained on intent's own question ("Intent: …" / "Which control does the intent mean?" with NONE), not webrunner's menu wording.actandretrieve: act or answer when the best non-NONE choice has at least this probability.verify: yes at p ≥ 0.9, no at p ≤ 0.1, unsure in between.
Each is the lowest threshold of at least 0.7 at which intent-bench precision reaches 95% (below).
Results
intent-bench (ai-ecoverse): 400 Mind2Web test_website steps with intents written by Claude Sonnet 5.5 (200 written
knowing the control's label, 200 blind), 88 RETRIEVE and 86 VERIFY items from 11 real pages; plus game-state
RETRIEVE/VERIFY from held-out webrunner game runs (Kittens Game, Drug Wars, Armchair Bike Touring), and 25 real-run hard
cases (calls where the tool was unsure and the agent then chose a ref). kev.js 0.6.0, q8f32, WebGPU in Chrome on an
Apple M4 Max.
| kev-0.8b-vision | this model | kev-4b-vision | kev-4b-vision-intent | |
|---|---|---|---|---|
| ACT top-1, 400 (blind / informed) | 61.8 (58.5 / 65.0) | 88.2 (85.0 / 91.5) | 82.0 (77.6 / 86.3) on 200 | 89.8 (87.0 / 92.5) |
| ACT: acts at p ≥ 0.7, precision | 16%, 98.4% | 90%, 95.3% | 69%, 94.9% | 91%, 94.5% |
| RETRIEVE top-1 | 63.6 | 81.8 | 86.4 | 87.5 |
| VERIFY accuracy | 77.9 | 88.4 | 93.0 | 95.3 |
| game-state RETRIEVE / VERIFY | – | 99.5 / 91.7 | 97.4 / 90.3 | 99.1 / 96.7 |
| real-run hard cases | – | 8/25 | 8/25 | 10/25 |
| p50 latency ACT / RETRIEVE | 0.50 s / 0.15 s | 0.61 s / 0.16 s | 1.96 s / 0.70 s | 2.21 s / 0.99 s |
At the manifest thresholds, the model:
- acts on 90.2% of intents at 95.3% precision (4.2% of intents end in a wrong action);
- answers 71.6% of RETRIEVE items at 95.2% precision;
- answers 80.2% of VERIFY items at 97.1% precision (95.5% on the held-out game states).
Serving temperature: 1.25, fitted on held-out training-distribution rows.
Training
Two LoRA stages, each a rank-16 LoRA on all linear layers, warm-started from the previous adapter, plus the pointer head. Qwen3.5's stock vision tower is frozen and shipped unchanged.
- wr1: Kev-0.8B (
jaredpalmer/kev-0.8b@2256796, the shipped kev-0.8b-vision checkpoint), one epoch over 12.5K rows:- Multimodal-Mind2Web train steps in the meep-meep webrunner format (marked screenshot, ranked control menu, SHRUG when the target is missing);
- webrunner game traces labelled by Claude Sonnet 5.5;
- 1,500 rows of Kev's decision-v7 suite, replayed.
- v3 (this model): wr1 plus one epoch, lr 5e-5, batch 4 × 8 accumulation, on 16,809 rows:
- ACT: 3,000 Multimodal-Mind2Web train steps with intents written by Claude Sonnet 5.5 (informed and blind), the intent tool's lexical shortlist (24, 32 or 48 controls), NONE when the shortlist missed the target, 10% target-removed NONE copies;
- RETRIEVE/VERIFY: 4,315 items from 284 public pages on sites intent-bench does not use, with questions and claims written by Claude Sonnet 5.5;
- game state: 2,776 RETRIEVE/VERIFY items from webrunner game snapshots (held-out runs excluded);
- 2,000 webrunner-format Mind2Web rows and 1,500 Kev replay rows.
One NVIDIA H100, 27 minutes. Mind2Web's test_* splits were not used for training; intent-bench's ACT records come
from test_website. head.pt records the exact training arguments.
Limitations
- English web pages, with screenshots at 1024 × 576 carrying the intent tool's ref marks. Without the screenshot, ACT drops to about 80. A text-only sibling trained without images reaches 85.8 at 0.15 s.
- RETRIEVE is overconfident relative to ACT (about 85% precision at p ≥ 0.5). Prefer returning candidates below the threshold.
- Real-run hard cases (the calls a tool is unsure about) remain hard for every Kev size (≤ 10/25).
- With the older webrunner menu wording, it is slightly worse than wr1 (54.5 vs 57.3 on 300 Mind2Web steps). Use
question: "plain". - intent-bench's RETRIEVE and VERIFY sets are small (88 and 86 items): one item moves precision by about one point.
- A decision model for one step of a browser agent. It does not plan, and it should not act unsupervised on consequential operations (purchases, messages, deletions) without a confirmation step.
License
OpenRAIL-M (LICENSE, the BigScience Open RAIL-M License of August 18, 2022). The model is trained on
Multimodal-Mind2Web, released under OpenRAIL, so this model carries the same use-based restrictions. Any derivative
and any user of the model must comply with them as well (paragraph 5 of the license).
You agree not to use the model or its derivatives:
- (a) in any way that violates any applicable national, federal, state, local or international law or regulation;
- (b) for the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
- (c) to generate or disseminate verifiably false information and/or content with the purpose of harming others;
- (d) to generate or disseminate personal identifiable information that can be used to harm an individual;
- (e) to generate or disseminate information and/or content (e.g. images, code, posts, articles), and place it in any context (e.g. a bot generating tweets) without expressly and intelligibly disclaiming that it is machine generated;
- (f) to defame, disparage or otherwise harass others;
- (g) to impersonate or attempt to impersonate (e.g. deepfakes) others without their consent;
- (h) for fully automated decision making that adversely impacts an individual's legal rights or otherwise creates or modifies a binding, enforceable obligation;
- (i) for any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics;
- (j) to exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
- (k) for any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories;
- (l) to provide medical advice and medical results interpretation;
- (m) to generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment.
Research purpose. Mind2Web and Multimodal-Mind2Web were "collected and released solely for research purposes, with the goal of making the web more accessible via language technologies", and their authors "are strongly against any potential harmful use of the data or technology to any party". This model is released in the same spirit, as a research artifact for web agents.
Upstream licenses. The original Mind2Web is CC BY 4.0. It is
attributed here, and its citation is below. Kev-0.8B and
Qwen3.5-0.8B-Base are Apache-2.0, which permits derivatives under
other terms provided the Apache license and attribution are kept. OpenRAIL-M's grant is itself Apache-style, with use
restrictions added, so the two are compatible. The base model's Apache-2.0 text ships as
LICENSE-Apache-2.0.txt; see NOTICE. Text written by Claude was used as training
data. The captured pages are not redistributed.
Citation
@inproceedings{zheng2024seeact,
title={GPT-4V(ision) is a Generalist Web Agent, if Grounded},
author={Boyuan Zheng and Boyu Gou and Jihyung Kil and Huan Sun and Yu Su},
booktitle={Forty-first International Conference on Machine Learning},
year={2024},
url={https://openreview.net/forum?id=piecKJ2DlB}
}
@inproceedings{deng2023mind2web,
title={Mind2Web: Towards a Generalist Agent for the Web},
author={Xiang Deng and Yu Gu and Boyuan Zheng and Shijie Chen and Samuel Stevens and Boshi Wang and Huan Sun and Yu Su},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
url={https://openreview.net/forum?id=kiYqbO3wqw}
}