tiny-superfast-agentic-MLM-6.5m-v1: a 6.4M-parameter voice-command parser for laptops and Raspberry Pi
tiny-superfast-agentic-MLM-6.5m-v1 turns one short spoken or typed command into a structured command that a program can execute. It understands English and Hinglish (Roman-script Hindi-English), with some Devanagari Hindi, and covers 371 actions (audio, display, Wi-Fi, apps, files, timers, power, network, services, Raspberry Pi GPIO/I2C and more), and runs on a plain CPU with NumPy only: no GPU, no PyTorch and no internet needed at inference.
ENGLISH "set the volume to 40" -> set_volume {"value": 40}
HINGLISH "volume 40 kar do" -> set_volume {"value": 40}
ENGLISH "connect to wifi Redmi Note 12 password hello@123" -> connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"}
HINGLISH "pwd Tiger@2025 daalo aur Hostel Wing C se jud jao" -> connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"}
HINGLISH "10 min baad shutdown kar dena" -> shutdown {"amount": 10, "unit": "min"}
HINGLISH "board pin 11 pe led jalao" -> gpio_on {"pin": 11, "numbering": "board"}
HINDI "5 मिनट का टाइमर लगाओ" -> set_timer {"amount": 5, "unit": "min"}
| Parameters | 6,441,472 (≈6.4M; "6.5m" in the name is rounded up) |
| Model file | model/model.npz, 23.9 MB, float32. pytorch/model.pt is included for fine-tuning |
| Runs on | CPU only, fully offline: Windows, Linux, Raspberry Pi 5. Inference needs only numpy |
| Accuracy | 97.53% exact on 6,611 held-out-template commands (test_strict); 99.10% on validation; 100% on 461 unseen Wi-Fi phrasings |
| Speed | median 18 ms per command on a laptop i7-1360P, 30 ms on a Raspberry Pi 5 (10 min sustained load) |
| Load time | 0.13 s (laptop), 0.20 s (Raspberry Pi 5) |
| Bigger sibling | tiny-agentic-home-robotic-for-edge-device-v10 (23.9M params, MiniLM-based) |
About the name. "MLM" here means micro language model. It is not a masked language model like BERT: the architecture is a small Transformer encoder-decoder (seq2seq) that writes the command one token at a time.
Executor notice. The model only parses commands; it never runs anything. The executor in this repository is a sample for testing, not a finished product. It gates every action by risk (safe / caution / critical) and asks for confirmation before anything risky. Use it on a test machine and read the command before you run it.
Best use: a personal voice companion for small commands
The model is built to be the "understanding" step of a small, private, always-available assistant: you speak a short command, and your computer or Raspberry Pi does it. It replaces a large LLM for common, well-defined device commands, so the answer is instant (tens of milliseconds), private (nothing leaves the device) and free.
Attach any speech recognizer (ASR) in front of it and an executor behind it:
flowchart LR
U(["User speaks<br/>'volume 40 kar do'"]) --> A["ASR<br/>speech → text<br/>(Whisper, Vosk, ...)"]
A -->|"text"| M["<b>tiny-superfast-agentic-MLM-6.5m-v1</b><br/>text → structured command<br/>6.4M params · ~20–70 ms on CPU"]
M -->|"set_volume {value: 40}<br/>+ confidence"| E["Executor<br/>validate → risk gate → OS command"]
E --> O(["Output action<br/>volume set to 40%"])
E -. "clarify / unknown /<br/>low confidence" .-> U
If your viewer does not render Mermaid:
user voice ==> [ ASR ] --text--> [ tiny-superfast-agentic-MLM-6.5m-v1 ] --command JSON--> [ executor ] ==> output action
|
ask the user again <---- clarify / unknown / low confidence
| Stage | What it does | Example |
|---|---|---|
| 1. ASR (you bring it) | Turns speech into text. Any engine works: Whisper / faster-whisper, Vosk, a phone keyboard, or a chat box. | audio → volume 40 kar do |
| 2. This model | Turns text into one action from the 371-action catalog, plus its typed arguments and a confidence score. | → {"action": "set_volume", "value": 40}, confidence 0.997 |
| 3. Executor (sample included) | Checks the arguments, applies the risk gate, renders a whitelisted OS command for Windows or Raspberry Pi, and runs it or asks first. | → Pi: wpctl set-volume @DEFAULT_AUDIO_SINK@ 40% |
| 4. Output action | The device acts, and the result can be read back to the user (TTS). | volume is now 40% |
A good agent pattern: execute only when confidence is high and the risk gate allows it, ask the user when the result
is clarify, and hand unknown or low-confidence text to a bigger model or back to the user.
Minimal voice loop (example)
# pip install numpy faster-whisper (ASR is your choice; this is one option)
import os; os.environ.setdefault("OPENBLAS_NUM_THREADS", "4") # "2" on a Raspberry Pi 5, see Speed
import sys; sys.path[:0] = ["model", "executor"]
from runtime import CommandModel
import executor as EX
from faster_whisper import WhisperModel
asr = WhisperModel("small", device="cpu", compute_type="int8")
parser = CommandModel("model")
def handle(wav_path):
text = " ".join(s.text for s in asr.transcribe(wav_path)[0]).strip() # 1. speech -> text
out = parser.predict(text); meta = out.pop("_meta") # 2. text -> command
plan = EX.plan(out, None, text) # 3. command -> gated OS command
if plan.action == "clarify" or meta["confidence"] < 0.8:
return f"Sorry, can you say that again? ({plan.note or text})"
ok, why = EX.gate(plan, allow_caution=False) # safe actions only in this demo
return EX.run(plan, out, dry_run=not ok) # 4. act (dry run if gated)
The ASR part of this snippet is an illustration. The model was trained on typed/synthetic text, not on real ASR transcripts, so test it with your own ASR output (see Limitations).
What it understands
Every example is shown in ENGLISH, then the same command in HINGLISH. The outputs, confidences and executor plans are real outputs of this model and the included sample executor.
| Use case | Language | Command | Model output | Confidence | Risk | Raspberry Pi command |
|---|---|---|---|---|---|---|
| Audio | ENGLISH | set the volume to 40 |
set_volume {"value": 40} |
0.996 | safe | wpctl set-volume … 40% |
| Audio | HINGLISH | volume 40 kar do |
set_volume {"value": 40} |
0.997 | safe | wpctl set-volume … 40% |
| Wi-Fi | ENGLISH | connect to wifi Redmi Note 12 password hello@123 |
connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"} |
1.000 | caution | nmcli dev wifi connect 'Redmi Note 12' password hello@123 |
| Wi-Fi | HINGLISH | Redmi Note 12 wifi se connect karo password hello@123 |
same | 1.000 | caution | same |
| Wi-Fi, password first | ENGLISH | password Tiger@2025 then join Hostel Wing C |
connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"} |
1.000 | caution | nmcli dev wifi connect 'Hostel Wing C' … |
| Wi-Fi, password first | HINGLISH | pwd Tiger@2025 daalo aur Hostel Wing C se jud jao |
same | 1.000 | caution | same |
| Apps | ENGLISH | open chrome |
open_app {"app": "chrome"} |
0.990 | caution | chromium-browser & |
| Apps | HINGLISH | chrome kholo |
open_app {"app": "chrome"} |
0.986 | caution | chromium-browser & |
| Timers | ENGLISH | set a timer for 5 minutes |
set_timer {"amount": 5, "unit": "min"} |
0.998 | safe | built-in timer |
| Timers | HINGLISH | 5 minute ka timer lagao |
set_timer {"amount": 5, "unit": "min"} |
0.998 | safe | built-in timer |
| Power | ENGLISH | shut down in 10 minutes |
shutdown {"amount": 10, "unit": "min"} |
0.998 | critical | sudo shutdown -h +10 |
| Power | HINGLISH | 10 min baad shutdown kar dena |
shutdown {"amount": 10, "unit": "min"} |
0.998 | critical | sudo shutdown -h +10 |
| Negation | ENGLISH | don't shut down the laptop |
cancel_shutdown {} |
0.995 | caution | shutdown -c |
| Negation | HINGLISH | shutdown mat karo |
cancel_shutdown {} |
0.995 | caution | shutdown -c |
| Pi GPIO | ENGLISH | turn on board pin 11 |
gpio_on {"pin": 11, "numbering": "board"} |
0.999 | critical | sudo pinctrl set 17 op dh |
| Pi GPIO | HINGLISH | board pin 11 pe led jalao |
gpio_on {"pin": 11, "numbering": "board"} |
0.998 | critical | sudo pinctrl set 17 op dh |
| Pi GPIO, ambiguous pin | HINGLISH | pin no 11 high karo |
gpio_on {"pin": 11} |
0.993 | → clarify | executor asks: board pin 11 (= GPIO 17) or GPIO 11 (= board pin 23)? |
| Diagnostics | ENGLISH | what is the cpu temperature |
get_temperature {} |
0.995 | safe | vcgencmd measure_temp |
| Diagnostics | HINGLISH | cpu ka temperature batao |
get_temperature {} |
0.996 | safe | vcgencmd measure_temp |
| Files | HINGLISH | reports naam ka folder banao |
create_folder {"path": "reports"} |
0.999 | caution | mkdir -p -- reports |
| Out of scope | ENGLISH | tell me a joke |
unknown {} |
0.993 | safe | nothing |
| Out of scope (miss) | HINGLISH | kya haal hai bhai |
recent_files {} |
0.604 | safe | (low confidence: ask or hand off) |
The last row is a real miss: small talk went to a harmless action, but with low confidence. This is why the agent should check the confidence before acting.
Test results
Accuracy
Held-out evaluation after training (raw model output, before the runtime "snap" repair). "Exact" means the action and every argument match; "Action" means only the action matches.
| Set | n | Exact | Action | What it tests |
|---|---|---|---|---|
| validation | 10,413 | 99.10% | 99.29% | same templates as training, unseen rows |
| test | 10,618 | 98.37% | 98.52% | held-out templates |
| test_strict | 6,611 | 97.53% | 97.70% | held-out templates with no paraphrase sibling in training (the honest generalisation number) |
| curated | 371 | 99.46% | 99.73% | one held-out prompt per action |
| B_typo | 3,000 | 99.17% | 99.50% | typos |
| wifi_holdout | 461 | 100.00% | 100.00% | unseen Wi-Fi connect phrasings |
| gap11_holdout | 696 | 99.28% | 100.00% | unseen templates: password-first Wi-Fi, CPU clock vs usage, board↔BCM pins, fan vs mouse speed |
| A_challenge | 1,152 | 97.40% | 98.44% | external challenge set |
| B_golden_final | 384 | 97.40% | 97.40% | external golden set, including out-of-scope prompts |
| B_golden_dev | 383 | 95.04% | 95.82% | external golden set |
| A_practical-dev | 1,321 | 95.31% | 96.44% | external practical set |
| A_practical-holdout | 428 | 92.06% | 94.16% | external practical set; some misses are label-convention differences |
The same model was also run on the laptop and on the Raspberry Pi 5 (2,340 prompts, with the runtime snap). Both devices produce identical predictions: 97.82% exact, 98.46% action.
Speed and temperature
10 minutes of continuous single-stream inference on each device. Full per-pass and per-minute tables are in
eval/BENCHMARK.md.
| Laptop (Intel i7-1360P) | Raspberry Pi 5 (16 GB) | |
|---|---|---|
| Latency p50 / p95 / p99 | 18.4 / 168.6 / 239.6 ms | 30.2 / 243.9 / 281.6 ms |
| Throughput, cool → after 10 min | 29.6 → 18.5 commands/s | 13.9 → 13.9 commands/s |
| Temperature idle → peak | 52.6 → 91.9 °C | 38.5 → 59.8 °C |
| Throttling | slows down ~39% once hot | none (2400 MHz throughout, get_throttled=0x0) |
| Best BLAS threads | 4 | 2 |
- Latency grows with the length of the command the model writes: short commands take ~10–35 ms, Wi-Fi commands with an SSID and a password take ~100–250 ms.
- Set the BLAS thread count. The matrices are small, so NumPy's default (one thread per core) is up to 4x slower.
Use
OPENBLAS_NUM_THREADS=4on x86 laptops andOPENBLAS_NUM_THREADS=2on the Raspberry Pi 5 (quickstart.pydoes this for you).
Test devices
| Laptop | Raspberry Pi 5 | |
|---|---|---|
| CPU | Intel Core i7-1360P (4P + 8E cores) | Broadcom BCM2712, 4× Cortex-A76 @ 2.4 GHz |
| RAM | 32 GB | 16 GB |
| OS | Windows 11 | Raspberry Pi OS (Debian 13), kernel 6.12 |
| Python / NumPy | 3.13 / 2.5.3 | 3.13 / 2.2.4 |
| Cooling | built-in fan, AC power | active cooler (PWM fan) |
Download and test
1. Get the files
pip install -U huggingface_hub
hf download sraivante/tiny-superfast-agentic-MLM-6.5m-v1 --local-dir tiny-agentic-6.5m
cd tiny-agentic-6.5m
2. Parse commands (nothing is executed)
pip install numpy
python quickstart.py "volume 40 kar do" "connect to wifi Redmi Note 12 password hello@123"
python quickstart.py # interactive
Each line prints the command, the confidence, the latency, the risk level and the OS command the sample executor would run on this machine.
3. Test UI and checks (Python 3.9+)
| Windows | Raspberry Pi / Linux | |
|---|---|---|
| Test UI in the browser (http://localhost:8000) | run.bat |
./run.sh |
| Terminal agent | run.bat cli |
./run.sh cli |
| Smoke test (model + executor, nothing executed) | run.bat test / run.bat test --full |
./run.sh test / ./run.sh test --full |
| HTML report of all 371 actions | run.bat report |
./run.sh report |
The first run creates .venv and installs numpy, fastapi and uvicorn. report really runs the safe,
read-only actions on your machine (for example reading the volume or the IP address) and dry-runs everything else.
In the UI, Run all is always a dry run. Execute on a row runs that one command: safe actions run directly, caution actions need the "allow caution" box, and critical actions ask you to type the action name. The UI listens on 127.0.0.1 only and has no login, so do not expose it to a network.
4. Reproduce the speed test
cd eval
sh bench_run_all.sh python ../model <dir with eval .jsonl sets> <out_dir> default 1 2 4
How the model works
- Input: one command, up to 128 characters, normalised and encoded character by character (200-symbol vocabulary), so typos, code-mixing and odd spellings are handled without a word tokenizer.
- Network: Transformer encoder-decoder, d_model 256, 8 heads, FFN 768, 5 encoder + 3 decoder layers (pre-LN).
- Output: a short canonical command such as
connect_wifi ssid="…" password="…", written token by token (1,052-token output vocabulary). - Grammar-constrained decoding: at every step only tokens that keep the output valid are allowed. The action must come from the 371-action catalog, the slots must belong to that action, and text slots (SSIDs, passwords, file names, apps) must be copied from the input. The model cannot invent an action or a password that you did not say.
- Snap: a small post-decode repair extends copied SSIDs and passwords to whole tokens and recovers the original letter case from the input.
- Training: from random initialisation (no pretrained base), 1,230,162 synthetic English/Hinglish rows, 18 epochs on one NVIDIA A100 (bf16, batch 1024, lr 3e-3), about 36 minutes.
- Catalog metadata: each action carries
risk(safe/caution/critical),needs_sudoandplatforms, so the executor can decide what to run, confirm or refuse.
Safety
- The model does not refuse anything. The executor is the policy: whitelisted command templates, per-shell quoting, a risk gate and clarification questions.
- Credential guard: if a parsed Wi-Fi password or network name looks like a slip (a filler word such as "hai" or "pehle", the password equal to the SSID, the password inside the SSID), the executor asks again instead of connecting.
- Bare pin numbers: "pin no 11" can mean physical pin 11 or GPIO 11. For any action that drives a pin, the executor asks which one unless the user said gpio, bcm, board, physical or header. Power and ground header pins are always refused.
- Critical actions (shutdown, GPIO writes,
run_command, ...) need the action name typed as confirmation.
Limitations
- Training data is synthetic, with no real ASR transcripts. Test with your own speech recognizer, and expect lower accuracy on noisy transcripts.
- English and Roman-script Hinglish dominate the training data. Devanagari Hindi is only 0.26% of it (3,232 rows).
Spot checks work (
क्रोम खोलो→open_app chrome,वाईफाई बंद करो→wifi_off,5 मिनट का टाइमर लगाओ→set_timer), but there is no Devanagari test set, so its accuracy is not measured. - One command per utterance. It is a parser, not a multi-step planner.
- Small talk and out-of-scope requests usually come back as
unknown, but not always (seekya haal hai bhaiabove). Use the confidence score. - Known confusions:
scan_lanvsdiscover_devices(a labelling ambiguity), and bare GPIO pin numbers (handled by the executor, see Safety). - The sample executor is not a finished product. Some Windows actions have no command-line equivalent and only open
the right Settings page. Many Raspberry Pi actions need tools such as
pinctrl,nmcliori2c-tools.
Files
| Path | What it is |
|---|---|
model/model.npz, model/config.json |
The model weights (NumPy) and architecture |
model/runtime.py, tokenizer.py, grammar.py, normalize.py, snap.py, catalog.py |
NumPy inference code |
model/catalog.json |
The 371 actions with slots, risk, sudo and platform metadata |
pytorch/model.pt, pytorch/model.py |
PyTorch checkpoint ({"state", "cfg"}) and model definition, for fine-tuning |
quickstart.py |
Command-line demo (parse + show plan, nothing executed) |
executor/ |
Sample executor: Windows and Raspberry Pi command templates, risk gate, credential and pin guards |
ui/, run.sh, run.bat |
Local test UI and terminal agent |
eval/ |
Smoke test, curated and Wi-Fi evaluation prompts, HTML report generator, benchmark script and results |
tools/install_model.py |
Swap in a re-trained model.npz (backs up the current one) |
Load the PyTorch checkpoint:
import torch; from pytorch.model import Seq2Seq
ck = torch.load("pytorch/model.pt", map_location="cpu"); c = ck["cfg"]
m = Seq2Seq(c["n_in"], c["n_out"], d=c["d"], h=c["h"], f=c["f"], n_enc=c["n_enc"], n_dec=c["n_dec"],
max_in=c["max_in"], max_out=c["max_out"])
m.load_state_dict(ck["state"]) # 6,441,472 parameters
License and credits
Code and model weights: Apache-2.0, Copyright (c) 2026 sraivante. Trained from scratch; no pretrained weights are used. The training dataset is synthetic and is not published. Provided as-is, without warranty. You are responsible for anything you let an executor do.