tiny-superfast-agentic-MLM-6.5m-v1: a 6.4M-parameter voice-command parser for laptops and Raspberry Pi

tiny-superfast-agentic-MLM-6.5m-v1 turns one short spoken or typed command into a structured command that a program can execute. It understands English and Hinglish (Roman-script Hindi-English), with some Devanagari Hindi, and covers 371 actions (audio, display, Wi-Fi, apps, files, timers, power, network, services, Raspberry Pi GPIO/I2C and more), and runs on a plain CPU with NumPy only: no GPU, no PyTorch and no internet needed at inference.

ENGLISH   "set the volume to 40"                              ->  set_volume {"value": 40}
HINGLISH  "volume 40 kar do"                                  ->  set_volume {"value": 40}
ENGLISH   "connect to wifi Redmi Note 12 password hello@123"  ->  connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"}
HINGLISH  "pwd Tiger@2025 daalo aur Hostel Wing C se jud jao" ->  connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"}
HINGLISH  "10 min baad shutdown kar dena"                     ->  shutdown {"amount": 10, "unit": "min"}
HINGLISH  "board pin 11 pe led jalao"                         ->  gpio_on {"pin": 11, "numbering": "board"}
HINDI     "5 मिनट का टाइमर लगाओ"                              ->  set_timer {"amount": 5, "unit": "min"}
Parameters 6,441,472 (≈6.4M; "6.5m" in the name is rounded up)
Model file model/model.npz, 23.9 MB, float32. pytorch/model.pt is included for fine-tuning
Runs on CPU only, fully offline: Windows, Linux, Raspberry Pi 5. Inference needs only numpy
Accuracy 97.53% exact on 6,611 held-out-template commands (test_strict); 99.10% on validation; 100% on 461 unseen Wi-Fi phrasings
Speed median 18 ms per command on a laptop i7-1360P, 30 ms on a Raspberry Pi 5 (10 min sustained load)
Load time 0.13 s (laptop), 0.20 s (Raspberry Pi 5)
Bigger sibling tiny-agentic-home-robotic-for-edge-device-v10 (23.9M params, MiniLM-based)

About the name. "MLM" here means micro language model. It is not a masked language model like BERT: the architecture is a small Transformer encoder-decoder (seq2seq) that writes the command one token at a time.

Executor notice. The model only parses commands; it never runs anything. The executor in this repository is a sample for testing, not a finished product. It gates every action by risk (safe / caution / critical) and asks for confirmation before anything risky. Use it on a test machine and read the command before you run it.

Best use: a personal voice companion for small commands

The model is built to be the "understanding" step of a small, private, always-available assistant: you speak a short command, and your computer or Raspberry Pi does it. It replaces a large LLM for common, well-defined device commands, so the answer is instant (tens of milliseconds), private (nothing leaves the device) and free.

Attach any speech recognizer (ASR) in front of it and an executor behind it:

flowchart LR
    U(["User speaks<br/>'volume 40 kar do'"]) --> A["ASR<br/>speech → text<br/>(Whisper, Vosk, ...)"]
    A -->|"text"| M["<b>tiny-superfast-agentic-MLM-6.5m-v1</b><br/>text → structured command<br/>6.4M params · ~20–70 ms on CPU"]
    M -->|"set_volume {value: 40}<br/>+ confidence"| E["Executor<br/>validate → risk gate → OS command"]
    E --> O(["Output action<br/>volume set to 40%"])
    E -. "clarify / unknown /<br/>low confidence" .-> U

If your viewer does not render Mermaid:

user voice ==> [ ASR ] --text--> [ tiny-superfast-agentic-MLM-6.5m-v1 ] --command JSON--> [ executor ] ==> output action
                                                                                             |
                                          ask the user again  <----  clarify / unknown / low confidence
Stage What it does Example
1. ASR (you bring it) Turns speech into text. Any engine works: Whisper / faster-whisper, Vosk, a phone keyboard, or a chat box. audio → volume 40 kar do
2. This model Turns text into one action from the 371-action catalog, plus its typed arguments and a confidence score. → {"action": "set_volume", "value": 40}, confidence 0.997
3. Executor (sample included) Checks the arguments, applies the risk gate, renders a whitelisted OS command for Windows or Raspberry Pi, and runs it or asks first. → Pi: wpctl set-volume @DEFAULT_AUDIO_SINK@ 40%
4. Output action The device acts, and the result can be read back to the user (TTS). volume is now 40%

A good agent pattern: execute only when confidence is high and the risk gate allows it, ask the user when the result is clarify, and hand unknown or low-confidence text to a bigger model or back to the user.

Minimal voice loop (example)

# pip install numpy faster-whisper   (ASR is your choice; this is one option)
import os; os.environ.setdefault("OPENBLAS_NUM_THREADS", "4")      # "2" on a Raspberry Pi 5, see Speed
import sys; sys.path[:0] = ["model", "executor"]
from runtime import CommandModel
import executor as EX
from faster_whisper import WhisperModel

asr = WhisperModel("small", device="cpu", compute_type="int8")
parser = CommandModel("model")

def handle(wav_path):
    text = " ".join(s.text for s in asr.transcribe(wav_path)[0]).strip()   # 1. speech -> text
    out = parser.predict(text); meta = out.pop("_meta")                     # 2. text -> command
    plan = EX.plan(out, None, text)                                          # 3. command -> gated OS command
    if plan.action == "clarify" or meta["confidence"] < 0.8:
        return f"Sorry, can you say that again? ({plan.note or text})"
    ok, why = EX.gate(plan, allow_caution=False)                             # safe actions only in this demo
    return EX.run(plan, out, dry_run=not ok)                                 # 4. act (dry run if gated)

The ASR part of this snippet is an illustration. The model was trained on typed/synthetic text, not on real ASR transcripts, so test it with your own ASR output (see Limitations).

What it understands

Every example is shown in ENGLISH, then the same command in HINGLISH. The outputs, confidences and executor plans are real outputs of this model and the included sample executor.

Use case Language Command Model output Confidence Risk Raspberry Pi command
Audio ENGLISH set the volume to 40 set_volume {"value": 40} 0.996 safe wpctl set-volume … 40%
Audio HINGLISH volume 40 kar do set_volume {"value": 40} 0.997 safe wpctl set-volume … 40%
Wi-Fi ENGLISH connect to wifi Redmi Note 12 password hello@123 connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"} 1.000 caution nmcli dev wifi connect 'Redmi Note 12' password hello@123
Wi-Fi HINGLISH Redmi Note 12 wifi se connect karo password hello@123 same 1.000 caution same
Wi-Fi, password first ENGLISH password Tiger@2025 then join Hostel Wing C connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"} 1.000 caution nmcli dev wifi connect 'Hostel Wing C' …
Wi-Fi, password first HINGLISH pwd Tiger@2025 daalo aur Hostel Wing C se jud jao same 1.000 caution same
Apps ENGLISH open chrome open_app {"app": "chrome"} 0.990 caution chromium-browser &
Apps HINGLISH chrome kholo open_app {"app": "chrome"} 0.986 caution chromium-browser &
Timers ENGLISH set a timer for 5 minutes set_timer {"amount": 5, "unit": "min"} 0.998 safe built-in timer
Timers HINGLISH 5 minute ka timer lagao set_timer {"amount": 5, "unit": "min"} 0.998 safe built-in timer
Power ENGLISH shut down in 10 minutes shutdown {"amount": 10, "unit": "min"} 0.998 critical sudo shutdown -h +10
Power HINGLISH 10 min baad shutdown kar dena shutdown {"amount": 10, "unit": "min"} 0.998 critical sudo shutdown -h +10
Negation ENGLISH don't shut down the laptop cancel_shutdown {} 0.995 caution shutdown -c
Negation HINGLISH shutdown mat karo cancel_shutdown {} 0.995 caution shutdown -c
Pi GPIO ENGLISH turn on board pin 11 gpio_on {"pin": 11, "numbering": "board"} 0.999 critical sudo pinctrl set 17 op dh
Pi GPIO HINGLISH board pin 11 pe led jalao gpio_on {"pin": 11, "numbering": "board"} 0.998 critical sudo pinctrl set 17 op dh
Pi GPIO, ambiguous pin HINGLISH pin no 11 high karo gpio_on {"pin": 11} 0.993 → clarify executor asks: board pin 11 (= GPIO 17) or GPIO 11 (= board pin 23)?
Diagnostics ENGLISH what is the cpu temperature get_temperature {} 0.995 safe vcgencmd measure_temp
Diagnostics HINGLISH cpu ka temperature batao get_temperature {} 0.996 safe vcgencmd measure_temp
Files HINGLISH reports naam ka folder banao create_folder {"path": "reports"} 0.999 caution mkdir -p -- reports
Out of scope ENGLISH tell me a joke unknown {} 0.993 safe nothing
Out of scope (miss) HINGLISH kya haal hai bhai recent_files {} 0.604 safe (low confidence: ask or hand off)

The last row is a real miss: small talk went to a harmless action, but with low confidence. This is why the agent should check the confidence before acting.

Test results

Accuracy

Held-out evaluation after training (raw model output, before the runtime "snap" repair). "Exact" means the action and every argument match; "Action" means only the action matches.

Set n Exact Action What it tests
validation 10,413 99.10% 99.29% same templates as training, unseen rows
test 10,618 98.37% 98.52% held-out templates
test_strict 6,611 97.53% 97.70% held-out templates with no paraphrase sibling in training (the honest generalisation number)
curated 371 99.46% 99.73% one held-out prompt per action
B_typo 3,000 99.17% 99.50% typos
wifi_holdout 461 100.00% 100.00% unseen Wi-Fi connect phrasings
gap11_holdout 696 99.28% 100.00% unseen templates: password-first Wi-Fi, CPU clock vs usage, board↔BCM pins, fan vs mouse speed
A_challenge 1,152 97.40% 98.44% external challenge set
B_golden_final 384 97.40% 97.40% external golden set, including out-of-scope prompts
B_golden_dev 383 95.04% 95.82% external golden set
A_practical-dev 1,321 95.31% 96.44% external practical set
A_practical-holdout 428 92.06% 94.16% external practical set; some misses are label-convention differences

The same model was also run on the laptop and on the Raspberry Pi 5 (2,340 prompts, with the runtime snap). Both devices produce identical predictions: 97.82% exact, 98.46% action.

Speed and temperature

10 minutes of continuous single-stream inference on each device. Full per-pass and per-minute tables are in eval/BENCHMARK.md.

Laptop (Intel i7-1360P) Raspberry Pi 5 (16 GB)
Latency p50 / p95 / p99 18.4 / 168.6 / 239.6 ms 30.2 / 243.9 / 281.6 ms
Throughput, cool → after 10 min 29.6 → 18.5 commands/s 13.9 → 13.9 commands/s
Temperature idle → peak 52.6 → 91.9 °C 38.5 → 59.8 °C
Throttling slows down ~39% once hot none (2400 MHz throughout, get_throttled=0x0)
Best BLAS threads 4 2
  • Latency grows with the length of the command the model writes: short commands take ~10–35 ms, Wi-Fi commands with an SSID and a password take ~100–250 ms.
  • Set the BLAS thread count. The matrices are small, so NumPy's default (one thread per core) is up to 4x slower. Use OPENBLAS_NUM_THREADS=4 on x86 laptops and OPENBLAS_NUM_THREADS=2 on the Raspberry Pi 5 (quickstart.py does this for you).

Test devices

Laptop Raspberry Pi 5
CPU Intel Core i7-1360P (4P + 8E cores) Broadcom BCM2712, 4× Cortex-A76 @ 2.4 GHz
RAM 32 GB 16 GB
OS Windows 11 Raspberry Pi OS (Debian 13), kernel 6.12
Python / NumPy 3.13 / 2.5.3 3.13 / 2.2.4
Cooling built-in fan, AC power active cooler (PWM fan)

Download and test

1. Get the files

pip install -U huggingface_hub
hf download sraivante/tiny-superfast-agentic-MLM-6.5m-v1 --local-dir tiny-agentic-6.5m
cd tiny-agentic-6.5m

2. Parse commands (nothing is executed)

pip install numpy
python quickstart.py "volume 40 kar do" "connect to wifi Redmi Note 12 password hello@123"
python quickstart.py            # interactive

Each line prints the command, the confidence, the latency, the risk level and the OS command the sample executor would run on this machine.

3. Test UI and checks (Python 3.9+)

Windows Raspberry Pi / Linux
Test UI in the browser (http://localhost:8000) run.bat ./run.sh
Terminal agent run.bat cli ./run.sh cli
Smoke test (model + executor, nothing executed) run.bat test / run.bat test --full ./run.sh test / ./run.sh test --full
HTML report of all 371 actions run.bat report ./run.sh report

The first run creates .venv and installs numpy, fastapi and uvicorn. report really runs the safe, read-only actions on your machine (for example reading the volume or the IP address) and dry-runs everything else.

In the UI, Run all is always a dry run. Execute on a row runs that one command: safe actions run directly, caution actions need the "allow caution" box, and critical actions ask you to type the action name. The UI listens on 127.0.0.1 only and has no login, so do not expose it to a network.

4. Reproduce the speed test

cd eval
sh bench_run_all.sh python ../model <dir with eval .jsonl sets> <out_dir> default 1 2 4

How the model works

  • Input: one command, up to 128 characters, normalised and encoded character by character (200-symbol vocabulary), so typos, code-mixing and odd spellings are handled without a word tokenizer.
  • Network: Transformer encoder-decoder, d_model 256, 8 heads, FFN 768, 5 encoder + 3 decoder layers (pre-LN).
  • Output: a short canonical command such as connect_wifi ssid="…" password="…", written token by token (1,052-token output vocabulary).
  • Grammar-constrained decoding: at every step only tokens that keep the output valid are allowed. The action must come from the 371-action catalog, the slots must belong to that action, and text slots (SSIDs, passwords, file names, apps) must be copied from the input. The model cannot invent an action or a password that you did not say.
  • Snap: a small post-decode repair extends copied SSIDs and passwords to whole tokens and recovers the original letter case from the input.
  • Training: from random initialisation (no pretrained base), 1,230,162 synthetic English/Hinglish rows, 18 epochs on one NVIDIA A100 (bf16, batch 1024, lr 3e-3), about 36 minutes.
  • Catalog metadata: each action carries risk (safe/caution/critical), needs_sudo and platforms, so the executor can decide what to run, confirm or refuse.

Safety

  • The model does not refuse anything. The executor is the policy: whitelisted command templates, per-shell quoting, a risk gate and clarification questions.
  • Credential guard: if a parsed Wi-Fi password or network name looks like a slip (a filler word such as "hai" or "pehle", the password equal to the SSID, the password inside the SSID), the executor asks again instead of connecting.
  • Bare pin numbers: "pin no 11" can mean physical pin 11 or GPIO 11. For any action that drives a pin, the executor asks which one unless the user said gpio, bcm, board, physical or header. Power and ground header pins are always refused.
  • Critical actions (shutdown, GPIO writes, run_command, ...) need the action name typed as confirmation.

Limitations

  • Training data is synthetic, with no real ASR transcripts. Test with your own speech recognizer, and expect lower accuracy on noisy transcripts.
  • English and Roman-script Hinglish dominate the training data. Devanagari Hindi is only 0.26% of it (3,232 rows). Spot checks work (क्रोम खोलो → open_app chrome, वाईफाई बंद करो → wifi_off, 5 मिनट का टाइमर लगाओ → set_timer), but there is no Devanagari test set, so its accuracy is not measured.
  • One command per utterance. It is a parser, not a multi-step planner.
  • Small talk and out-of-scope requests usually come back as unknown, but not always (see kya haal hai bhai above). Use the confidence score.
  • Known confusions: scan_lan vs discover_devices (a labelling ambiguity), and bare GPIO pin numbers (handled by the executor, see Safety).
  • The sample executor is not a finished product. Some Windows actions have no command-line equivalent and only open the right Settings page. Many Raspberry Pi actions need tools such as pinctrl, nmcli or i2c-tools.

Files

Path What it is
model/model.npz, model/config.json The model weights (NumPy) and architecture
model/runtime.py, tokenizer.py, grammar.py, normalize.py, snap.py, catalog.py NumPy inference code
model/catalog.json The 371 actions with slots, risk, sudo and platform metadata
pytorch/model.pt, pytorch/model.py PyTorch checkpoint ({"state", "cfg"}) and model definition, for fine-tuning
quickstart.py Command-line demo (parse + show plan, nothing executed)
executor/ Sample executor: Windows and Raspberry Pi command templates, risk gate, credential and pin guards
ui/, run.sh, run.bat Local test UI and terminal agent
eval/ Smoke test, curated and Wi-Fi evaluation prompts, HTML report generator, benchmark script and results
tools/install_model.py Swap in a re-trained model.npz (backs up the current one)

Load the PyTorch checkpoint:

import torch; from pytorch.model import Seq2Seq
ck = torch.load("pytorch/model.pt", map_location="cpu"); c = ck["cfg"]
m = Seq2Seq(c["n_in"], c["n_out"], d=c["d"], h=c["h"], f=c["f"], n_enc=c["n_enc"], n_dec=c["n_dec"],
            max_in=c["max_in"], max_out=c["max_out"])
m.load_state_dict(ck["state"])   # 6,441,472 parameters

License and credits

Code and model weights: Apache-2.0, Copyright (c) 2026 sraivante. Trained from scratch; no pretrained weights are used. The training dataset is synthetic and is not published. Provided as-is, without warranty. You are responsible for anything you let an executor do.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support