--- title: MiniCPM5-1B-Agent emoji: ๐Ÿ› ๏ธ colorFrom: indigo colorTo: yellow sdk: docker app_port: 7860 pinned: false # Build Small Hackathon TRACK + MERIT-BADGE tags. Kebab-case best-guesses - confirm the exact strings against # the Field Guide (https://build-small-hackathon-field-guide.hf.space/) before final submit; the intent is right. tags: - backyard-ai # TRACK: local, self-hosted AI (swap if you pick the other track) - well-tuned # full fine-tune of MiniCPM5-1B, published on the Hub - llama-champion # served on the llama.cpp runtime - off-the-grid # runs fully local on a CPU, no cloud model APIs --- # ๐Ÿ› ๏ธ MiniCPM5-1B-Agent **A tiny agentic coding agent that runs the whole write โ†’ run โ†’ read โ†’ debug โ†’ verify loop on a free CPU.** ![MiniCPM5-1B-Agent demo](minicpm5-1b-agent-demo.gif) A full fine-tune of [`openbmb/MiniCPM5-1B`](https://huggingface.co/openbmb/MiniCPM5-1B) (1B params), served as a Q8_0 GGUF on llama.cpp, no GPU. Give it a task; it reasons in ``, then uses `bash` / `write` / `read` / `edit` in a sandbox to build, run, and fix code, and renders the result (charts, images, live HTML) inline in the chat. Multi-turn: files and history persist across messages. It is also exposed as an **MCP tool** (`run_coding_task` at `/gradio_api/mcp/`). ## What it is Most coding agents are 70B+ behind a cloud API. This is the opposite: a **1B** model doing the *real* agentic loop on a **2-vCPU CPU Space**, no GPU. It writes a file, runs it, reads the output, debugs, and shows you the artifact, the same loop a big agent runs, shrunk to something you could host in your own backyard. ## How it was built - **Data (`train_v4`, 45,762 rows):** the proven v2 backbone (retail-filtered teacher mixes + real-usage agent traces) kept whole, plus ~3,538 curated additions, gated to a small served tool vocab and solution-aware MinHash-deduped. Bundled on the model repo under `dataset/`. - **SFT:** full fine-tune (not LoRA, the long agentic mix needs the capacity) of the abliterated base, 1 epoch, 24k context, fits in ~15-18 GB VRAM (direct Liger fused cross-entropy + mem-efficient SDPA). - **DPO (on-policy):** run the SFT model over the training prompts and capture its OWN behaviour. *chosen* = a valid `` tool call; *rejected* = its real miss (rambling in `` / answering with no call). ~649 pairs. This rewards ACTING over stalling. - **Serving:** Q8_0 GGUF on llama.cpp; a two-phase decode bounds the `` separately from the action so the model acts instead of looping; produced files render inline (charts, images, sandboxed live-HTML iframes). ## Try it - "Write a Python script that makes a bar chart of 30, 45, 25 (A, B, C), save chart.png, then run it." โ†’ writes + runs it; the PNG renders inline. - "Write an HTML page quote.html with a button that shows a random quote each click (hard-coded, no internet)." โ†’ writes the file; renders live in a sandboxed iframe. It is a tiny 1B on a free CPU: expect **~4 min per simple turn**, longer for multi-step tasks (the demo video shows it working end-to-end, so it can be judged even if a live run is slow). ## Model, dataset & full reproduction โ†’ **[Luminia/MiniCPM5-1B-Agent-GGUF](https://huggingface.co/Luminia/MiniCPM5-1B-Agent-GGUF)** (model card = the full data mix, SFT/DPO recipe, eval, and exact reproduce commands; v4 dataset bundled under `dataset/`). ๐Ÿ’ป **Code on GitHub:** [Katehuuh/MiniCPM5-1B-Agent](https://github.com/Katehuuh/MiniCPM5-1B-Agent) (the Space + the full training pipeline; code reviewed with OpenAI Codex). *Built for the Build Small Hackathon ยท OpenBMB + OpenAI Codex tracks.*