Instructions to use trl-lab/qwen3.5-2b-grpo-bird with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use trl-lab/qwen3.5-2b-grpo-bird with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="trl-lab/qwen3.5-2b-grpo-bird") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("trl-lab/qwen3.5-2b-grpo-bird") model = AutoModelForMultimodalLM.from_pretrained("trl-lab/qwen3.5-2b-grpo-bird", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use trl-lab/qwen3.5-2b-grpo-bird with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "trl-lab/qwen3.5-2b-grpo-bird" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "trl-lab/qwen3.5-2b-grpo-bird", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/trl-lab/qwen3.5-2b-grpo-bird
- SGLang
How to use trl-lab/qwen3.5-2b-grpo-bird with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "trl-lab/qwen3.5-2b-grpo-bird" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "trl-lab/qwen3.5-2b-grpo-bird", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "trl-lab/qwen3.5-2b-grpo-bird" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "trl-lab/qwen3.5-2b-grpo-bird", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use trl-lab/qwen3.5-2b-grpo-bird with Docker Model Runner:
docker model run hf.co/trl-lab/qwen3.5-2b-grpo-bird
MBIRD: a 2B text-to-SQL agent trained on BIRD
Project page · MSQaLe · All three models · Citation
MBIRD is one of the three text-to-SQL models from the paper SQaLe: a large realistic dataset to empower small specialised text-to-SQL models. It is Qwen3.5-2B trained with GRPO on the BIRD training split, using the same recipe as MSQaLe, the paper's model trained on SQaLe. The paper uses it to measure what this training achieves on a widely used benchmark corpus. Like its siblings, it works as an agent. It explores the database through tools, tests its query and then submits it, and the schema is never part of its prompt.
Quickstart
The model is an agent, so it needs a loop that executes its tool calls against a database and feeds the results back. sqale_agent.py in this repository is a self-contained reference loop that reproduces the training environment: the same prompt, tools, turn budget and final-answer fallback. It opens your database read-only.
pip install vllm huggingface_hub
hf download trl-lab/qwen3.5-2b-grpo-bird sqale_agent.py --local-dir .
python sqale_agent.py --model trl-lab/qwen3.5-2b-grpo-bird \
--db path/to/database.sqlite \
--question "Which supplier shipped the most back-ordered items last quarter?" \
--trace
From Python, one agent can answer many questions with a single loaded engine:
from sqale_agent import SQaLeAgent, vllm_generate
agent = SQaLeAgent(*vllm_generate("trl-lab/qwen3.5-2b-grpo-bird"))
episode = agent.run("How many customers placed more than three orders in 2024?", "shop.sqlite")
print(episode["sql"]) # the submitted query
print(episode["ended_by"]) # submit_sql, final_turn, fallback_to_last_working_query, ...
To build your own loop, follow the protocol in The agent loop below.
Highlights
- 54.7% on BIRD dev, the highest of the three models and +35.4 points over the untrained base model (19.3%). BIRD dev is in-distribution for this model.
- 54.0% on the SQaLe test set (+15.3 points over the base model) and 23.7% on EHRSQL (+15.5 points), level with MSQaLe.
- Reinforcement learning only. It is trained from the base checkpoint with no supervised warm start, no reasoning traces distilled from a larger teacher, and no test-time scaffolding at evaluation.
Accuracy on 100 moderate questions from the SQaLe test set against model size, with the average compute per question as colour. MSQaLe, MBIRD and MSynSQL are Qwen3.5-2B trained with the same recipe on SQaLe, BIRD train and SynSQL-2.5M; the red arrow is the gain of MSQaLe over the untrained Qwen3.5-2B. Figure from the paper's appendix.
The SQaLe model family
The paper trains Qwen3.5-2B three times with one recipe, once per training corpus, to measure how much of a small model's ability comes from the data it is trained on. Base model, agentic environment, tool set, reward, GRPO objective, step count, batch size and evaluation protocol are identical across the three. Only the training corpus and the strength of the length curriculum differ, so the gaps between the models come from the data.
| Model | Training corpus | Length curriculum |
|---|---|---|
| MSQaLe | SQaLe | full, four length buckets |
| MBIRD (this model) | BIRD train | none |
| MSynSQL | SynSQL-2.5M | milder, three length buckets |
Evaluation
Execution accuracy (%) on 300 questions per benchmark, as reported in the paper. Every model runs in the agentic environment used for training, with the schema withheld and no test-time scaffolding. The SQaLe test split shares no schema with SQaLe's training split and is evaluated over the full schema. EHRSQL is a single-domain benchmark over electronic health records.
| Model | SQaLe test | BIRD dev | EHRSQL |
|---|---|---|---|
| MSQaLe | 66.3 | 52.3 | 23.7 |
| MBIRD (this model) | 54.0 | 54.7 | 23.7 |
| MSynSQL | 50.7 | 44.3 | 13.3 |
| Qwen3.5-2B (untrained) | 38.7 | 19.3 | 8.2 |
| Qwen3.6-27B (untrained) | 76.0 | 69.3 | 55.0 |
Compared with MSQaLe
On the SQaLe test set, MSQaLe leads MBIRD by 12.3 points at full schema size. The paper's analyses trace much of this gap to two behaviours that BIRD's small databases do not train.
Searching large schemas
The SQaLe test questions are evaluated at three database sizes: the gold tables only, the gold tables plus 32 distractor tables, and the full schema. Brackets give the share of episodes in which the model opened every gold table.
| Tables in the database | MSQaLe | MBIRD | MSynSQL |
|---|---|---|---|
| Gold tables only | 72.7 (95.3) | 64.7 (98.7) | 59.3 (98.3) |
| + 32 distractor tables | 68.0 (96.7) | 61.0 (90.7) | 52.3 (92.7) |
| Full schema | 66.3 (89.7) | 54.0 (85.0) | 50.7 (83.3) |
BIRD databases hold 5–10 tables. With only the gold tables present, MBIRD finds all of them in 98.7% of episodes. At full SQaLe schema size it opens every gold table in 85.0% of episodes, against 89.7% for MSQaLe. A gap of 8.0 points remains even when only the gold tables are present, which the paper relates to the wider domain coverage of SQaLe's schemas.
Looking up values
Many questions depend on literals that exist only in the table values. Before submitting, MBIRD has looked up 40% of the string literals the gold answer depends on, against 88% for MSQaLe. When the question states the literal, both models carry it into their SQL equally often (91.5% and 91.7%). When it does not, MBIRD searches the data for 27% of them and uses the correct value in 41% of cases, against 61% and 55% for MSQaLe.
The agent loop
What an episode looks like
Question Which supplier shipped the most back-ordered items last quarter?
Turn 1 {"tool": "list_tables"}
→ suppliers, shipments, order_items, …
Turn 2 {"tool": "join_path", "args": {"table_a": "suppliers", "table_b": "order_items"}}
→ suppliers → shipments → order_items
⋮
Turn 7 {"tool": "submit_sql", "args": {"sql": "SELECT s.name …"}}
Each turn opens with a <think> block in which the model plans its next call, and a turn may hold several calls.
Using your own harness
The following details from training matter for accuracy if you build your own loop.
- Use the training prompt verbatim.
SYSTEM_PROMPTandUSER_TEMPLATEinsqale_agent.pyare the exact strings used in training. The system prompt holds the task, the tool list and an execution-plan scaffold. The user turn states that the schema is not shown and gives the question. - Tool calls are bare JSON. After its
<think>block the model writes one or more{"tool": "<name>", "args": {...}}objects with no prose and no code fences. Keep thinking enabled. - Return results as the next user turn, formatted as
Tool result:followed by the result and the remaining budget, for example[4 tool turns left before you must answer]. - Keep earlier reasoning in context. During training every earlier
<think>block stayed in the context. The Qwen chat template drops those blocks when a conversation is rebuilt from a message list, so append each completion to the prompt string instead, assqale_agent.pydoes. - Six tool turns, then a final turn. After the sixth tool turn, or after four failed
run_querycalls in a row, send the exploration-limit message and take the next answer as final. If that answer is empty or does not execute, use the last query that ran without error. - Sampling. Use temperature 0.6, top-p 0.95 and top-k 20, which are the defaults in
generation_config.json, and stop on<|im_end|>and<|endoftext|>. An episode is budgeted at 12,288 tokens, so a context of 16,384 tokens leaves headroom.
| Tool | Returns |
|---|---|
list_tables |
every table name |
describe_table(table_name) |
the table's CREATE TABLE statement |
foreign_keys(table_name) |
declared foreign keys as child.col -> parent.col; table_name is optional |
join_path(table_a, table_b) |
the shortest chain of foreign-key joins between two tables |
sample_rows(table_name, limit) |
a few rows of a table |
distinct_values(table_name, column_name, limit) |
the distinct values of a column |
run_query(sql) |
the first rows of a read-only query |
submit_sql(sql) |
records the final answer and ends the episode |
Training
Environment. Each training question becomes a tool-use episode over a live SQLite database, and the schema is never placed in the prompt. A graph-aware sampler decides which tables the episode's database holds. It draws a target size between 4 and 64 tables, starts from the gold tables and grows along foreign keys until the target is reached, then shuffles the table order. BIRD databases are small, so the sampler often returns the full schema. Early in training, a submit gate refuses submit_sql until the same query has been run with run_query. The gate's probability decays from 1.0 to 0 over the first 80 steps, and it is never active at evaluation.
Reward. A query whose result set matches the gold result scores 3.0. A query that executes with a wrong result scores 1.0 + 0.9·F1 + b, where F1 is the multiset F1 against the gold result and b adds 0.1 times the share of joins along declared foreign keys and 0.1 for agreeing with the most common result in the group. A query that parses but does not execute scores 0.5, and anything else scores 0. The executable tier tops out at 2.1, so no partial credit ranks a wrong query above a correct one. Correctness is decided by comparing result sets, with no LLM judge.
No length curriculum. BIRD's gold queries span a narrower range of lengths than SQaLe's, so MBIRD is trained with uniform sampling. This is the only training setting besides the corpus in which it differs from MSQaLe.
Optimisation. GRPO from the Qwen3.5-2B base checkpoint, with 18 questions × 8 candidates = 144 episodes per update for 1,800 steps. This repository holds the last checkpoint. BIRD dev was monitored during training and not used to select the checkpoint.
All hyperparameters (identical for the three models except the curriculum)
| Hyperparameter | Value |
|---|---|
| Base model / initial policy | Qwen3.5-2B (base checkpoint) |
| Reference policy | base model, re-anchored every 111 updates |
| Training corpus | BIRD train |
| Schema in prompt | no (withheld; tools only) |
| Protocol | execution-plan scaffold, 7 tools + submit_sql |
| Tool turns / final turn | 6 / 1 |
| Submit gate probability | 1.0 → 0.0 over the first 80 steps |
| Consecutive query-error cap | 4 |
| Tokens per tool turn / final turn / episode | 2,048 / 3,072 / 12,288 |
| Questions per step B / candidates per question G | 18 / 8 |
| Rollout temperature | 1.0, no top-p / top-k truncation |
| Rollout backend | asynchronous in-process vLLM |
| Optimiser | AdamW (8-bit where the engine shares a card) |
| Learning rate | 1 × 10-5, cosine schedule, 5% linear warm-up |
| Rollout steps | 1,800, one optimiser update per step |
| Surrogate | token-level ratio, PPO clip |
| Clip ε (low / high) | 0.2 / 0.28 (DAPO clip-higher) |
| Importance-ratio cap | 2.0 |
| Loss aggregation | token sum / 1024 |
| Advantage normalisation | none; clipped to ±2.0 |
| Negative-advantage weight | 0.7 |
| KL estimator / target / β init / β max | k3 / 0.005 / 0.02 / 1.0 |
| Dynamic sampling | groups with identical rewards resampled, up to 2 rounds (DAPO) |
| Reward | 3.0 equivalent · 1.0 + 0.9·F1 executable · 0.5 parsable · 0 invalid |
| Reward bonuses | 0.1 × FK-supported join share · 0.1 group consensus |
| Column-subset credit | enabled, discount 0.9 |
| Length curriculum | none (uniform sampling) |
| Monitoring | BIRD dev, 300 samples, every 50 steps, schema withheld |
| Checkpoint | last (step 1,800); no selection on BIRD dev |
Intended use and limitations
- MBIRD answers natural-language questions over SQLite databases through tool access. It is a research model released with the paper, mainly as the BIRD-trained reference for MSQaLe.
- Its BIRD dev score is in-distribution. On databases with many tables it finds the relevant tables less reliably than MSQaLe, which is the better choice of the three for large schemas.
- It expects the agent loop described above. Putting the DDL in a single prompt and asking for SQL is a different setting from the one it was trained in.
- It was trained and evaluated on SQLite only.
- At 2B parameters it still trails much larger models. The untrained Qwen3.6-27B scores 76.0 / 69.3 / 55.0 on the three benchmarks above, and EHRSQL's clinical questions remain hard for every 2B model in the paper.
- The model's SQL runs against your data, so give it read-only access, as
sqale_agent.pydoes. - The checkpoint is a full
Qwen3_5ForConditionalGenerationexport and carries Qwen3.5-2B's vision tower unchanged. Only the language model was trained, and only text was evaluated.
Citation
If you use this model, please cite the SQaLe paper:
@misc{wolff2026sqale,
title = {{SQaLe}: A Large Realistic Dataset to Empower Small Specialised Text-to-{SQL} Models},
author = {Wolff, Cornelius and Gomm, Daniel and Hulsebos, Madelon},
year = {2026}
}
Authors: Cornelius Wolff and Daniel Gomm (University of Amsterdam, Centrum Wiskunde & Informatica), Madelon Hulsebos (Centrum Wiskunde & Informatica). Questions and feedback are welcome in the Community tab of this repository.
The model builds on Qwen3.5-2B and is trained on BIRD. The evaluation also uses EHRSQL and the SQaLe test set, whose schemas come from SchemaPile. We thank their authors for making them available.
- Downloads last month
- 1,198