Ο0.5 β Piper cube-stack (behavior-cloning baseline)
A Ο0.5 (openpi) vision-language-action policy fine-tuned to "pick up the red cube and stack it on the blue cube" with a single-arm Piper, in the mjlab (MuJoCo) simulator.
This is the behavior-cloning baseline from a RECAP (Ο*0.6) study: Ο0.5 fine-tuned from pi05_base
on 200 scripted demonstrations (128 success / 72 fail). Geometric stack-success on eval β 48%
(25β50 episodes, 700-step horizon). It clones the demos cleanly β picks, transports, stacks β but
inherits their failure modes. Full visual writeup of the pipeline:
https://claude.ai/code/artifact/7482c197-98a9-4bfd-9415-d6ed16310829
Contents
params/β Ο0.5 inference weights (orbax checkpoint).assets/local/piper_stack_act_v4_lr21/norm_stats.jsonβ state/action normalization (bundled; required)._CHECKPOINT_METADATA.
Inference-only: the optimizer state (train_state/) is intentionally omitted, so this is ~12 GB instead
of ~42 GB. Self-contained the same way openpi's own pi05_base checkpoint is.
Observation / action spec
- Inputs: three RGB cameras @ 224Γ224 β
scene_top,wrist_left,wrist_rightβ plus a 7-DoF joint state, at 50 Hz. - Output: 7-DoF action chunks,
action_horizon = 10. - Prompt:
pick up the red cube and stack it on the blue cube
Run it locally
You need two repos:
- Serving (openpi):
axiboai/piperx-openpi, branchsagar_pi06β provides thepi05_piper_stackconfig and the Piper input/output transforms. - Simulator / client (mjlab):
vla-mjlabβ provides thepiper_stacktask and thepiper_stack_actobservation layout.
1. Download the model
huggingface-cli download axiboai/pi05_piper_stack_bc --local-dir ./pi05_piper_stack_bc
2. Serve the policy (openpi env, on a GPU)
# in axiboai/piperx-openpi (branch sagar_pi06)
uv run python scripts/serve_policy.py policy:checkpoint \
--policy.config pi05_piper_stack \
--policy.dir ./pi05_piper_stack_bc
# -> websocket policy server on :8000
3. Drive the mjlab sim against it (vla-mjlab env)
# in vla-mjlab; serve + sim can share a host, or split over an SSH tunnel on :8000
python -m scripts.collect_recap_rollouts \
--repo-id local/try_bc --root /tmp/try_bc --overwrite \
--policy-backend openpi --policy-host localhost --policy-port 8000 \
--obs-layout piper_stack_act \
--task "pick up the red cube and stack it on the blue cube" \
--num-episodes 25 --num-envs 1 --fps 50 --legacy-cube-spawn --max-steps 700 \
--action-horizon 10 --chunk-consume 10 --device cuda
# prints success_rate; saves rollout videos under --root
Notes
- Baseline only (~48%). Imitation learning; no RL. It's the reference point a RECAP-fine-tuned policy is compared against.
- Training data:
axiboai/piper_stack_act_v4_lr21(200 eps, LeRobot v2.1). - Success = red cube ends within xy β€ 2 cm and z β€ 1.2 cm of the blue cube top, held stable.