Ο€0.5 β€” Piper cube-stack (behavior-cloning baseline)

A Ο€0.5 (openpi) vision-language-action policy fine-tuned to "pick up the red cube and stack it on the blue cube" with a single-arm Piper, in the mjlab (MuJoCo) simulator.

This is the behavior-cloning baseline from a RECAP (Ο€*0.6) study: Ο€0.5 fine-tuned from pi05_base on 200 scripted demonstrations (128 success / 72 fail). Geometric stack-success on eval β‰ˆ 48% (25–50 episodes, 700-step horizon). It clones the demos cleanly β€” picks, transports, stacks β€” but inherits their failure modes. Full visual writeup of the pipeline: https://claude.ai/code/artifact/7482c197-98a9-4bfd-9415-d6ed16310829

Contents

  • params/ β€” Ο€0.5 inference weights (orbax checkpoint).
  • assets/local/piper_stack_act_v4_lr21/norm_stats.json β€” state/action normalization (bundled; required).
  • _CHECKPOINT_METADATA.

Inference-only: the optimizer state (train_state/) is intentionally omitted, so this is ~12 GB instead of ~42 GB. Self-contained the same way openpi's own pi05_base checkpoint is.

Observation / action spec

  • Inputs: three RGB cameras @ 224Γ—224 β€” scene_top, wrist_left, wrist_right β€” plus a 7-DoF joint state, at 50 Hz.
  • Output: 7-DoF action chunks, action_horizon = 10.
  • Prompt: pick up the red cube and stack it on the blue cube

Run it locally

You need two repos:

  1. Serving (openpi): axiboai/piperx-openpi, branch sagar_pi06 β€” provides the pi05_piper_stack config and the Piper input/output transforms.
  2. Simulator / client (mjlab): vla-mjlab β€” provides the piper_stack task and the piper_stack_act observation layout.

1. Download the model

huggingface-cli download axiboai/pi05_piper_stack_bc --local-dir ./pi05_piper_stack_bc

2. Serve the policy (openpi env, on a GPU)

# in axiboai/piperx-openpi (branch sagar_pi06)
uv run python scripts/serve_policy.py policy:checkpoint \
  --policy.config pi05_piper_stack \
  --policy.dir ./pi05_piper_stack_bc
# -> websocket policy server on :8000

3. Drive the mjlab sim against it (vla-mjlab env)

# in vla-mjlab; serve + sim can share a host, or split over an SSH tunnel on :8000
python -m scripts.collect_recap_rollouts \
  --repo-id local/try_bc --root /tmp/try_bc --overwrite \
  --policy-backend openpi --policy-host localhost --policy-port 8000 \
  --obs-layout piper_stack_act \
  --task "pick up the red cube and stack it on the blue cube" \
  --num-episodes 25 --num-envs 1 --fps 50 --legacy-cube-spawn --max-steps 700 \
  --action-horizon 10 --chunk-consume 10 --device cuda
# prints success_rate; saves rollout videos under --root

Notes

  • Baseline only (~48%). Imitation learning; no RL. It's the reference point a RECAP-fine-tuned policy is compared against.
  • Training data: axiboai/piper_stack_act_v4_lr21 (200 eps, LeRobot v2.1).
  • Success = red cube ends within xy ≀ 2 cm and z ≀ 1.2 cm of the blue cube top, held stable.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading