pi0.5 fine-tuned on HARP VLA train37

This is the validated final LeRobot policy checkpoint from the joint 37-task HARP VLA fine-tuning run. It predicts absolute Franka joint-position targets plus gripper state (qpos_target_abs, action dimension 8).

Repository: zimplex/harp-vla-train37-pi05
Final step: 100,000
Final training job: 131710
Final audit job: 131714

Dataset

Field Value
Public source zhouqh/harp_vla_data
Dataset card license MIT at the pinned revision
Pinned source revision 46691c4098ae2ca46a131745d1fe386d63a3fc33
Conversion LeRobot v3, qpos_target_abs, 20 Hz, PyAV video backend
Robot / renderer / schema Franka / NYX / vla_steps_v2_qpos_target
Coverage 37 source tasks, 36 language instructions, 300 episodes/task
Size 11,100 episodes; 4,373,192 frames
Camera payload head RGB + right-wrist RGB; recorded depth groups are empty by contract
Dataset completion marker SHA-256 f9d5121d9f7a210918e5d3dd394d915b00a413d1659538d01729e8f8f6b46db6
Converted data tree SHA-256 65dfc626682a26bfb64478d8b9da9b008aafe2d5536c97b52a21b2c320e8c055
Task manifest SHA-256 b0f46008dfe238c08be3c31f24d86ef80e155e7c5b9b8e75b99c8d73812aa8c6

Training protocol

All five policies used the same optimizer-step and example budget as the earlier HR-Bench train19 sweep. Every policy consumed exactly 6,400,000 training examples (6,400,000 for this model), approximately 1.46 passes over the 4,373,192-frame dataset. This is not equal-epoch training.

Field Value
Model / policy type pi05
Initialization lerobot/pi05_base
Pinned base revision or digest b211f3d44c36b6acfcf7ae94a64e8e96f75a64ba
Nodes × GPUs/node 4 × 8 NVIDIA B200
Batch per process 2
Global batch 64
Optimizer steps 100,000
Save cadence every 5,000 steps
Total examples 6,400,000
Seed 1000
Image augmentation disabled
Mixed precision flag use_amp=false
Data workers 4 per process; prefetch factor 4
cuDNN deterministic false
W&B offline; project harp_vla_train37_gcp
Training budget policy same_optimizer_and_example_budget_as_hr_v2_train19

The pinned lerobot/pi05_base compatibility view only renamed the disabled relative_actions_processor registry entry to delta_actions_processor; relative-action processing remained disabled.

Pinned software runtime

Component Version / revision
LeRobot source 26ff40ddd784280efc133a8e5af1a76e5ac731c2 plus the recorded HARP entrypoint
Python 3.12
PyTorch / TorchVision 2.8.0 / 0.23.0
Diffusers / Datasets 0.35.2 / 4.4.1
PyAV 15.1.0

Optimizer and scheduler

{
  "optimizer": {
    "betas": [
      0.9,
      0.95
    ],
    "eps": 1e-08,
    "grad_clip_norm": 1.0,
    "lr": 2.5e-05,
    "type": "adamw",
    "weight_decay": 0.01
  },
  "scheduler": {
    "decay_lr": 2.5e-06,
    "num_decay_steps": 30000,
    "num_warmup_steps": 1000,
    "peak_lr": 2.5e-05,
    "type": "cosine_decay_with_warmup"
  }
}

Policy architecture

Config field Value
n_obs_steps 1
paligemma_variant gemma_2b
action_expert_variant gemma_300m
dtype bfloat16
chunk_size 50
n_action_steps 50
max_state_dim 32
max_action_dim 32
num_inference_steps 10
tokenizer_max_length 200
use_relative_actions false
gradient_checkpointing false
freeze_vision_encoder false
train_expert_only false
normalization_mapping {"ACTION": "QUANTILES", "STATE": "QUANTILES", "VISUAL": "IDENTITY"}

Inputs and outputs

Feature Direction Type Shape
observation.images.base_0_rgb input VISUAL 3 × 224 × 224
observation.images.left_wrist_0_rgb input VISUAL 3 × 224 × 224
observation.images.right_wrist_0_rgb input VISUAL 3 × 224 × 224
observation.state input STATE 32
action output ACTION 8

The exact, machine-readable training and policy configurations are included as train_config.json and config.json. They are the files emitted by the final checkpoint; their hashes were checked against the immutable completion marker.

Files and loading

The repository root contains the complete pretrained_model/ payload from the LeRobot checkpoint (8.71 GiB), including policy weights, processor state, train_config.json, and config.json. checkpoint_manifest.json records every uploaded checkpoint file's byte size and SHA-256; SHA256SUMS provides the same digests in standard text form. training_provenance.json records the immutable run, dataset, source, job, and audit contract.

Optimizer and RNG files from training_state/ are intentionally not published; this public repository is an inference checkpoint, not a resume bundle.

from lerobot.policies.factory import make_policy_from_pretrained

policy = make_policy_from_pretrained("zimplex/harp-vla-train37-pi05", device="cuda")
policy.eval()

Use a LeRobot checkout compatible with the configuration included here.

Reproducibility and validation

Field Value
Immutable run ID harp_vla_train37_full_deadline_20261001T151243Z_2747a40
Original protocol source commit 2747a401bd29a070452aeab815d99847d7b507bb
Final attempt source commit da1a0456fe16bab22bdf3759b6577b335c683680
HARP race-safe entrypoint SHA-256 bf6ac26d9d4487ffe3258fd482a4c3561d5effaab9fc501442367a159631e88e
Pinned LeRobot trainer SHA-256 5ae8b0b8f3312f14f054364b56d21f39e3843d650162443286e7342bc489d8c8
Training environment marker SHA-256 dd775f736fa809107dc69b653d5c0eef4679e6958fc9b8950fe6aaa3835c2bab
Completion marker SHA-256 91a8787b74571ea89cd587815ac589b8032a1700b3ba9122486f519a36200246
train_config.json SHA-256 5058523ccedc9a86589905c24d713a2ddcb6d3da9acefce3fa71cbd1e27492c9
config.json SHA-256 32e0f9ad96881b50f951f7f0fe028a01aa1d2e0e1fc0c5f91ded783e9d9e656d

The final audit required Slurm COMPLETED/0:0, exact marker/config hashes, the pinned dataset and topology, and finite values in every final Safetensors tensor. This model passed with zero audit errors.

Limitations

  • This release contains training artifacts, not HARP benchmark evaluation scores. No downstream success rate is claimed by this model card.
  • The checkpoint is specific to the converted 20 Hz absolute-qpos convention and its feature/normalization schema.
  • No license is asserted by this card; users must comply with the licenses and terms of the base model/backbone, dataset, LeRobot, and other dependencies.
Downloads last month
12
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for zimplex/harp-vla-train37-pi05

Finetuned
(737)
this model

Dataset used to train zimplex/harp-vla-train37-pi05