ACT fine-tuned on HARP VLA train37

This is the validated final LeRobot policy checkpoint from the joint 37-task HARP VLA fine-tuning run. It predicts absolute Franka joint-position targets plus gripper state (qpos_target_abs, action dimension 8).

Repository: zimplex/harp-vla-train37-act
Final step: 100,000
Final training job: 135212
Final audit job: 135213

The ACT allocation that reached step 80,000 hit its 36-hour wall-time. Training was resumed exactly from the complete 080000 model, optimizer, and RNG state, and the repository was published only after step 100,000 and the replacement final audit passed.

Dataset

Field Value
Public source zhouqh/harp_vla_data
Dataset card license MIT at the pinned revision
Pinned source revision 46691c4098ae2ca46a131745d1fe386d63a3fc33
Conversion LeRobot v3, qpos_target_abs, 20 Hz, PyAV video backend
Robot / renderer / schema Franka / NYX / vla_steps_v2_qpos_target
Coverage 37 source tasks, 36 language instructions, 300 episodes/task
Size 11,100 episodes; 4,373,192 frames
Camera payload head RGB + right-wrist RGB; recorded depth groups are empty by contract
Dataset completion marker SHA-256 f9d5121d9f7a210918e5d3dd394d915b00a413d1659538d01729e8f8f6b46db6
Converted data tree SHA-256 65dfc626682a26bfb64478d8b9da9b008aafe2d5536c97b52a21b2c320e8c055
Task manifest SHA-256 b0f46008dfe238c08be3c31f24d86ef80e155e7c5b9b8e75b99c8d73812aa8c6

Training protocol

All five policies used the same optimizer-step and example budget as the earlier HR-Bench train19 sweep. Every policy consumed exactly 6,400,000 training examples (6,400,000 for this model), approximately 1.46 passes over the 4,373,192-frame dataset. This is not equal-epoch training.

Field Value
Model / policy type act
Initialization ACT initialized from scratch with a TorchVision ResNet-18 ImageNet-1K backbone
Pinned base revision or digest f37072fd47e89c5e827621c5baffa7500819f7896bbacec160b1a16c560e07ec
Nodes × GPUs/node 2 × 8 NVIDIA B200
Batch per process 4
Global batch 64
Optimizer steps 100,000
Save cadence every 5,000 steps
Total examples 6,400,000
Seed 1000
Image augmentation disabled
Mixed precision flag use_amp=false
Data workers 4 per process; prefetch factor 4
cuDNN deterministic false
W&B offline; project harp_vla_train37_gcp
Training budget policy same_optimizer_and_example_budget_as_hr_v2_train19

The ACT policy was initialized from scratch apart from the pinned TorchVision ResNet-18 IMAGENET1K_V1 visual backbone.

Pinned software runtime

Component Version / revision
LeRobot source 26ff40ddd784280efc133a8e5af1a76e5ac731c2 plus the recorded HARP entrypoint
Python 3.12
PyTorch / TorchVision 2.8.0 / 0.23.0
Diffusers / Datasets 0.35.2 / 4.4.1
PyAV 15.1.0

Optimizer and scheduler

{
  "optimizer": {
    "betas": [
      0.9,
      0.999
    ],
    "eps": 1e-08,
    "grad_clip_norm": 10.0,
    "lr": 1e-05,
    "type": "adamw",
    "weight_decay": 0.0001
  },
  "scheduler": null
}

Policy architecture

Config field Value
n_obs_steps 1
chunk_size 100
n_action_steps 100
vision_backbone resnet18
pretrained_backbone_weights ResNet18_Weights.IMAGENET1K_V1
dim_model 512
n_heads 8
dim_feedforward 3200
n_encoder_layers 4
n_decoder_layers 1
use_vae true
latent_dim 32
n_vae_encoder_layers 4
dropout 0.1
kl_weight 10.0
normalization_mapping {"ACTION": "MEAN_STD", "STATE": "MEAN_STD", "VISUAL": "MEAN_STD"}

Inputs and outputs

Feature Direction Type Shape
observation.image input VISUAL 3 × 256 × 256
observation.state input STATE 8
observation.wrist_image input VISUAL 3 × 256 × 256
action output ACTION 8

The exact, machine-readable training and policy configurations are included as train_config.json and config.json. They are the files emitted by the final checkpoint; their hashes were checked against the immutable completion marker.

Files and loading

The repository root contains the complete pretrained_model/ payload from the LeRobot checkpoint (0.19 GiB), including policy weights, processor state, train_config.json, and config.json. checkpoint_manifest.json records every uploaded checkpoint file's byte size and SHA-256; SHA256SUMS provides the same digests in standard text form. training_provenance.json records the immutable run, dataset, source, job, and audit contract.

Optimizer and RNG files from training_state/ are intentionally not published; this public repository is an inference checkpoint, not a resume bundle.

from lerobot.policies.factory import make_policy_from_pretrained

policy = make_policy_from_pretrained("zimplex/harp-vla-train37-act", device="cuda")
policy.eval()

Use a LeRobot checkout compatible with the configuration included here.

Reproducibility and validation

Field Value
Immutable run ID harp_vla_train37_full_deadline_20261001T151243Z_2747a40
Original protocol source commit 2747a401bd29a070452aeab815d99847d7b507bb
Final attempt source commit da1a0456fe16bab22bdf3759b6577b335c683680
HARP race-safe entrypoint SHA-256 bf6ac26d9d4487ffe3258fd482a4c3561d5effaab9fc501442367a159631e88e
Pinned LeRobot trainer SHA-256 5ae8b0b8f3312f14f054364b56d21f39e3843d650162443286e7342bc489d8c8
Training environment marker SHA-256 dd775f736fa809107dc69b653d5c0eef4679e6958fc9b8950fe6aaa3835c2bab
Completion marker SHA-256 623c7df20eaccee389645658c8012e54d507a458e7ea47fabbb7cff0c6fa19d2
train_config.json SHA-256 1672b4f12bf697fbb055d5097d824f52d5589ad77c236f4e1784fbe11382cfef
config.json SHA-256 870f7913e685bb62ddd52124c1620a4c66d5ac9be5e6fd2658a872c9a78fecc7

The final audit required Slurm COMPLETED/0:0, exact marker/config hashes, the pinned dataset and topology, and finite values in every final Safetensors tensor. This model passed with zero audit errors.

Limitations

  • This release contains training artifacts, not HARP benchmark evaluation scores. No downstream success rate is claimed by this model card.
  • The checkpoint is specific to the converted 20 Hz absolute-qpos convention and its feature/normalization schema.
  • No license is asserted by this card; users must comply with the licenses and terms of the base model/backbone, dataset, LeRobot, and other dependencies.
Downloads last month
18
Safetensors
Model size
51.7M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train zimplex/harp-vla-train37-act