Instructions to use zimplex/harp-vla-train37-act with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use zimplex/harp-vla-train37-act with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
ACT fine-tuned on HARP VLA train37
This is the validated final LeRobot policy checkpoint from the joint 37-task
HARP VLA fine-tuning run. It predicts absolute Franka joint-position targets
plus gripper state (qpos_target_abs, action dimension 8).
Repository: zimplex/harp-vla-train37-act
Final step: 100,000
Final training job: 135212
Final audit job: 135213
The ACT allocation that reached step 80,000 hit its 36-hour wall-time. Training was resumed exactly from the complete 080000 model, optimizer, and RNG state, and the repository was published only after step 100,000 and the replacement final audit passed.
Dataset
| Field | Value |
|---|---|
| Public source | zhouqh/harp_vla_data |
| Dataset card license | MIT at the pinned revision |
| Pinned source revision | 46691c4098ae2ca46a131745d1fe386d63a3fc33 |
| Conversion | LeRobot v3, qpos_target_abs, 20 Hz, PyAV video backend |
| Robot / renderer / schema | Franka / NYX / vla_steps_v2_qpos_target |
| Coverage | 37 source tasks, 36 language instructions, 300 episodes/task |
| Size | 11,100 episodes; 4,373,192 frames |
| Camera payload | head RGB + right-wrist RGB; recorded depth groups are empty by contract |
| Dataset completion marker SHA-256 | f9d5121d9f7a210918e5d3dd394d915b00a413d1659538d01729e8f8f6b46db6 |
| Converted data tree SHA-256 | 65dfc626682a26bfb64478d8b9da9b008aafe2d5536c97b52a21b2c320e8c055 |
| Task manifest SHA-256 | b0f46008dfe238c08be3c31f24d86ef80e155e7c5b9b8e75b99c8d73812aa8c6 |
Training protocol
All five policies used the same optimizer-step and example budget as the earlier HR-Bench train19 sweep. Every policy consumed exactly 6,400,000 training examples (6,400,000 for this model), approximately 1.46 passes over the 4,373,192-frame dataset. This is not equal-epoch training.
| Field | Value |
|---|---|
| Model / policy type | act |
| Initialization | ACT initialized from scratch with a TorchVision ResNet-18 ImageNet-1K backbone |
| Pinned base revision or digest | f37072fd47e89c5e827621c5baffa7500819f7896bbacec160b1a16c560e07ec |
| Nodes × GPUs/node | 2 × 8 NVIDIA B200 |
| Batch per process | 4 |
| Global batch | 64 |
| Optimizer steps | 100,000 |
| Save cadence | every 5,000 steps |
| Total examples | 6,400,000 |
| Seed | 1000 |
| Image augmentation | disabled |
| Mixed precision flag | use_amp=false |
| Data workers | 4 per process; prefetch factor 4 |
| cuDNN deterministic | false |
| W&B | offline; project harp_vla_train37_gcp |
| Training budget policy | same_optimizer_and_example_budget_as_hr_v2_train19 |
The ACT policy was initialized from scratch apart from the pinned TorchVision ResNet-18 IMAGENET1K_V1 visual backbone.
Pinned software runtime
| Component | Version / revision |
|---|---|
| LeRobot source | 26ff40ddd784280efc133a8e5af1a76e5ac731c2 plus the recorded HARP entrypoint |
| Python | 3.12 |
| PyTorch / TorchVision | 2.8.0 / 0.23.0 |
| Diffusers / Datasets | 0.35.2 / 4.4.1 |
| PyAV | 15.1.0 |
Optimizer and scheduler
{
"optimizer": {
"betas": [
0.9,
0.999
],
"eps": 1e-08,
"grad_clip_norm": 10.0,
"lr": 1e-05,
"type": "adamw",
"weight_decay": 0.0001
},
"scheduler": null
}
Policy architecture
| Config field | Value |
|---|---|
n_obs_steps |
1 |
chunk_size |
100 |
n_action_steps |
100 |
vision_backbone |
resnet18 |
pretrained_backbone_weights |
ResNet18_Weights.IMAGENET1K_V1 |
dim_model |
512 |
n_heads |
8 |
dim_feedforward |
3200 |
n_encoder_layers |
4 |
n_decoder_layers |
1 |
use_vae |
true |
latent_dim |
32 |
n_vae_encoder_layers |
4 |
dropout |
0.1 |
kl_weight |
10.0 |
normalization_mapping |
{"ACTION": "MEAN_STD", "STATE": "MEAN_STD", "VISUAL": "MEAN_STD"} |
Inputs and outputs
| Feature | Direction | Type | Shape |
|---|---|---|---|
observation.image |
input | VISUAL | 3 × 256 × 256 |
observation.state |
input | STATE | 8 |
observation.wrist_image |
input | VISUAL | 3 × 256 × 256 |
action |
output | ACTION | 8 |
The exact, machine-readable training and policy configurations are included as
train_config.json and config.json. They are the files emitted by the final
checkpoint; their hashes were checked against the immutable completion marker.
Files and loading
The repository root contains the complete pretrained_model/ payload from the
LeRobot checkpoint (0.19 GiB), including policy
weights, processor state, train_config.json, and config.json.
checkpoint_manifest.json records every uploaded checkpoint file's byte size
and SHA-256; SHA256SUMS provides the same digests in standard text form.
training_provenance.json records the immutable run, dataset,
source, job, and audit contract.
Optimizer and RNG files from training_state/ are intentionally not published;
this public repository is an inference checkpoint, not a resume bundle.
from lerobot.policies.factory import make_policy_from_pretrained
policy = make_policy_from_pretrained("zimplex/harp-vla-train37-act", device="cuda")
policy.eval()
Use a LeRobot checkout compatible with the configuration included here.
Reproducibility and validation
| Field | Value |
|---|---|
| Immutable run ID | harp_vla_train37_full_deadline_20261001T151243Z_2747a40 |
| Original protocol source commit | 2747a401bd29a070452aeab815d99847d7b507bb |
| Final attempt source commit | da1a0456fe16bab22bdf3759b6577b335c683680 |
| HARP race-safe entrypoint SHA-256 | bf6ac26d9d4487ffe3258fd482a4c3561d5effaab9fc501442367a159631e88e |
| Pinned LeRobot trainer SHA-256 | 5ae8b0b8f3312f14f054364b56d21f39e3843d650162443286e7342bc489d8c8 |
| Training environment marker SHA-256 | dd775f736fa809107dc69b653d5c0eef4679e6958fc9b8950fe6aaa3835c2bab |
| Completion marker SHA-256 | 623c7df20eaccee389645658c8012e54d507a458e7ea47fabbb7cff0c6fa19d2 |
train_config.json SHA-256 |
1672b4f12bf697fbb055d5097d824f52d5589ad77c236f4e1784fbe11382cfef |
config.json SHA-256 |
870f7913e685bb62ddd52124c1620a4c66d5ac9be5e6fd2658a872c9a78fecc7 |
The final audit required Slurm COMPLETED/0:0, exact marker/config hashes, the
pinned dataset and topology, and finite values in every final Safetensors
tensor. This model passed with zero audit errors.
Limitations
- This release contains training artifacts, not HARP benchmark evaluation scores. No downstream success rate is claimed by this model card.
- The checkpoint is specific to the converted 20 Hz absolute-qpos convention and its feature/normalization schema.
- No license is asserted by this card; users must comply with the licenses and terms of the base model/backbone, dataset, LeRobot, and other dependencies.
- Downloads last month
- 18