3D LLM Diffusion fk_k8_lam4

This repository is the reproducible release for the fk_k8_lam4 raw LeMat-GenBench run reported in the project GitHub repository. It releases the actual generator checkpoint, the MACE formation-energy guidance model, the FK sampling configuration, all 2,500 generated CIFs, and the GenBench summary.

fk_k8_lam4 is a sampling run name, not a separate neural checkpoint. It uses v12_multiprop/best.pt with force-kernel (FK) steering: 8 particles, lambda 4.0, sigma range [0.03, 0.20], and 4 resampling checkpoints.

Released Artifacts

File or directory Purpose SHA256
best.pt v12_multiprop generator checkpoint used by FK 5995f3086844e8b1ff0d8d76b029f11024b1a319ce10ff186c3f4e003ecf8056
mace_eform_mp20_run-7.model Formation-energy guidance model used by FK bf6040a923e22b67add51557c69d0b7f2a3d0b1830280c2a1f914ad72b665001
artifacts/prompts.jsonl Fixed 2,500 evaluation prompts 41c7f660585da40e8ab60b5ffdaa619e470cd0edf5da0fb45764bca18fcd4f78
artifacts/text_z_bench.pt Fixed prompt embeddings 2b00a3f16cbe1166cc824391df28ae1183126a7853cc189816b22822683bc652
evaluation_cifs.tar.gz Archive containing all 2,500 raw generated CIFs evaluated by LeMat 96537fe54dc95248c27c9367d3cf5cc27e9d5e40fbf900327bf6d245931e8e87
artifacts/generation_manifest.json Exact sampling configuration See SHA256SUMS
artifacts/genbench_direct.json Formal GenBench result See SHA256SUMS

Direct LeMat-GenBench Result

The run uses 2,500 fixed prompts, 100 EDM steps, CFG scale 2.0, s_churn=2.0, and seed 11. CIFs were submitted without pre-relaxation and evaluated with the LeMat ORB, MACE, and UMA relaxation ensemble.

Valid Unique Novel Stable Metastable Mean E_hull (eV) Relaxation displacement RMSE (A)
86.08% 84.99% 79.18% 4.74% 29.88% 0.2981 1.1246

This direct raw-generator run passes 1 of the 7 project thresholds (novelty). Relaxation displacement RMSE is LeMat's per-index Cartesian displacement from the submitted structure to its MLIP-relaxed structure. It is not an RMSD to a reference crystal structure.

Reproduction Scope

reproducibility/run_fk_generation.sh regenerates the FK sampling protocol after an MP20-compatible training CSV is supplied. The current sampler derives empirical atom-count and allowed-element priors from that CSV. The original MP20 train/validation split and training text embeddings are not redistributed.

The full evaluated CIF pool is included in evaluation_cifs.tar.gz, so the published GenBench submission can be audited without rerunning generation. The source snapshot, environment lock, artifact hashes, and provenance are provided for the inference-side protocol. Retraining v12_multiprop is outside this release scope.

Source And License

The main implementation and project documentation are at Richardyangfan78/3D_LLM_Diffusion. The code and released artifacts are provided under the MIT license in this repository. Third-party code and dependencies retain their notices in NOTICE.

Limitations

This is a research candidate-generation run. The direct benchmark result is not DFT validation, does not establish synthesizability, and should not be confused with the separately curated MLIP-relaxed hybrid delivery result.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support