MitoInteract / recovery /SOURCE_ATLAS_GATE.md
Ethan Troy
feat: add source-aware multi-task transfer gate
04bd3b9
|
Raw History Blame Contribute Delete
6.09 kB

Source Atlas Transfer Gate

Verdict

Do not fund source-aware neural pretraining from this configuration.

A fixed CPU-only multi-task model trained on strict exact Kd, Ki, IC50, and EC50 observations performed worse than an otherwise identical Kd-only model on every homology-grouped development fold. The benchmark test partition was not evaluated.

This result rejects the tested linear transfer mechanism and removes the current justification for a paid neural pretraining run. It does not prove that every possible multi-task architecture must fail.

Exact-only relational contract

The BindingDB 202607 source-aware envelope contained 87,315 records. Source Atlas admitted only records satisfying all of the following:

  • relation exactly =;
  • positive finite value in nM;
  • one protein chain and a nonempty normalized sequence;
  • canonical protein, ligand, and pair identities matching recomputation;
  • valid canonical SMILES;
  • successful assay join.
Task Strict exact records
Kd 2,432
Ki 21,339
IC50 47,829
EC50 3,731
Total 75,331

The relational contract contains 2,061 proteins, 40,477 ligands, and 64,055 canonical pairs. Only 4,041 pairs have more than one measurement type, and only two pairs have all four types. Measurement types remain explicit and were never relabeled or pooled into a generic target.

The 11,984 excluded records comprise 11,344 non-exact measurements and 640 records without a successful assay join. Censored records remain available in the original source-aware dataset but were not used in this gate.

Joint homology filtering

All 2,061 exact-task proteins were jointly clustered with MMseqs2 using:

minimum sequence identity: 0.5
coverage:                  0.8
coverage mode:             0
alignment mode:            3

This produced 1,123 clusters. Thirty-four clusters contain at least one of the 43 proteins in the benchmark cold-protein test partition. Every Atlas record and development row in those clusters was reserved and excluded.

This stricter filter revealed that the original exact-protein split contains 121 development proteins homologous at the 50/80 threshold to a test protein. Those proteins account for 292 development observations. Removing them left:

safe development observations: 1,765
safe development clusters:       187
safe Atlas records:            71,473

The original test labels were not used for fitting, model selection, or evaluation.

Fixed transfer experiment

The cheap gate intentionally used a fixed configuration rather than searching hyperparameters:

  • protein features: log sequence length and 20 amino-acid fractions;
  • ligand features: 512-bit radius-2 Morgan fingerprint and eight physical descriptors;
  • model: Ridge with alpha=100;
  • multi-task architecture: one shared feature block plus one task-specific block for each of Kd, Ki, IC50, and EC50;
  • task weights: inverse-frequency weights giving each task equal total weight;
  • evaluation: five outer folds grouped by joint MMseqs2 clusters;
  • bootstrap: 2,000 cluster-level replicates with seed 42.

For every outer fold, all source records from the held protein clusters were removed from training. The multi-task and Kd-only models used the same feature contract and fixed regularization.

Results

Model OOF RMSE OOF MAE R² Pearson
Kd-only Ridge 1.7202 1.3941 0.0235 0.2792
Exact multi-task Ridge 1.9369 1.5390 -0.2381 0.2537

The multi-task model increased RMSE by 0.2167. Expressed as the predeclared positive-improvement quantity, the result was:

Kd-only RMSE - multi-task RMSE: -0.2167
95% cluster-bootstrap interval: [-0.3441, -0.0893]
probability of any improvement:  0.0%

All five folds worsened. Three folds exceeded the allowed 0.10 catastrophic regression tolerance.

Decision

reject_neural_pretraining_not_justified

No GPU run was launched, no test prediction was generated, and no checkpoint was produced. Runtime was 86.8 seconds on CPU at zero cloud cost.

The generated 75,331-record relational dataset, MMseqs2 databases, and 1,765 prediction rows remain outside Git. Only source code, tests, hashes, and the checked reports are versioned.

What remains plausible

This result makes naive shared transfer from exact Ki, IC50, and EC50 a poor next investment. It does not rule out:

  1. genuinely independent exact-Kd data with compatible assay provenance;
  2. structure-aware protein-ligand features that add information absent from sequence composition and fingerprints;
  3. a future relation-aware censored objective, but only after a local control demonstrates benefit;
  4. uncertainty-aware prediction and abstention around the strong LightGBM control.

The next model experiment should begin with new information, not another parameterization of the same BindingDB-derived targets.

Reproduce

Generated data stays under ignored artifact/cache directories:

cd recovery
uv run python scripts/prepare_source_atlas.py \
  --source-records artifacts/bindingdb-202607-source-aware-final/source_records.jsonl \
  --output-dir /root/.cache/mitointeract/source-atlas-v1

mmseqs easy-cluster \
  /root/.cache/mitointeract/source-atlas-v1/proteins.fasta \
  /root/.cache/mitointeract/source-atlas-v1/mmseqs-identity-50-coverage-80/clusters \
  /root/.cache/mitointeract/source-atlas-v1/mmseqs-identity-50-coverage-80/tmp \
  --min-seq-id 0.5 -c 0.8 --cov-mode 0 --alignment-mode 3 --threads 8

uv run python scripts/run_source_atlas_gate.py \
  --atlas-dir /root/.cache/mitointeract/source-atlas-v1 \
  --clusters /root/.cache/mitointeract/source-atlas-v1/mmseqs-identity-50-coverage-80/clusters_cluster.tsv \
  --sample artifacts/bindingdb-202607-benchmark-v2/sample.jsonl \
  --manifest artifacts/bindingdb-202607-benchmark-v2/split-cold_protein_exact.jsonl \
  --output-dir /root/.cache/mitointeract/source-atlas-gate-v1 \
  --alpha 100 --outer-splits 5 --bootstrap-iterations 2000 --seed 42