Optifold
research notes.

Current development results, their limits, and the next tests.

Download the evidence snapshot Source records checked September 5, 2026

Development results.

Named secondary model0.7584

Development AUROC

Weighted hydropathy + disorder-promoting fraction. 166 development proteins: 97 passing and 69 not passing the study gate.

Recorded estimator: exhaustive leave-pair-out over 6,693 discordant pairs. Frozen August 28; refit September 2 returned the same value.

Frozen primary feature0.6891

Development AUROC

Negative disorder-promoting fraction, evaluated on 81 previously used development proteins.

Recorded August 28. The operational rule predicts FOLDABLE when the fraction is ≤ 0.5641.

These are not a head-to-head improvement claim.

These development estimates use different populations and provisional structural labels. Neither is a confirmatory result on untouched prospective targets.

The secondary model was declared in advance of confirmatory holdout evaluation. Its apparent advantage over hydropathy alone was not statistically established in the recorded comparison. The original single-feature rule remains the primary analysis.

The combined formula.

Optifold combines mean hydropathy (H) and the disorder-promoting fraction (D), after scaling each with frozen development-set statistics.

S=0.67z(H)+0.33z(D)S = 0.67\,z(H) + 0.33\,z(-D)
Scaled hydropathy
z(H)=H+0.3440830.620907z(H)=\frac{H+0.344083}{0.620907}
Scaled −D
z(D)=0.509481D0.081716z(-D)=\frac{0.509481-D}{0.081716}

H is mean Kyte–Doolittle hydropathy over the sequence. D is the fraction of A, R, G, Q, S, P, E and K. It is a composition proxy, not a prediction of where disordered regions occur.

Round H and D to four decimals, apply the fixed scaling and weights, then round the score to six decimals. The scaling is never recomputed from the sequence being entered.

For 7SQG, H = −0.2009 and D = 0.5029 produce an Optifold score of 0.181081. Higher scores rank toward passing the tested folding gate; no combined-score cutoff or probability calibration has been established.

This is weighted_hydropathy_disorder_v1, frozen August 28 and unchanged by the September 2 refit. It remains the study’s named secondary model; D ≤ 0.5641 remains the original primary analysis.

Try the combined score

What AUROC tells us.

AUROC measures how well a score ranks passing targets above non-passing targets. A value of 0.5 represents chance-level ranking; 1.0 represents perfect ranking on the evaluated population.

An AUROC of 0.7584 does not mean 75.84% of recommendations are correct. Sensitivity, specificity, and the cost of a wrong decision depend on the threshold and the population. The practical question is how many useful folds are missed and how many unsuccessful attempts are avoided.

Chance ranking
Perfect ranking

0.5 0.7584, development estimate 1.0

Experimental conditions.

Each full-matrix target is tested at three MSA depths—32, 128, and 256—and four recycle counts—0, 1, 3, and 8. AlphaFold2 v2.3.2 uses template-free monomer model 3 with a fixed seed.

12 settings per full matrix
MSA depth0 recycles1 recycle3 recycles8 recycles
3232 / 032 / 132 / 332 / 8
128128 / 0128 / 1128 / 3128 / 8
256256 / 0256 / 1256 / 3256 / 8

No additional MSA sequences are supplied beyond the selected depth, so the depth setting controls the actual experiment. This is a reduced-MSA study; failing all twelve conditions does not establish failure under stock AlphaFold2.

The accuracy gate

A usable structure must meet all five preregistered checks against the experimental reference:

  • TM-score ≥ 0.50
  • C-alpha lDDT ≥ 0.60
  • C-alpha RMSD ≤ 5.0 Å
  • Aligned coverage ≥ 0.95
  • Sequence identity ≥ 0.99

AlphaFold’s own confidence, pLDDT, is kept outside this accuracy label. Primary analyses also require at least 30 observed C-alpha residues and at least 50% experimental coverage.

What remains open.

Prospective performance

The clean holdout has 118 frozen targets, but only 36 meet the full-depth feasibility requirement: at least 256 genuine deduplicated MSA rows, including the query. These are the deep-MSA subset, not a random subset of all 118. Any confirmatory result needs to state that narrower scope.

Publication-grade scoring

Chain review alone does not resolve scoring validation. The current local scoring path uses an in-house C-alpha lDDT implementation rather than the pinned publication toolchain.

Savings and regret

Wall-clock savings depend on the execution regime, including compilation. Savings must name their baseline and be reported with the failure rate of the recommended policy.

Public folding access

This site calculates sequence scores. It does not start GPU jobs or return new predicted structures.

Source notes.

The downloadable snapshot names the records and amendments behind each headline figure.

  • Combined model, fixed scaling and rounding: tools/optifold.py, SECONDARY and secondary_score; records/q1_feature_prespecification_v1.json, amendments 5 and 6.
  • Feature and threshold: q1_feature_prespecification_v1, amendment 2.
  • Named secondary and development result: amendment 5.
  • Unchanged refit and remaining scoring requirement: amendment 6.
  • Full-matrix feasibility of the holdout: amendments 7 and 8.
  • Experimental structure illustration: RCSB PDB 7SQG, chain A .
Calculate the Optifold score