Structure prediction tools now provide protein-ligand geometry for pairs no laboratory experiment has resolved thus far. This data is typically provided without calibrated, practically usable per-system measures of whether a given complex is correct and how much to trust it. Agreement between independently constructed methods is an established signal for this type of problem. What has been missing is a way to derive what a given level of agreement is worth. We present PLI-Parallax, which deposits predicted geometry together with experimental data needed to calibrate it. Chai-1, two Boltz-2 configurations, and the docking engine smina were run over shared inputs across an experimentally resolved crystal tier of 19,350 complexes and a corpus tier of 31,746 cross-docked pairs without experimental ground truth on predicted receptors, yielding 307,314,646 residue-to-ligand-atom distance records, which the stored coordinates allow a consumer to recompute at a cutoff of their own choosing. Their mutual agreement is fitted against observed accuracy where experimental data permits it and carried to where it does not. This way 30,567 systems carry predicted label accuracy together with a split-conformal interval. On the crystal tier that interval covers observed accuracy at the stated rate. On the corpus tier it ranks systems by expected label quality, since both distributions differ. The deposit is accompanied by 646 evaluation configurations across seven data split families, and 631 of them report how far their training and test entities separate under a two-sample test. The 906 protein accessions were partitioned to control sequence leakage, so a protein-cold split here tests generalisation across sequence space. The two tiers support evaluation against experimental data and, where this is absent, supervision weighted by how far the configurations agree.
Klamt, T.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0
