AlphaFold2 provides highly accurate protein structure predictions (including “atomic accuracy” scale improvements) in CASP14 and shows that the method’s predicted confidence (pLDDT) meaningfully tracks measured local accuracy on a large set of newer PDB structures.
On the CASP14 dataset of 87 protein domains, the paper reports a median backbone accuracy of 0.96 Å r.m.s.d.95 (95% confidence interval 0.85–1.16 Å) for AlphaFold, compared with a 2.8 Å r.m.s.d.95 (95% confidence interval 2.7–4.0 Å) for the next-best method, measured on CASP domains (reported also with an all-atom metric summary).
Skeptical note: CASP is designed as a blind assessment, but the paper’s conclusion is still conditional on (i) the benchmark’s protein mix, (ii) how model selection and confidence were used for that evaluation, and (iii) the appropriateness of the chosen metrics as surrogates for downstream biological correctness (metrics quantify geometry; they do not directly measure function).
The paper extends evaluation beyond CASP by analyzing 10,795 protein chains from newer PDB structures (filtered to exclude structures too similar to the training set; see Methods filtering details). It reports that the measured local backbone accuracy lDDT-Cα is predicted by pLDDT with a fitted linear relationship: lDDT-Cα = 0.997 × pLDDT − 1.17 with Pearson r = 0.76, and a 95% confidence region from 10,000 bootstrap samples.
Interpretation vs inference: The paper treats this correlation as validation that pLDDT can support “confident use” of predictions. That is a reasonable inference for geometry-based correctness, but it does not guarantee correct ligand binding, allosteric states, or conformational ensembles—those require additional experimental or modeling modalities beyond static geometry scores.
The same newer PDB analysis reports a quantitative link between predicted pTM and the full-chain TM-score. It gives the linear fit: TM-score = 0.98 × pTM + 0.07 with Pearson r = 0.85 for the 10,795 chains, again with a bootstrap-derived 95% interval shading.
Known limitations (explicitly stated by authors): The paper reports that accuracy decreases substantially when median alignment depth drops below ~30 sequences, showing a threshold-like effect; it also reports weaker performance for proteins with few intra-chain (homotypic) contacts compared to heterotypic contacts (e.g., bridging domains in complexes), with homomer cases often better. These are important because they indicate that “sequence-only structure prediction” is partially “conditioned” on the availability and type of evolutionary/co-complex information encoded in MSAs.
What remains uncertain / how this could be falsified: The paper’s quantitative validations focus on static geometric agreement on benchmark-like targets. Strong falsification would require showing that predicted structures with high confidence (high pLDDT/pTM) systematically misrepresent experimentally observed conformational states, ligand-bound states, or interface geometries across diverse protein families—especially in regimes where MSAs are shallow or when the relevant geometry is determined primarily by heterotypic contacts not captured by homotypic reasoning.
The paper’s core claim is well-supported by (i) a large benchmark-style blind evaluation showing a large median accuracy gap with reported uncertainty, and (ii) a sizeable externalized test set where confidence predictions are quantitatively calibrated to geometry-based accuracy. The main scientific blind spot is not “whether coordinates can be accurate,” but “how far these static accuracy proxies transfer” to dynamics, ensembles, ligand effects, and context-dependent complexes—domains not resolved by the paper’s geometry-only evaluation framework.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.