A comparison of human protein variants shows that each approach can miss important effects for different reasons.
Researchers compared missense variant data from multiplexed laboratory tests with predictions from five widely used computational tools. They found that disagreements were not random, but reflected the different biological signals and assumptions used by the two approaches.
The computer tools relied heavily on evolutionary conservation and basic structural features. Laboratory tests could capture more context-specific effects, but their results were limited when the test did not represent the biology of a disease or produced substantial experimental noise.
Where the methods differ
The study analysed missense variant data from 37 human proteins and compared the results with five state-of-the-art computational variant effect predictors. The researchers found that the tools tended to overestimate harmful effects at buried and water-repelling parts of proteins, while underestimating effects in disordered regions and at charged sites on a protein’s surface.
The laboratory assays more accurately captured some context-specific mechanisms. However, they could miss harmful variants when the assay did not reflect the biology involved in a disease, and their results could be affected by high experimental noise. Protein features, assay design and variant type all influenced whether the two approaches agreed.
Why the disagreement matters
Genetic variant interpretation often combines computer predictions with laboratory evidence. The study provides a framework for understanding why those sources may conflict, rather than treating disagreement as random or assuming that one approach is always more reliable.
The findings support combining the approaches with attention to the underlying protein mechanism, the biological context represented by an assay and the kinds of variants being assessed.
Evidence and caveats
This was a comparative analysis of missense variant measurements from multiplexed assays covering 37 human proteins and predictions from five computational tools. It identifies systematic patterns of agreement and disagreement across these data, including examples involving clinically relevant variants.
The abstract does not report the total number of variants or describe the detailed design of each assay. The laboratory results can be noisy or fail to model disease biology, while the computational predictions depend strongly on sequence conservation and basic structural information. These limitations mean that neither method captures every disease-relevant effect in every context.
// Source
Nature Communications · 2026 · DOI: 10.1038/s41467-026-77211-x
Authors: Benjamin Livesey, Joseph A. Marsh
Institutions: University of Edinburgh, Institute of Genetics and Cancer