Analysis of structural errors in the AlphaFold DB v4 and v6
Abstract
3D protein structures predicted using machine-learning techniques have become routinely used in research. One of the major sources of the predicted structures is the AlphaFold DB, providing predictions for over 240 million proteins. For a proper interpretation of predicted protein structures, it is crucial to consider the prediction confidence. Regions with the highest confidence in predicted structures are expected to be modelled with high accuracy and should be suitable for any application that benefits from high accuracy (e.g. characterizing binding sites). Recently (October 2025), a new version of predictions in the AlphaFold DB was released (v6). In this paper, we analyse the number of structural errors in the newest version of predictions (v6) for Swiss-Prot sequences and compare the results with the last major prediction version (v4). In the analysis, we understand structural error as any inconsistency between the set of expected bonds according to the Chemical Component Dictionary (CCD) and the set of bonds actually estimated by the RDKit tool. Thus, we checked both for clashes (cases where atoms which should not be bonded are too close and so a bond between them is estimated) and for broken bonds (cases where atoms which should be bonded are too distant and so no bond between them is estimated). The results show a notable decrease in the number of structural errors, especially in the highest-confidence regions. Within v4, we detected 17 498 structural errors, including 6 385 in the highest-confidence regions. Meanwhile, we found 3 932 in v6, of which 259 were in the regions with the highest confidence. However, the decrease in the number of structural errors in v6 originates in the complete absence of two specific error patterns affecting histidine. In fact, this pair of patterns constituted the majority of structural errors in v4 predictions. At the same time, the number of other structural errors remains essentially the same. Despite the remarkable decline in the number of structural errors, the highest-confidence regions still contain some. Thus, before conducting accuracy-sensitive research, an inspection for structural errors is necessary. In addition, we show that most of the structural errors affecting residue side chains in both prediction versions can be corrected by removing the erroneous atoms and applying the missing-atom reconstruction technique implemented in PDB2PQR, commonly used in the post-processing of experimentally determined protein structures. The best success rates in correcting side-chain structural errors were attained in the highest-confidence regions, amounting to 99.64% for v4 and 87.29% for v6. The overall success rate across all confidence regions for v4 and v6 was 94.90% and 69.01%, respectively. We show that predicted protein structures in the AlphaFold DB contain structural errors, notably even in the highest-confidence regions. While the two histidine-specific error patterns constituting the majority of structural errors in v4 predictions have suddenly disappeared in v6 predictions, the number of other structural errors has persisted essentially the same. We also demonstrate that most of the side-chain structural errors can be resolved by deleting the erroneous atoms and reconstructing them using the PDB2PQR tool.
// Source
Authors: Lukáš Bohuš, Tomáš Svoboda, Ondřej Schindler
Institutions: Masaryk University, Central European Institute of Technology, Central European Institute of Technology – Masaryk University