Recent developments in the field of evolutionary covariance and machine learning have enabled the precise prediction of residue-residue contacts and increasingly accurate inter-residue distance predictions. Access to this accurate covariance information has played a pivotal role in the latest advances observed in the field of protein bioinformatics, particularly the improvement of prediction of protein folds by ab initio protein modelling.
X-ray crystallographic experiments typically entail building a model that satisfies the experimental observations. However, experimental limitations can lead to unavoidable uncertainties during model building resulting in regions that require validation and potentially further refinement. Many metrics are available for model validation, but most are limited to the consideration of the physico-chemical aspects of the model or its match to the map.
Here we present new validation metrics based on the availability of accurate inter-residue distance predictions, which are compared with the distances observed in the emerging model. These metrics were fed into a support vector machine classifier that was trained to detect model errors based on historical data from the EM Validation Challenges. Further analysis of the possible register errors is done by performing an alignment of the predicted contact map and the map inferred from the contacts observed in the model. Regions of the model where the maximum contact overlap is achieved through a sequence register different to that observed in the model are flagged and the optimal sequence register can then be used to fix the register error.
Results suggest that both the detection of model errors and the correction of sequence register errors is possible, even in challenging cases, through the use of the trained classifier in conjunction with the contact map alignment. The approach, implemented in IRIS and ConKit, thus provides a new tool for protein structure validation that is orthogonal to existing methods.