A study of structure-based AI systems found that adding extra training goals can help with new peptides, while important prediction challenges persist.
T-cell receptors recognize short protein fragments, called peptides, when those fragments are displayed by immune proteins known as MHC. Computational tools aim to predict these interactions, but the study found that graph-based AI systems often perform poorly when asked about peptides absent from their training data.
The researchers examined whether the interaction features at the binding site and uncertainty in predicted molecular structures affected this problem. They report that adding extra training objectives to the classifier improved its ability to generalize to new peptides, while also showing that important challenges remain.
Findings on unfamiliar peptides
The researchers assessed graph neural network classifiers that predict whether a T-cell receptor will bind to a peptide displayed by a major histocompatibility complex, or MHC. They focused on systems using computationally predicted molecular structures.
They found that these classifiers generally had poor accuracy on samples containing peptides not seen during training. Their analysis examined two possible influences: how sensitive the classifiers were to features of the T-cell receptor–peptide–MHC contact site, and how uncertainty in the predicted structures affected performance.
Based on this analysis, the researchers designed classifiers with auxiliary training objectives—additional tasks used during training—and report that this improved generalization to novel peptides. The study also describes strengths and weaknesses among current graph-based approaches to this problem.