Data-centric training enables meaningful interaction learning in protein-ligand binding affinity prediction
Abstract
Predicting protein–ligand binding affinity remains a fundamental challenge in drug discovery, where deep learning models have become central. However, these models often struggle with protein and ligand biases, and their generalization to unseen targets remains limited. Here, we introduce a convolutional neural network (CNN)-based scoring function trained with principled data-centric strategies as a controlled case study to investigate model bias, diagnostic evidence of interaction learning, and the comparative rigor of out-of-distribution evaluation protocols. We propose molecular dropout, a novel regularization technique that, combined with random rotation augmentation, mitigates rotational variance and reduces bias toward isolated protein or ligand features. Together, these strategies promote the learning of meaningful interaction patterns using only experimental structures and measured affinities, without requiring artificial data such as cross-docked poses, decoy generation, or assumed binding labels. In-depth analyses show that the model achieves low false-negative rates for strong binders, supporting its applicability in virtual screening. On the PDBbind v2016 core set, the model achieves a Pearson correlation (R P ) of 0.839, and under the stringent Pfam-CV out-of-distribution protocol, it reaches R P = 0.573, performing competitively with existing methods on both benchmarks. Among the out-of-distribution protocols evaluated, Pfam-CV is the more rigorous benchmark. Analysis of interaction fingerprint similarity between training and test sets suggests that performance degradation under more stringent splitting schemes is largely driven by insufficient coverage of diverse binding landscapes in the available training data, rather than by an intrinsic inability of the model to capture meaningful intermolecular interaction patterns. Scientific contribution This work demonstrates that a straightforward CNN architecture, trained with principled data-centric strategies on experimental structures alone, can achieve competitive performance across diverse benchmarks without requiring complex preprocessing pipelines, synthetic data generation, or elaborate architectural designs. Ablation studies and in-depth analyses provide converging evidence that the model captures genuine protein–ligand interactions rather than merely exploiting molecular biases or dataset artifacts. Furthermore, analysis of interaction fingerprint similarity across splitting schemes suggests that the generalization gap observed under stringent out-of-distribution benchmarks reflects insufficient diversity of binding profiles in the available training data. This finding points to data coverage, rather than model capacity, as the main bottleneck to generalization.
// Source
Authors: Matheus Silva, Lincon Vidal, Isabella Alvim Guedes, Camila de Magalhães, Fábio Lima Custódio, Laurent E. Dardenne
Institutions: Universidade Federal do Rio de Janeiro, Laboratório Nacional de Computação Científica