Biologyarticle2026-08-02

Homology aware benchmarking of deep learning models for CAZy enzyme family classification

Open access0 citations

Abstract

The CAZy database organizes carbohydrate-active enzymes into more than 800 families with distinct biochemical roles, making automatic family assignment from sequence a high-value task in metagenomics and enzyme engineering. Although recent deep learning studies have reported strong classification performance, we identify a critical evaluation bias that has not been systematically quantified in this setting: commonly used random train/test splits allow homologous sequences to appear in both training and evaluation sets, thereby inflating apparent performance by 5.9–12.0%. We refer to this effect as homology leakage and quantify its magnitude across five representative classification approaches. To enable more rigorous evaluation, we construct a homology-aware benchmark comprising 30,000 sequences, 60 families, and 6 classes, split using MMseqs2 clustering at 20% sequence identity. Under this protocol, the method most strongly affected by leakage is homology-based inference itself ( \(\Delta \) F1 \(= +0.076\) ), whose performance would be reported as 0.894 under a random split but drops to 0.818 under homology-aware evaluation. Against this benchmark, our proposed model, ESM-2+Mean+MTL (two-stage selective fine-tuning with joint family/class supervision), achieves a family Macro-F1 of 0.805 ± 0.004, narrowing the gap to homology-based inference while providing calibrated uncertainty estimates (ECE \(= 0.028\) ) and compatibility with 8 GB consumer GPU hardware. These findings support homology-aware evaluation as a more reliable standard for protein sequence classification.

// Source

View paper (DOI)Open access versionOpenAlexDiscover Artificial IntelligencePublished 2026-08-02

Authors: Ahmet Haşim Yurttakal, Hasan Erbay

Institutions: Afyon Kocatepe University, Ostim Technical University