Biologypreprint2026-08-30

Frozen Protein Language Model Representations Capture Functional-Effect Signals Across Sodium-Channel Variants

Open access0 citations

Abstract

Missense variants in voltage-gated sodium channel genes are associated with neurological dis-eases such as infantile epilepsy or Dravet syndrome, but most computational variant predictiontools focus on pathogenicity rather than the functional effect of the mutation. Pathogenic vari-ants can cause either gain-of-function (GoF) or loss-of-function (LoF) effects, making functionalprediction important for understanding how a mutation actually impacts channel activity. In thisstudy, I constructed a benchmark of 348 sodium channel variants labeled as GoF or LoF acrossmultiple SCN paralogs and evaluated whether frozen ESM-2 protein language model embeddingsencode functional direction. In the held-out SCN1A analysis, frozen WT+mutant ESM nearest-neighbor classification achieved its highest balanced accuracy at k = 3, with a balanced accuracyof 0.764 and ROC-AUC of 0.787, compared with 0.671 and 0.733, respectively, for the strongestSCION-derived random-forest baseline. In a separate global embedding-geometry analysis using all348 validated variants as queries, the nearest cross-gene ESM neighbor shared the same GoF/LoFfunctional direction 66.95% of the time, compared with a 48.92% random cross-gene expectation.A separate embedding-geometry analysis showed that nearest cross-gene ESM neighbors shared thesame GoF/LoF mechanism 66.95% of the time, above a 48.92% random cross-gene null expectation,suggesting that ESM neighborhoods are non-randomly enriched for shared functional mechanism.AlphaMissense pathogenicity scores did not distinguish SCN1A GoF from LoF variants, supportingthe distinction between predicting whether a variant is damaging and predicting how it changesprotein function. Supervised adaptation produced inconsistent improvements across held-out genes,suggesting that the strongest transferable signal was already present in frozen ESM representations.Overall, these results suggest that protein language model embeddings capture functional effect in-formation beyond generic pathogenicity and engineered variant descriptors, supporting their usefor more specific interpretation of sodium channel variants.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-30

Authors: Rohith Sriram