An Explainable Machine Learning Framework for Hypothesis Generation in Biochemical Methane Potential Prediction
Abstract
Abstract Biochemical methane potential (BMP) assays are widely used to evaluate feedstocks for anaerobic digestion, but they are slow and resource-intensive. This study evaluated whether explainable machine learning can support transparent screening and hypothesis generation from compositional feedstock data. A Random Forest model was trained on 127 solid and semi-solid feedstocks from a public dataset (Dry Matter, DM ≥ 15%). For the primary BMP per dry matter setting, repeated nested five-fold cross-validation yielded R 2 = 0.530, RMSE = 44.5 Nm 3 CH 4 /t DM, and MAE = 33.6 Nm 3 CH 4 /t DM; target-sensitivity analyses gave lower performance when DM and VS were excluded from the predictor set or when BMP per VS was modeled directly, with R 2 values of 0.423 and 0.377, respectively. The evaluated models showed broadly similar, moderate performance with the available dataset and descriptor set. Lignin remained a leading predictor across Random Forest, Gradient Boosting, and Extra Trees, whereas the ordering of other leading descriptors was model-dependent. SHAP visualization and stratified analyses suggested a possible DM-stratified Lignin-BMP pattern, but the adjusted interaction was not significant and bootstrap uncertainty spanned zero; therefore, this pattern was treated as exploratory. Accordingly, the framework may support preliminary feedstock screening, prioritization of confirmatory experiments, and hypothesis generation within the observed compositional domain, but it is not intended for causal inference or field-ready engineering prediction.
// Source
Authors: Daniel Oluwagbotemi Fasheun, Folorunsho Bright Omage, Viridiana Santana Ferreira-Leitão
Institutions: Universidade Estadual de Campinas (UNICAMP), Universidade Federal do Rio de Janeiro, Instituto Nacional de Tecnologia