A novel clustering-regression machine learning framework for biomass classification and biochemical composition prediction from elemental composition.

Journal: Bioresource technology
Published Date:

Abstract

Biomass elemental and biochemical compositions determine its conversion behavior and utilization potential. However, a standardized classification system based on these intrinsic characteristics is lacking, and different biomass types follow distinct processing pathways. In comparison, biomass biochemical analysis is time-consuming and costly, limiting its large-scale evaluation and application. To address these challenges, this study developed a novel clustering-regression machine learning (ML) framework that predicts the main biochemical components of different biomass types based on their elemental composition. First, Principal Component Analysis (PCA) achieved effective dimensionality reduction, with a total explained variance of 93.2 %. Using a PCA-assisted clustering model, biomass was effectively clustered into lipid-/protein-rich and lignocellulosic groups with a high silhouette score of 0.605. Subsequently, regression model 1 predicted protein and lipid contents in lipid-/protein-rich biomass, achieving average test R2 of 0.88 and average validation R2 of 0.81. For lignocellulosic biomass, regression model 2 predicted fibres and lignin contents, achieving average test R2 of 0.77 and average validation R2 of 0.66. Clustering and regression models all underwent thorough validation and testing, showing strong generalization. To make the integrated clustering-regression framework more user-friendly, a software application is established. Users can input elemental composition to identify biomass clusters and predict the corresponding main component contents. This study provides a reliable predictive tool for screening suitable feedstocks for targeted product manufacturing.

Authors

Keywords

No keywords available for this article.