Plant biomass estimation is critical due to the variability of different environmental factors and crop management practices associated with it. The assessment is largely impacted by the accurate prediction of different environmental sustainability indicators. A robust model to predict sustainability indicators is a must for the biomass community. This study proposes a robust model for biomass sustainability prediction by analyzing sustainability indicators using machine learning models. The prospect of ensemble learning was also investigated to analyze the regression problem. All experiments were carried out on a crop residue data from the Ohio state. Ten machine learning models, namely, linear regression, ridge regression, multilayer perceptron, k-nearest neighbors, support vector machine, decision tree, gradient boosting, random forest, stacking and voting, were analyzed to estimate three biomass sustainability indicators, namely soil erosion factor, soil conditioning index, and organic matter factor. The performance of the model was assessed using cross-correlation (R2), root mean squared error and mean absolute error metrics. The results showed that Random Forest was the best performing model to assess sustainability indicators. The analyzed model can now serve as a guide for assessing sustainability indicators in real time.
翻译:植物生物量估算因环境因素和作物管理实践的变异性而至关重要。其评估在很大程度上受限于不同环境可持续性指标的准确预测。构建稳健的可持续性指标预测模型是生物质研究领域的必要任务。本研究通过运用机器学习模型分析可持续性指标,提出了一个用于生物质可持续性预测的稳健模型。同时探讨了集成学习方法在回归问题分析中的应用潜力。所有实验均基于俄亥俄州的作物残留数据展开。研究分析了十种机器学习模型——线性回归、岭回归、多层感知机、K近邻、支持向量机、决策树、梯度提升、随机森林、堆叠集成与投票集成——以估算三种生物质可持续性指标:土壤侵蚀因子、土壤调节指数及有机质因子。模型性能通过交叉相关系数(R²)、均方根误差和平均绝对误差进行评估。结果表明,随机森林在评估可持续性指标方面表现最优。该分析模型可为实时评估可持续性指标提供指导。