Cropland maps are a core and critical component of remote-sensing-based agricultural monitoring, providing dense and up-to-date information about agricultural development. Machine learning is an effective tool for large-scale agricultural mapping, but relies on geo-referenced ground-truth data for model training and testing, which can be scarce or time-consuming to obtain. In this study, we explore the usefulness of combining a global cropland dataset and a hand-labeled dataset to train machine learning models for generating a new cropland map for Nigeria in 2020 at 10 m resolution. We provide the models with pixel-wise time series input data from remote sensing sources such as Sentinel-1 and 2, ERA5 climate data, and DEM data, in addition to binary labels indicating cropland presence. We manually labeled 1827 evenly distributed pixels across Nigeria, splitting them into 50\% training, 25\% validation, and 25\% test sets used to fit the models and test our output map. We evaluate and compare the performance of single- and multi-headed Long Short-Term Memory (LSTM) neural network classifiers, a Random Forest classifier, and three existing 10 m resolution global land cover maps (Google's Dynamic World, ESRI's Land Cover, and ESA's WorldCover) on our proposed test set. Given the regional variations in cropland appearance, we additionally experimented with excluding or sub-setting the global crowd-sourced Geowiki cropland dataset, to empirically assess the trade-off between data quantity and data quality in terms of the similarity to the target data distribution of Nigeria. We find that the existing WorldCover map performs the best with an F1-score of 0.825 and accuracy of 0.870 on the test set, followed by a single-headed LSTM model trained with our hand-labeled training samples and the Geowiki data points in Nigeria, with a F1-score of 0.814 and accuracy of 0.842.
翻译:农田地图是遥感农业监测的核心关键组成部分,可提供关于农业发展的密集且最新信息。机器学习是大规模农业制图的有效工具,但依赖于用于模型训练和测试的地理参考实地数据,这类数据可能稀缺或耗时难以获取。本研究探索了结合全球农田数据集与人工标注数据集,训练机器学习模型生成尼日利亚2020年10米分辨率新农田地图的有效性。我们为模型提供来自遥感数据源(如Sentinel-1和2、ERA5气候数据和DEM数据)的逐像素时间序列输入数据,以及指示农田存在的二分类标签。我们在尼日利亚范围内人工标注了1827个均匀分布的像素,将其按50%训练集、25%验证集和25%测试集划分,用于拟合模型并验证输出地图。通过评估和比较单头与多头长短期记忆网络(LSTM)分类器、随机森林分类器,以及三张现有10米分辨率全球土地覆盖图(谷歌动态世界、ESRI土地覆盖、ESA世界覆盖)在我们提出的测试集上的性能,我们发现:考虑到农田外观的区域差异,我们进一步实验了排除或子集化全球众包Geowiki农田数据集的方法,以实证评估数据量与数据质量(就与尼日利亚目标数据分布的相似度而言)之间的权衡。结果表明,现有世界覆盖图性能最佳,测试集F1分数为0.825,准确率为0.870;紧随其后的是使用人工标注训练样本与尼日利亚Geowiki数据点训练的单头LSTM模型,其F1分数为0.814,准确率为0.842。