Epigraphy increasingly turns to modern artificial intelligence (AI) technologies such as machine learning (ML) for extracting insights from ancient inscriptions. However, scarce labeled data for training ML algorithms severely limits current techniques, especially for ancient scripts like Old Aramaic. Our research pioneers an innovative methodology for generating synthetic training data tailored to Old Aramaic letters. Our pipeline synthesizes photo-realistic Aramaic letter datasets, incorporating textural features, lighting, damage, and augmentations to mimic real-world inscription diversity. Despite minimal real examples, we engineer a dataset of 250,000 training and 25,000 validation images covering the 22 letter classes in the Aramaic alphabet. This comprehensive corpus provides a robust volume of data for training a residual neural network (ResNet) to classify highly degraded Aramaic letters. The ResNet model demonstrates high accuracy in classifying real images from the 8th century BCE Hadad statue inscription. Additional experiments validate performance on varying materials and styles, proving effective generalization. Our results validate the model's capabilities in handling diverse real-world scenarios, proving the viability of our synthetic data approach and avoiding the dependence on scarce training data that has constrained epigraphic analysis. Our innovative framework elevates interpretation accuracy on damaged inscriptions, thus enhancing knowledge extraction from these historical resources.
翻译:金石学日益依赖现代人工智能技术(如机器学习)从古代铭文中提取信息。然而,用于训练机器学习算法的标注数据稀缺严重限制了现有技术,尤其是对古阿拉姆语等古代文字的研究。本研究开创性地提出一种为古阿拉姆语字母量身定制的合成训练数据生成方法。我们的流程可合成具有照片真实感的阿拉姆语字母数据集,整合纹理特征、光照、破损及数据增强技术,以模拟真实铭文的多样性。尽管真实样本极少,我们仍构建了包含25万张训练图像和2.5万张验证图像的数据集,覆盖阿拉姆语字母表中的22个字母类别。这一大型语料库为训练残差神经网络分类高度退化的阿拉姆语字母提供了充足数据。该残差神经网络模型在分类公元前8世纪哈达德雕像铭文的真实图像时展现出高精度。附加实验验证了模型在不同材质和风格上的性能,证明其具备有效的泛化能力。我们的结果证实了模型处理多种真实场景的能力,验证了合成数据方法的可行性,并摆脱了制约金石学分析的训练数据稀缺困境。该创新框架提升了破损铭文的解读准确度,从而增强了对这些历史资源的知识提取能力。