The extraction of key information from receipts is a complex task that involves the recognition and extraction of text from scanned receipts. This process is crucial as it enables the retrieval of essential content and organizing it into structured documents for easy access and analysis. In this paper, we present AMuRD, a novel multilingual human-annotated dataset specifically designed for information extraction from receipts. This dataset comprises $47,720$ samples and addresses the key challenges in information extraction and item classification - the two critical aspects of data analysis in the retail industry. Each sample includes annotations for item names and attributes such as price, brand, and more. This detailed annotation facilitates a comprehensive understanding of each item on the receipt. Furthermore, the dataset provides classification into $44$ distinct product categories. This classification feature allows for a more organized and efficient analysis of the items, enhancing the usability of the dataset for various applications. In our study, we evaluated various language model architectures, e.g., by fine-tuning LLaMA models on the AMuRD dataset. Our approach yielded exceptional results, with an F1 score of 97.43\% and accuracy of 94.99\% in information extraction and classification, and an even higher F1 score of 98.51\% and accuracy of 97.06\% observed in specific tasks. The dataset and code are publicly accessible for further researchhttps://github.com/Update-For-Integrated-Business-AI/AMuRD.
翻译:摘要:从收据中提取关键信息是一项复杂任务,涉及从扫描收据中识别和提取文本。该流程至关重要,因为它能检索核心内容并将其整理成结构化文档,以便于访问和分析。本文提出了AMuRD——一个面向收据信息抽取的新型多语言人工标注数据集。该数据集包含47,720个样本,旨在解决零售业数据分析的两个关键方面:信息抽取与物品分类。每个样本均包含物品名称及其属性(如价格、品牌等)的标注,从而实现对收据中每件物品的全面理解。此外,数据集还提供了44个不同产品类别的分类标注,该分类特性有助于更有序高效地分析物品,增强了数据集在各类应用中的可用性。在本研究中,我们评估了多种语言模型架构,例如通过在AMuRD数据集上微调LLaMA模型。我们的方法取得了卓越成果:信息抽取与分类任务的F1分数达97.43%,准确率达94.99%;在特定任务中F1分数更高达98.51%,准确率达97.06%。数据集与代码已公开,可访问https://github.com/Update-For-Integrated-Business-AI/AMuRD 供进一步研究使用。