In this paper, we propose Docprompt for document question answering tasks with powerful zero-shot and few-shot performance. We proposed a novel weakly supervised data generation method, a novel multl-stage training method and a novel understanding model \& generation model ensemble method. We achieved state-of-the-art performance on 4 document question answering tasks. This method greatly improves the delivery efficiency and model performance of document question answering customer projects, reducing annotation costs and labor costs. Our demo can be found at https://huggingface.co/spaces/PaddlePaddle/ERNIE-Layout.
翻译:本文提出DocPrompt方法,用于文档问答任务,具备强大的零样本与少样本性能。我们提出了一种新颖的弱监督数据生成方法、一种创新的多阶段训练方法以及一种集成理解模型与生成模型的创新方法。在四个文档问答任务上取得了最优性能。该方法显著提升了文档问答客户项目的交付效率与模型性能,降低了标注成本与人力成本。我们的演示可在 https://huggingface.co/spaces/PaddlePaddle/ERNIE-Layout 处获取。