Formality plays a significant role in language communication, especially in low-resource languages such as Hindi, Japanese and Korean. These languages utilise formal and informal expressions to convey messages based on social contexts and relationships. When a language translation technique is used to translate from a source language that does not pertain the formality (e.g. English) to a target language that does, there is a missing information on formality that could be a challenge in producing an accurate outcome. This research explores how this issue should be resolved when machine learning methods are used to translate from English to languages with formality, using Hindi as the example data. This was done by training a bilingual model in a formality-controlled setting and comparing its performance with a pre-trained multilingual model in a similar setting. Since there are not a lot of training data with ground truth, automated annotation techniques were employed to increase the data size. The primary modeling approach involved leveraging transformer models, which have demonstrated effectiveness in various natural language processing tasks. We evaluate the official formality accuracy(ACC) by comparing the predicted masked tokens with the ground truth. This metric provides a quantitative measure of how well the translations align with the desired outputs. Our study showcases a versatile translation strategy that considers the nuances of formality in the target language, catering to diverse language communication needs and scenarios.
翻译:形式性在语言交流中起着重要作用,尤其是在印地语、日语和韩语等低资源语言中。这些语言根据社会语境和关系使用正式与非正式表达传递信息。当使用语言翻译技术将不包含形式性的源语言(如英语)翻译成包含形式性的目标语言时,形式性信息的缺失可能导致翻译结果准确性不足。本研究探讨了如何在使用机器学习方法进行英译含形式性语言(以印地语为例)的翻译时解决这一问题。具体方法为:在形式性受控条件下训练双语模型,并与预训练多语言模型在相似条件下的性能进行对比。由于缺乏大量带标注真值的训练数据,我们采用自动标注技术扩充数据集规模。主要建模方法涉及利用Transformer模型(该模型已在多种自然语言处理任务中展现出有效性)。我们通过将预测掩码标记与真值进行对比,评估官方形式性准确率(ACC),该指标提供了翻译结果与预期输出匹配程度的定量测量。本研究展示了一种兼顾目标语言形式性细微差异的通用翻译策略,可满足多样化的语言交流需求与场景。