We tackle the problem of neural machine translation of mathematical formulae between ambiguous presentation languages and unambiguous content languages. Compared to neural machine translation on natural language, mathematical formulae have a much smaller vocabulary and much longer sequences of symbols, while their translation requires extreme precision to satisfy mathematical information needs. In this work, we perform the tasks of translating from LaTeX to Mathematica as well as from LaTeX to semantic LaTeX. While recurrent, recursive, and transformer networks struggle with preserving all contained information, we find that convolutional sequence-to-sequence networks achieve 95.1% and 90.7% exact matches, respectively.
翻译:我们解决了数学公式在歧义表示语言与无歧义内容语言之间的神经机器翻译问题。与自然语言的神经机器翻译相比,数学公式的词汇量显著较小,符号序列长度更长,且其翻译需满足数学信息需求的极高精度要求。本研究开展了从LaTeX到Mathematica以及从LaTeX到语义LaTeX的翻译任务。尽管循环网络、递归网络和Transformer网络难以完整保留所有包含信息,我们发现卷积序列到序列网络分别实现了95.1%和90.7%的精确匹配率。