Is tractable tokenization for humans also tractable for machine learning models? This study investigates relations between tractable tokenization for humans (e.g., appropriateness and readability) and one for models of machine learning (e.g., performance on an NLP task). We compared six tokenization methods on the Japanese commonsense question-answering dataset (JCommmonsenseQA in JGLUE). We tokenized question texts of the QA dataset with different tokenizers and compared the performance of human annotators and machine-learning models. Besides,we analyze relationships among the performance, appropriateness of tokenization, and response time to questions. This paper provides a quantitative investigation result that shows the tractable tokenizations for humans and machine learning models are not necessarily the same as each other.
翻译:人类可处理的分词(例如,适当性与可读性)对机器学习模型是否同样可处理?本研究探讨了人类可处理分词(如适当性与可读性)与机器学习模型可处理分词(如自然语言处理任务性能)之间的关系。我们针对日本常识问答数据集(JGLUE中的JCommmonsenseQA)比较了六种分词方法。采用不同分词器对问答数据集中的问题文本进行分词,并比较了人类标注员与机器学习模型的性能表现。此外,我们还分析了性能、分词适当性及问题响应时间之间的关系。本文通过定量研究结果表明,人类与机器学习模型的可处理分词并非必然一致。