In this paper, we investigate the effect of addressing difficult samples from a given text dataset on the downstream text classification task. We define difficult samples as being non-obvious cases for text classification by analysing them in the semantic embedding space; specifically - (i) semantically similar samples that belong to different classes and (ii) semantically dissimilar samples that belong to the same class. We propose a penalty function to measure the overall difficulty score of every sample in the dataset. We conduct exhaustive experiments on 13 standard datasets to show a consistent improvement of up to 9% and discuss qualitative results to show effectiveness of our approach in identifying difficult samples for a text classification model.
翻译:本文研究了从给定文本数据集中处理困难样本对下游文本分类任务的影响。我们将困难样本定义为通过语义嵌入空间分析时,在文本分类中属于非明显情形:具体包括两类——(i)语义相似但属于不同类别的样本,以及(ii)语义相异但属于同一类别的样本。我们提出了一种惩罚函数来评估数据集中每个样本的整体困难度分数。我们在13个标准数据集上进行了详尽的实验,结果显示性能持续提升最高达9%,并讨论了定性结果,以证明我们的方法在识别文本分类模型困难样本方面的有效性。