In this paper, we describe our participation in the MESINESP Task of the BioASQ biomedical semantic indexing challenge. The participating system follows an approach based solely on conventional information retrieval tools. We have evaluated various alternatives for extracting index terms from IBECS/LILACS documents in order to be stored in an Apache Lucene index. Those indexed representations are queried using the contents of the article to be annotated and a ranked list of candidate labels is created from the retrieved documents. We also have evaluated a sort of limited Label Powerset approach which creates meta-labels joining pairs of DeCS labels with high co-occurrence scores, and an alternative method based on label profile matching. Results obtained in official runs seem to confirm the suitability of this approach for languages like Spanish.
翻译:本文描述了我们在 BioASQ 生物医学语义索引挑战赛的 MESINESP 任务中的参与情况。所采用的系统基于纯粹的传统信息检索工具。我们评估了从 IBECS/LILACS 文档中提取索引词以存储在 Apache Lucene 索引中的多种替代方案。这些索引表示通过待标注文章的内容进行查询,并从检索到的文档中生成候选标签的排序列表。我们还评估了一种有限的 Label Powerset 方法,该方法通过结合高共现得分的 DeCS 标签对来创建元标签,以及一种基于标签轮廓匹配的替代方法。官方运行结果似乎证实了这种方法适用于西班牙语等语言。