Cross-national comparison of research funding projects is increasingly important for science policy and strategic planning, but language differences remain a major obstacle. In particular, KAKENHI project descriptions are written primarily in Japanese, whereas projects from major overseas funding agencies, such as NSF, NIH, and UKRI, are documented in English. This study investigates whether multilingual sentence embeddings can support meaningful cross-lingual comparison of research funding projects, with particular attention to the semantic effects of translating Japanese texts into English. For each KAKENHI project, we construct two representations: the original Japanese text and its machine-translated English version, both embedded in a shared semantic space using a multilingual Sentence-BERT model. We then compare their distances and nearest-neighbor relationships with respect to projects from English-language funding agencies. The results show that the Japanese and translated English representations of the same KAKENHI project are, on average, located closer to one another than to native English projects, indicating substantial cross-lingual alignment. However, the overlap of nearest neighbors between the two representations is limited, averaging 2.9 out of 10. This suggests that multilingual embeddings capture semantic similarity across languages to a meaningful extent, while language differences and translation still affect the local structure of the embedding space. These findings suggest that multilingual embeddings provide a useful basis for large-scale exploratory comparison of funding projects across countries and agencies. At the same time, they offer an empirical reference for assessing semantic drift when Japanese research project data are translated into English for international analysis.
翻译:跨国科研资助项目的比较对于科学政策制定和战略规划日益重要,但语言差异仍是主要障碍。特别是KAKENHI项目描述主要使用日语撰写,而美国国家科学基金会(NSF)、美国国立卫生研究院(NIH)和英国研究与创新署(UKRI)等主要海外资助机构的项目则使用英语记录。本研究探讨多语言句嵌入能否支持科研资助项目有意义的跨语言比较,重点关注将日语文本翻译成英语对语义的影响。针对每个KAKENHI项目,我们构建两种表征形式:原始日语文本及其机器翻译的英语版本,两者均通过多语言Sentence-BERT模型嵌入到共享语义空间中。随后比较它们与英语资助机构项目之间的距离和最近邻关系。结果表明:同一KAKENHI项目的日语表征与翻译英语表征的平均距离,比其与原生英语项目的距离更近,表明存在显著的跨语言对齐。然而,两种表征的最近邻重叠有限,平均仅达10个中的2.9个。这说明多语言嵌入虽能在有意义的程度上捕获跨语言语义相似性,但语言差异与翻译仍会影响嵌入空间的局部结构。这些发现表明,多语言嵌入为跨国跨机构资助项目的大规模探索性比较提供了有效基础,同时为评估日语研究项目数据在翻译为英语进行国际分析时的语义漂移提供了实证参考。