Cross-national comparison of research funding projects is increasingly important for science policy and strategic planning, but language differences remain a major obstacle. In particular, KAKENHI project descriptions are written primarily in Japanese, whereas projects from major overseas funding agencies, such as NSF, NIH, and UKRI, are documented in English. This study investigates whether multilingual sentence embeddings can support meaningful cross-lingual comparison of research funding projects, with particular attention to the semantic effects of translating Japanese texts into English. For each KAKENHI project, we construct two representations: the original Japanese text and its machine-translated English version, both embedded in a shared semantic space using a multilingual Sentence-BERT model. We then compare their distances and nearest-neighbor relationships with respect to projects from English-language funding agencies. The results show that the Japanese and translated English representations of the same KAKENHI project are, on average, located closer to one another than to native English projects, indicating substantial cross-lingual alignment. However, the overlap of nearest neighbors between the two representations is limited, averaging 2.9 out of 10. This suggests that multilingual embeddings capture semantic similarity across languages to a meaningful extent, while language differences and translation still affect the local structure of the embedding space. These findings suggest that multilingual embeddings provide a useful basis for large-scale exploratory comparison of funding projects across countries and agencies. At the same time, they offer an empirical reference for assessing semantic drift when Japanese research project data are translated into English for international analysis.


翻译:跨国科研资助项目的比较对于科学政策制定和战略规划日益重要,但语言差异仍是主要障碍。特别是KAKENHI项目描述主要使用日语撰写,而美国国家科学基金会(NSF)、美国国立卫生研究院(NIH)和英国研究与创新署(UKRI)等主要海外资助机构的项目则使用英语记录。本研究探讨多语言句嵌入能否支持科研资助项目有意义的跨语言比较,重点关注将日语文本翻译成英语对语义的影响。针对每个KAKENHI项目,我们构建两种表征形式:原始日语文本及其机器翻译的英语版本,两者均通过多语言Sentence-BERT模型嵌入到共享语义空间中。随后比较它们与英语资助机构项目之间的距离和最近邻关系。结果表明:同一KAKENHI项目的日语表征与翻译英语表征的平均距离,比其与原生英语项目的距离更近,表明存在显著的跨语言对齐。然而,两种表征的最近邻重叠有限,平均仅达10个中的2.9个。这说明多语言嵌入虽能在有意义的程度上捕获跨语言语义相似性,但语言差异与翻译仍会影响嵌入空间的局部结构。这些发现表明,多语言嵌入为跨国跨机构资助项目的大规模探索性比较提供了有效基础,同时为评估日语研究项目数据在翻译为英语进行国际分析时的语义漂移提供了实证参考。

0
下载
关闭预览

相关内容

金融领域自然语言处理研究资源大列表
专知
13+阅读 · 2020年2月27日
中文对比英文自然语言处理NLP的区别综述
AINLP
18+阅读 · 2019年3月20日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
BERT相关论文、文章和代码资源汇总
AINLP
19+阅读 · 2018年11月17日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
8+阅读 · 7月19日
相关VIP内容
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员