This paper explores the current trending research areas in the field of Computer Science (CS) and investigates the factors contributing to their emergence. Leveraging a comprehensive dataset comprising papers, citations, and funding information, we employ advanced machine learning techniques, including Decision Tree and Logistic Regression models, to predict trending research areas. Our analysis reveals that the number of references cited in research papers (Reference Count) plays a pivotal role in determining trending research areas making reference counts the most relevant factor that drives trend in the CS field. Additionally, the influence of NSF grants and patents on trending topics has increased over time. The Logistic Regression model outperforms the Decision Tree model in predicting trends, exhibiting higher accuracy, precision, recall, and F1 score. By surpassing a random guess baseline, our data-driven approach demonstrates higher accuracy and efficacy in identifying trending research areas. The results offer valuable insights into the trending research areas, providing researchers and institutions with a data-driven foundation for decision-making and future research direction.
翻译:本文探索了计算机科学(CS)领域当前的热点研究区域,并剖析了其新兴趋势的成因。通过利用涵盖论文、引用及经费信息的综合数据集,我们采用包括决策树和逻辑回归模型在内的先进机器学习技术,对热点研究领域进行预测。分析表明,研究论文中的参考文献数量(引用计数)在判定热点研究领域时起关键作用,使引用计数成为驱动计算机科学领域趋势的最相关因素。此外,美国国家科学基金会(NSF)资助项目与专利对热点课题的影响力随时间推移而增强。逻辑回归模型在预测趋势方面优于决策树模型,展现出更高的准确率、精确率、召回率和F1分数。在超越随机猜测基准的基础上,我们的数据驱动方法在识别热点研究领域方面表现出更高的准确性和有效性。研究结果为研究者与机构提供了关于热点领域的重要洞见,为决策制定与未来研究方向奠定了数据驱动基础。