Investors are interested in predicting future success of startup companies, preferably using publicly available data which can be gathered using free online sources. Using public-only data has been shown to work, but there is still much room for improvement. Two of the best performing prediction experiments use 17 and 49 features respectively, mostly numeric and categorical in nature. In this paper, we significantly expand and diversify both the sources and the number of features (to 171) to achieve better prediction. Data collected from Crunchbase, the Google Search API, and Twitter (now X) are used to predict whether a company will raise a round of funding within a fixed time horizon. Much of the new features are textual and the Twitter subset include linguistic metrics such as measures of passive voice and parts-of-speech. A total of ten machine learning models are also evaluated for best performance. The adaptable model can be used to predict funding 1-5 years into the future, with a variable cutoff threshold to favor either precision or recall. Prediction with comparable assumptions generally achieves F scores above 0.730 which outperforms previous attempts in the literature (0.531), and does so with fewer examples. Furthermore, we find that the vast majority of the performance impact comes from the top 18 of 171 features which are mostly generic company observations, including the best performing individual feature which is the free-form text description of the company.
翻译:投资者对利用公开可用数据(可通过免费在线资源获取)预测初创企业未来成功的能力颇感兴趣。已有研究表明仅使用公开数据可行,但提升空间依然广阔。两项表现最佳的预测实验分别使用了17个和49个特征,且多为数值型与类别型特征。本文中,我们显著扩展并多样化数据来源及特征数量(增至171个),以实现更优预测。通过从Crunchbase、谷歌搜索API及推特(现更名为X)收集的数据,预测企业是否能在固定时间窗口内完成新一轮融资。新增特征多为文本类型,其中推特子集包含被动语态、词性等语言度量指标。我们还评估了十种机器学习模型以确定最佳性能。该可适配模型可用于预测未来1-5年的融资情况,并可通过可变阈值偏向于查准率或查全率。在可比假设条件下,预测F值普遍超过0.730,优于既往文献中的0.531,且所用样本量更少。此外,我们发现绝大部分性能贡献来自171个特征中的前18个——这些多为通用企业观测指标,其中表现最佳的单一特征即为企业自由格式文本描述。