The next generation of telescopes will yield a substantial increase in the availability of high-resolution spectroscopic data for thousands of exoplanets. The sheer volume of data and number of planets to be analyzed greatly motivate the development of new, fast and efficient methods for flagging interesting planets for reobservation and detailed analysis. We advocate the application of machine learning (ML) techniques for anomaly (novelty) detection to exoplanet transit spectra, with the goal of identifying planets with unusual chemical composition and even searching for unknown biosignatures. We successfully demonstrate the feasibility of two popular anomaly detection methods (Local Outlier Factor and One Class Support Vector Machine) on a large public database of synthetic spectra. We consider several test cases, each with different levels of instrumental noise. In each case, we use ROC curves to quantify and compare the performance of the two ML techniques.
翻译:下一代望远镜将大幅增加对数千颗系外行星的高分辨率光谱数据的获取。庞大的数据量和待分析行星数量极大地推动了开发快速高效新方法的需求,以便标记出值得重新观测和详细分析的有趣行星。我们主张将机器学习(ML)技术应用于系外行星凌星光谱的异常(新奇性)检测,旨在识别具有异常化学组成的行星,甚至搜寻未知的生物特征。我们成功地在大型合成光谱公共数据库上验证了两种流行的异常检测方法(局部异常因子和单类支持向量机)的可行性。我们考虑了多个测试案例,每个案例具有不同水平的仪器噪声。在每个案例中,我们使用ROC曲线来量化并比较这两种机器学习技术的性能。