Small molecules in biological samples are studied to provide information about disease states, environmental toxins, natural product drug discovery, and many other applications. The primary window into the composition of small molecule mixtures is tandem mass spectrometry (MS2), which produces data that are of high sensitivity and part per million resolution. We adopt multi-scale sinusoidal embeddings of the mass data in MS2 designed to meet the challenge of learning from the full resolution of MS2 data. Using these embeddings, we provide a new state of the art model for spectral library search, the standard task for initial evaluation of MS2 data. We also introduce a new task, chemical property prediction from MS2 data, that has natural applications in high-throughput MS2 experiments and show that an average $R^2$ of 80\% for novel compounds can be achieved across 10 chemical properties prioritized by medicinal chemists. We use dimensionality reduction techniques and experiments with different floating point resolutions to show the essential role multi-scale sinusoidal embeddings play in learning from MS2 data.
翻译:生物样本中的小分子研究可为疾病状态、环境毒素、天然产物药物发现等多种应用提供信息。分析小分子混合物组成的主要手段是串联质谱(MS2),该技术可产生高灵敏度及百万分之一级分辨率的数据。我们采用针对MS2数据的多尺度正弦嵌入方法,以应对从全分辨率MS2数据中学习的挑战。基于这些嵌入方法,我们提出了光谱库搜索任务(MS2数据初步评估的标准任务)的最新最佳模型。同时引入一项新任务——基于MS2数据的化学性质预测,该任务在高通量MS2实验中具有天然应用价值,并证明对药物化学家优先关注的10项化学性质,新化合物的平均预测决定系数(R²)可达80%。通过降维技术及不同浮点精度实验,我们揭示了多尺度正弦嵌入在MS2数据学习中的关键作用。