Training specific deep learning models for particular tasks is common across various domains within seismology. However, this approach encounters two limitations: inadequate labeled data for certain tasks and limited generalization across regions. To address these challenges, we develop SeisCLIP, a seismology foundation model trained through contrastive learning from multi-modal data. It consists of a transformer encoder for extracting crucial features from time-frequency seismic spectrum and an MLP encoder for integrating the phase and source information of the same event. These encoders are jointly pre-trained on a vast dataset and the spectrum encoder is subsequently fine-tuned on smaller datasets for various downstream tasks. Notably, SeisCLIP's performance surpasses that of baseline methods in event classification, localization, and focal mechanism analysis tasks, employing distinct datasets from different regions. In conclusion, SeisCLIP holds significant potential as a foundational model in the field of seismology, paving the way for innovative directions in foundation-model-based seismology research.
翻译:针对地震学各领域中特定任务训练专用深度学习模型是常见做法,但该方法面临两个局限:某些任务缺乏足量标注数据,以及模型跨区域泛化能力有限。为解决这些问题,我们开发了SeisCLIP——一种通过多模态数据对比学习训练的地震学基础模型。该模型由两个编码器组成:用于提取时频地震频谱关键特征的Transformer编码器,以及用于整合同一事件震相和震源信息的MLP编码器。这些编码器在大规模数据集上联合预训练后,频谱编码器可针对不同下游任务在较小数据集上进行微调。值得注意的是,在事件分类、定位和震源机制分析任务中,SeisCLIP使用不同地区的独立数据集均展现出超越基线方法的性能。总之,SeisCLIP作为地震学领域的基础模型具有重要潜力,为基础模型驱动的地震学研究开辟了创新方向。