In this work, we develop language models for the Sanskrit language, namely Bidirectional Encoder Representations from Transformers (BERT) and its variants: A Lite BERT (ALBERT), and Robustly Optimized BERT (RoBERTa) using Devanagari Sanskrit text corpus. Then we extracted the features for the given text from these models. We applied the dimensional reduction and clustering techniques on the features to generate an extractive summary for a given Sanskrit document. Along with the extractive text summarization techniques, we have also created and released a Sanskrit Devanagari text corpus publicly.
翻译:在本研究中,我们利用天城体梵文文本语料库,开发了梵文语言模型,即基于Transformer的双向编码器表示(BERT)及其变体:轻量版BERT(ALBERT)和鲁棒优化版BERT(RoBERTa)。随后,我们从这些模型中提取了给定文本的特征。我们应用降维和聚类技术对特征进行处理,以生成给定梵文文档的抽取式摘要。除抽取式文本摘要技术外,我们还创建并公开发布了一个梵文天城体文本语料库。