Energy-Based Models with Applications to Speech and Language Processing

Energy-Based Models (EBMs) are an important class of probabilistic models, also known as random fields and undirected graphical models. EBMs are un-normalized and thus radically different from other popular self-normalized probabilistic models such as hidden Markov models (HMMs), autoregressive models, generative adversarial nets (GANs) and variational auto-encoders (VAEs). Over the past years, EBMs have attracted increasing interest not only from the core machine learning community, but also from application domains such as speech, vision, natural language processing (NLP) and so on, due to significant theoretical and algorithmic progress. The sequential nature of speech and language also presents special challenges and needs a different treatment from processing fix-dimensional data (e.g., images). Therefore, the purpose of this monograph is to present a systematic introduction to energy-based models, including both algorithmic progress and applications in speech and language processing. First, the basics of EBMs are introduced, including classic models, recent models parameterized by neural networks, sampling methods, and various learning methods from the classic learning algorithms to the most advanced ones. Then, the application of EBMs in three different scenarios is presented, i.e., for modeling marginal, conditional and joint distributions, respectively. 1) EBMs for sequential data with applications in language modeling, where the main focus is on the marginal distribution of a sequence itself; 2) EBMs for modeling conditional distributions of target sequences given observation sequences, with applications in speech recognition, sequence labeling and text generation; 3) EBMs for modeling joint distributions of both sequences of observations and targets, and their applications in semi-supervised learning and calibrated natural language understanding.

翻译：基于能量的模型（EBMs）是一类重要的概率模型，亦称为随机场与无向图模型。EBM属于非归一化模型，因此与隐马尔可夫模型（HMMs）、自回归模型、生成对抗网络（GANs）及变分自编码器（VAEs）等流行的自归一化概率模型有本质区别。近年来，由于理论与算法方面的显著进展，EBM不仅吸引了核心机器学习领域的广泛关注，还引起了语音、视觉、自然语言处理（NLP）等应用领域的兴趣。语音与语言的序列特性带来了特殊挑战，需要与处理固定维度数据（如图像）不同的方法。因此，本专著的目的是系统介绍基于能量的模型，涵盖算法进展及其在语音与语言处理中的应用。首先，介绍EBM的基础知识，包括经典模型、近期基于神经网络参数化的模型、采样方法以及从经典学习算法到最先进方法的各类学习方法。随后，分别阐述EBM在三种不同场景中的应用，即用于建模边际分布、条件分布与联合分布：1）面向序列数据的EBM及其在语言建模中的应用（主要关注序列本身的边际分布）；2）面向给定观测序列的目标序列条件分布建模的EBM，应用于语音识别、序列标注及文本生成；3）面向观测序列与目标序列联合分布建模的EBM，及其在半监督学习与校准型自然语言理解中的应用。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

【亚马逊-WWW2020】不解析,生成!用于面向任务的语义分析的序列到序列体系结构，Don't Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing

专知会员服务

15+阅读 · 2020年2月1日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日