This paper is concerned with the computational complexity of learning the Hidden Markov Model (HMM). Although HMMs are some of the most widely used tools in sequential and time series modeling, they are cryptographically hard to learn in the standard setting where one has access to i.i.d. samples of observation sequences. In this paper, we depart from this setup and consider an interactive access model, in which the algorithm can query for samples from the conditional distributions of the HMMs. We show that interactive access to the HMM enables computationally efficient learning algorithms, thereby bypassing cryptographic hardness. Specifically, we obtain efficient algorithms for learning HMMs in two settings: (a) An easier setting where we have query access to the exact conditional probabilities. Here our algorithm runs in polynomial time and makes polynomially many queries to approximate any HMM in total variation distance. (b) A harder setting where we can only obtain samples from the conditional distributions. Here the performance of the algorithm depends on a new parameter, called the fidelity of the HMM. We show that this captures cryptographically hard instances and previously known positive results. We also show that these results extend to a broader class of distributions with latent low rank structure. Our algorithms can be viewed as generalizations and robustifications of Angluin's $L^*$ algorithm for learning deterministic finite automata from membership queries.
翻译:本文关注学习隐马尔可夫模型(HMM)的计算复杂度。尽管HMM是序列和时间序列建模中最广泛使用的工具之一,但在可获取独立同分布观测序列样本的标准设定下,其学习问题具有密码学意义上的难度。本文突破传统框架,考虑一种交互式访问模式——算法可查询HMM条件分布的样本。我们证明,对HMM的交互式访问能够实现计算高效的学习算法,从而绕过密码学困难性。具体而言,我们在两种设定下获得了高效学习HMM的算法:(a)较易设定——可查询精确条件概率,此时算法能在多项式时间内通过多项式次查询逼近总变差距离下的任意HMM;(b)较难设定——仅能获取条件分布样本,此时算法性能取决于称为HMM保真度的新参数。研究表明,该参数可刻画密码学困难实例及先前已知的正面结果。我们进一步证明这些结果可推广至具有隐式低秩结构的更广泛分布类。本文算法可视为Angluin基于成员查询学习确定型有限自动机的$L^*$算法的泛化与鲁棒化版本。