Variable-length Markov chains (VLMCs) are a flexible class of higher-order Markov models that admit a natural representation as context trees. Existing Bayesian methods for specifying prior distributions on tree structures rely on branching processes, but these suffer from a fundamental limitation. The connection between branching probabilities at individual nodes and the structural properties of the induced tree distribution is not straightforward, making it difficult to construct priors encoding specific structural beliefs. We address this limitation by introducing a novel representation of prior distributions on tree space based on context-tree functions. By directly specifying weights for individual contexts through a function on nodes, our approach provides an intuitive mechanism for incorporating structural hypotheses into the prior. This class of distributions maintains computational tractability, allowing marginal likelihoods and posterior mode trees to be computed exactly via generalizations of the Context Tree Weighting (CTW) and Context Tree Maximizing (CTM) algorithms. Exact Bayes factor computation enables rigorous model comparison and hypothesis testing. We demonstrate the flexibility and effectiveness of our approach through simulation studies comparing different prior specifications, and develop practical algorithms for selecting the maximal depth and performing model selection based on Bayes factors.
翻译:变长马尔可夫链(VLMCs)是一类灵活的高阶马尔可夫模型,其自然表示为上下文树。现有基于树结构的先验分布贝叶斯方法依赖于分支过程,但存在根本性局限:单个节点的分支概率与诱导树分布的结构特性之间的关联并不直接,这使得构建编码特定结构信念的先验变得困难。我们通过提出一种基于上下文树函数的树空间先验分布新表示来解决此局限。通过节点函数直接为单个上下文指定权重,我们的方法为将结构假设纳入先验提供了直观机制。该类分布保持计算可处理性,可通过上下文树加权(CTW)和上下文树最大化(CTM)算法的泛化形式精确计算边际似然和后验众数树。精确贝叶斯因子计算能够实现严格的模型比较与假设检验。我们通过比较不同先验设定的仿真研究展示所提方法的灵活性与有效性,并开发基于贝叶斯因子选择最大深度及执行模型选择的实际算法。