Minimum Description Length (MDL) estimators, using two-part codes for universal coding, are analyzed. For general parametric families under certain regularity conditions, we introduce a two-part code whose regret is close to the minimax regret, where regret of a code with respect to a target family M is the difference between the code length of the code and the ideal code length achieved by an element in M. This is a generalization of the result for exponential families by Gr\"unwald. Our code is constructed by using an augmented structure of M with a bundle of local exponential families for data description, which is not needed for exponential families. This result gives a tight upper bound on risk and loss of the MDL estimators based on the theory introduced by Barron and Cover in 1991. Further, we show that we can apply the result to mixture families, which are a typical example of non-exponential families.
翻译:最小描述长度(MDL)估计器采用两部分编码进行通用编码,本文对此进行了分析。针对满足特定正则条件的一般参数族,我们引入了一种两部分编码,其遗憾值接近极小极大遗憾(其中编码相对于目标族M的遗憾值,是指该编码的码长与M中某元素所能达到的理想码长之差)。这一结论是对Grünwald关于指数族结果的推广。我们的编码通过使用局部指数族纤维丛增强M的结构来构造数据描述,而指数族本身并不需要这种结构。该结果基于Barron与Cover在1991年提出的理论,为MDL估计器的风险与损失提供了紧致上界。此外,我们证明该结论可应用于混合族——这是非指数族中的一个典型范例。