Fully Probabilistic design (FPD) is a powerful framework offering an elegant and unifying account of stochastic control, learning and decision-making. Here we introduce a generalized FPD framework, which we term as Tsallis FPD. Tsallis FPD uses Tsallis divergence in place of the Kullback-Leibler divergence that defines the standard FPD cost term. Tsallis divergence is a natural generalization of the KL divergence, rooted in non-extensive statistical mechanics and providing flexibility towards modeling stochastic processes with non-Gaussian tail behavior. After formulating Tsallis FPD, we develop a constructive proof of convergence by formulating a fixed point iteration. The construction takes the form of a double iteration scheme that performs a sequence of backwards inductions, rather than a single pass down the stages that constitutes the proven approach for classical FPD. We prove that this construction asymptotically converges to a fixed point and that this fixed point is an optimal solution to Tsallis FPD.
翻译:完全概率设计(FPD)是一个强大的框架,为随机控制、学习和决策制定提供了优雅且统一的描述。本文提出了一种广义的FPD框架,我们称之为Tsallis FPD。Tsallis FPD使用Tsallis散度替代定义标准FPD成本项的Kullback-Leibler散度。Tsallis散度是KL散度的自然推广,根植于非广延统计力学,为模拟具有非高斯尾部行为的随机过程提供了灵活性。在阐述Tsallis FPD后,我们通过构建定点迭代法给出了一个建设性的收敛证明。该构造采用双迭代方案的形式,执行一系列反向归纳,而不是构成经典FPD已证明方法的单次阶段传递。我们证明,该构造渐近收敛于一个定点,且该定点是Tsallis FPD的最优解。