We present a sampling-based trajectory optimization method derived from the maximum entropy formulation of Differential Dynamic Programming with Tsallis entropy. This method can be seen as a generalization of the legacy work with Shannon entropy, which leads to a Gaussian optimal control policy for exploration during optimization. With the Tsallis entropy, the optimal control policy takes the form of $q$-Gaussian, which further encourages exploration with its heavy-tailed shape. Moreover, in our formulation, the exploration variance, which was scaled by a fixed constant inverse temperature in the original formulation with Shannon entropy, is automatically scaled based on the value function of the trajectory. Due to this property, our algorithms can promote exploration when necessary, that is, the cost of the trajectory is high, rather than using the same scaling factor. The simulation results demonstrate the properties of the proposed algorithm described above.
翻译:我们提出了一种基于Tsallis熵的微分动态规划最大熵公式的采样轨迹优化方法。该方法可视为香农熵经典工作的推广,其导致优化过程中用于探索的高斯最优控制策略。采用Tsallis熵后,最优控制策略呈现$q$-高斯分布形式,凭借其重尾特性进一步促进了探索。此外,在我们的公式中,原本在香农熵原始公式中由固定常数逆温度缩放的探索方差,现在根据轨迹的价值函数自动调整。这一特性使得算法能够在必要时(即轨迹成本较高时)促进探索,而非采用统一缩放因子。仿真结果验证了所提算法的上述特性。