The quality of the inferences we make from pathogen sequence data is determined by the number and composition of pathogen sequences that make up the sample used to drive that inference. However, there remains limited guidance on how to best structure and power studies when the end goal is phylogenetic inference. One question that we can attempt to answer with molecular data is whether some people are more likely to transmit a pathogen than others. Here we present an estimator to quantify differential transmission, as measured by the ratio of reproductive numbers between people with different characteristics, using transmission pairs linked by molecular data, along with a sample size calculation for this estimator. We also provide extensions to our method to correct for imperfect identification of transmission linked pairs, overdispersion in the transmission process, and group imbalance. We validate this method via simulation and provide tools to implement it in an R package, phylosamp.
翻译:从病原体序列数据中进行的推断质量,取决于构成该推断样本的病原体序列数量与组成。然而,当最终目标是系统发育推断时,关于如何最优地设计实验和计算统计功效的指导仍十分有限。利用分子数据可尝试回答的一个问题是:某些人是否比其他人更可能传播病原体。本文提出一种基于分子数据关联传播对的估计量,用于量化不同特征人群间繁殖数比值所衡量的差异性传播,并给出该估计量的样本量计算方法。我们还扩展了该方法,以修正传播关联对的不完美识别、传播过程中的过度离散以及群体不平衡等问题。通过模拟验证了该方法的有效性,并在R语言包phylosamp中提供了实现工具。