The Huge Object model is a distribution testing model in which we are given access to independent samples from an unknown distribution over the set of strings $\{0,1\}^n$, but are only allowed to query a few bits from the samples. We investigate the problem of testing whether a distribution is supported on $m$ elements in this model. It turns out that the behavior of this property is surprisingly intricate, especially when also considering the question of adaptivity. We prove lower and upper bounds for both adaptive and non-adaptive algorithms in the one-sided and two-sided error regime. Our bounds are tight when $m$ is fixed to a constant (and the distance parameter $\varepsilon$ is the only variable). For the general case, our bounds are at most $O(\log m)$ apart. In particular, our results show a surprising $O(\log \varepsilon^{-1})$ gap between the number of queries required for non-adaptive testing as compared to adaptive testing. For one sided error testing, we also show that an $O(\log m)$ gap between the number of samples and the number of queries is necessary. Our results utilize a wide variety of combinatorial and probabilistic methods.
翻译:大对象模型是一种分布测试模型,在此模型中,我们可以访问未知分布(该分布定义于字符串集合 $\{0,1\}^n$)的独立样本,但仅允许查询样本中的少量比特位。我们研究在该模型中测试分布是否由 $m$ 个元素支持的问题。结果表明,该属性的行为异常复杂,尤其是在考虑自适应性问题时。我们针对单侧和双侧错误场景下的自适应与非自适应算法,给出了下界与上界。当 $m$ 固定为常数(且距离参数 $\varepsilon$ 为唯一变量)时,我们的界是紧的。对于一般情况,我们的界之间最多相差 $O(\log m)$。特别地,我们的结果揭示了非自适应测试所需查询次数与自适应测试之间存在 $O(\log \varepsilon^{-1})$ 的惊人差距。对于单侧错误测试,我们还证明样本数量与查询次数之间必须存在 $O(\log m)$ 的差距。我们的结果运用了多种组合学与概率方法。