Given a small random sample of $n$-bit strings labeled by an unknown Boolean function, which properties of this function can be tested computationally efficiently? We show an equivalence between properties that are efficiently testable from few samples and properties with structured symmetry, which depend only on the function's average values on an efficiently computable partition of the domain. Without the efficiency constraint, a similar characterization in terms of unstructured symmetry was obtained by Blais and Yoshida (2019). We also give a function testing analogue of the classic characterization of testable graph properties in terms of regular partitions, as well as a sublinear time and differentially private algorithm to compute concise summaries of such partitions of graphs. Finally, we tighten a recent characterization of the computational indistinguishability of product distributions, which encompasses the related task of efficiently testing which of two candidate functions labeled the observed samples. Essential to our proofs is the following observation of independent interest: Every randomized Boolean function, no matter how complex, admits a supersimulator: a randomized polynomial-size circuit whose output on random inputs cannot be efficiently distinguished from reality with constant advantage, even by polynomially larger distinguishers. This surprising fact is implicit in a theorem of Dwork et al. (2021) in the context of algorithmic fairness, but its complexity-theoretic implications were not previously explored. We give a new proof of this lemma using an iteration technique from the graph regularity literature, and we observe that a subtle quantifier switch allows it to powerfully circumvent known barriers to improving the landmark complexity-theoretic regularity lemma of Trevisan, Tulsiani, and Vadhan (2009).
翻译:给定一个由未知布尔函数标记的$n$比特字符串的小型随机样本,该函数的哪些属性可以高效地进行计算测试?我们证明了从少量样本中可高效测试的属性与具有结构化对称性的属性之间存在等价关系,这些属性仅取决于函数在域的一个高效可计算划分上的平均值。在无效率约束的情况下,Blais和Yoshida(2019)已基于非结构化对称性得到了类似的刻画。我们还给出了可测试图属性在正则划分方面的经典刻画的函数测试类比,并提出了一种次线性时间和差分隐私算法,用于计算此类图划分的简明摘要。最后,我们强化了最近关于乘积分布计算不可区分性的刻画,该刻画涵盖了高效测试两个候选函数中哪一个标记了观察样本的相关任务。我们证明的关键在于以下具有独立意义的观察:每个随机化布尔函数,无论其多么复杂,都允许一个超模拟器:一个随机化多项式大小的电路,其在随机输入上的输出无法以恒定优势被高效地与真实情况区分,即使区分器具有多项式更大的规模。这一惊人事实隐含在Dwork等人(2021)关于算法公平性的定理中,但其计算复杂性方面的含义此前未被探索。我们利用图正则性文献中的迭代技术给出了该引理的一个新证明,并观察到微妙的量词转换使其能够巧妙地绕过已知障碍,从而改进了Trevisan、Tulsiani和Vadhan(2009)里程碑式的计算复杂性正则性引理。