The Hilbert-Schmidt Independence Criterion (HSIC) and its joint-independence extension $d\mathrm{HSIC}$ are degenerate $V$-statistics whose data-dependent weighted-$χ^2$ null limits force a permutation calibration that multiplies the per-test cost by the number of permutations, in practice two orders of magnitude. Adapting the recent martingale MMD construction for two-sample testing to the (joint) independence problem, we introduce two studentised statistics whose null distributions are standard normal regardless of the data law, so that a single normal-quantile lookup replaces the permutation step entirely. The first, $m\mathrm{HSIC}$, is a self-normalised lower-triangular sum of the Hadamard product of two empirically centred Gram matrices. Under independence and bounded-fourth-moment kernels it converges to a standard normal. It is consistent against every fixed alternative, and runs at quadratic cost in the sample size without any sample split, matching the biased HSIC $V$-statistic. Our second statistic, $md\mathrm{HSIC}$, achieves finite-sample consistency with a single half-sample split: the centring is estimated on one half and the lower-triangular self-normalised martingale is run on the other, shrinking the conditional-mean residual to a quantity that is exponentially small in $d$, so the statistic is asymptotically standard normal at every fixed number of jointly tested variables, with a per-test cost that grows only linearly in $d$. On synthetic data with per-variable input dimension from $1$ to $500$ and between $2$ and $10$ jointly tested variables, both statistics match the empirical type-I error rate and test power of permutation-calibrated baselines while running $25$ to $60\times$ faster.


翻译:希尔伯特-施密特独立准则(HSIC)及其联合独立性扩展$d\mathrm{HSIC}$是退化的$V$-统计量,其依赖于数据的加权-$\chi^2$零分布迫使使用排列校准,使得每次检验的成本乘以排列次数,实际中通常高出两个数量级。通过将最近用于双样本检验的鞅MMD构造方法适配到(联合)独立性问题,我们引入了两种学生化统计量,其零分布为标准正态分布且与数据律无关,从而完全用单次正态分位数查找替代了排列步骤。第一个统计量$m\mathrm{HSIC}$是两个经验中心化格拉姆矩阵的哈达玛乘积的自归一化下三角和。在独立性及有界四阶矩核条件下,它收敛于标准正态分布。该统计量对每个固定备择假设具有一致性,且不进行样本分割即可实现与有偏HSIC $V$-统计量相当的样本量二次成本。第二个统计量$md\mathrm{HSIC}$通过单次半样本分割实现有限样本一致性:在半样本上估计中心化,在另一半样本上执行下三角自归一化鞅,将条件均值残差压缩至随$d$指数级减小的量,因此在任意固定联合检验变量数下该统计量渐近服从标准正态分布,且每次检验成本仅随$d$线性增长。在合成数据实验(每变量输入维度从1到500、联合检验变量数从2到10)中,两个统计量均达到与排列校准基线相当的经验I类错误率和检验功效,同时运行速度快25至60倍。

0
下载
关闭预览

相关内容

【剑桥大学博士论文】联邦自监督学习,141页pdf
专知会员服务
19+阅读 · 2024年6月15日
异常检测(Anomaly Detection)综述
极市平台
20+阅读 · 2020年10月24日
换个角度看GAN:另一种损失函数
机器之心
16+阅读 · 2019年1月1日
最新|深度离散哈希算法,可用于图像检索!
全球人工智能
14+阅读 · 2017年12月15日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月18日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
8+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
10+阅读 · 8月1日
相关VIP内容
【剑桥大学博士论文】联邦自监督学习,141页pdf
专知会员服务
19+阅读 · 2024年6月15日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员