Two-block data---two sets of variables measured on the same individuals, such as microbial taxa and metabolites---raise the question of how \emph{groups} of covariate variables relate to \emph{groups} of response variables. Co-clustering answers this for a single matrix, not for two variable blocks; existing two-block methods either cluster only one side or return signed factors rather than clusters. Starting from the multivariate linear regression $Y_1\approx M Y_2$, we give its non-negative coefficient matrix a tri-factorization $M=X_1ΘX_2$ (a tri-NMF), so that $X_1$ softly clusters the response variables, $X_2$ the covariate variables, and $Θ$ is a tested matrix of block correspondences. This makes the method the non-negative member of the reduced-rank regression (RRR) family, expressing RRR's low-rank class in a parts-based basis as NMF relates to PCA; the constraint can only restrict the fit, so predictive accuracy is not the aim; the co-clustering and tested correspondences are. We give multiplicative update rules, choose the two ranks by cross-validation, and develop a conditional Wald test for $Θ$ applied after basis selection; its size is nominal with fixed bases, conservative after re-estimation, and slightly above nominal under correlated responses, while a non-zero path's \emph{magnitude} stays conditional on the estimated bases. We illustrate the method---a tri-factorized non-negative RRR (NMF-RRR)---on four data sets spanning a permutation structure (Doubs, community ecology), a weak cross-structure under $p>n$ (nutrimouse, nutrigenomics), a pronounced one in a screened microbiome--metabolome study (FRANZOSA, where two microbial groups are jointly associated with each metabolite module), and a classification special case (Wine).
翻译:暂无翻译