When there are signals and noises, physicists try to identify signals by modeling them, whereas statisticians oppositely try to model noise to identify signals. In this study, we applied the statisticians' concept of signal detection of physics data with small-size samples and high dimensions without modeling the signals. Most of the data in nature, whether noises or signals, are assumed to be generated by dynamical systems; thus, there is essentially no distinction between these generating processes. We propose that the correlation length of a dynamical system and the number of samples are crucial for the practical definition of noise variables among the signal variables generated by such a system. Since variables with short-term correlations reach normal distributions faster as the number of samples decreases, they are regarded to be ``noise-like'' variables, whereas variables with opposite properties are ``signal-like'' variables. Normality tests are not effective for data of small-size samples with high dimensions. Therefore, we modeled noises on the basis of the property of a noise variable, that is, the uniformity of the histogram of the probability that a variable is a noise. We devised a method of detecting signal variables from the structural change of the histogram according to the decrease in the number of samples. We applied our method to the data generated by globally coupled map, which can produce time series data with different correlation lengths, and also applied to gene expression data, which are typical static data of small-size samples with high dimensions, and we successfully detected signal variables from them. Moreover, we verified the assumption that the gene expression data also potentially have a dynamical system as their generation model, and found that the assumption is compatible with the results of signal extraction.
翻译:当存在信号和噪声时,物理学家通过建立信号模型来识别信号,而统计学家则相反地通过建立噪声模型来识别信号。本研究采用统计学家的概念,在不建立信号模型的情况下,对具有小样本量和高维特征的物理数据进行信号检测。自然界中大多数数据(无论是噪声还是信号)都被认为是由动力系统生成的,因此这些生成过程之间本质上没有区别。我们提出,动力系统的相关长度和样本数量对于在该系统产生的信号变量中实际定义噪声变量至关重要。由于短程相关变量随着样本数量减少而更快趋近正态分布,因此它们被视为"类噪声"变量,而具有相反特性的变量则为"类信号"变量。正态性检验对于高维小样本数据并不有效。因此,我们基于噪声变量的特性(即变量属于噪声的概率直方图的均匀性)对噪声进行建模。我们设计了一种方法,通过分析随样本数量减少而产生的直方图结构变化来检测信号变量。我们将该方法应用于全局耦合映射产生的数据(可生成不同相关长度的时间序列数据),以及基因表达数据(典型的高维小样本静态数据),并成功从中检测出信号变量。此外,我们验证了基因表达数据也潜在地以动力系统作为其生成模型的假设,并发现该假设与信号提取结果相符。