Mining large datasets and obtaining calibrated predictions from tem is of immediate relevance and utility in reliable deep learning. In our work, we develop methods for Deep neural networks based inferences in such datasets like the Gene Expression. However, unlike typical Deep learning methods, our inferential technique, while achieving state-of-the-art performance in terms of accuracy, can also provide explanations, and report uncertainty estimates. We adopt the Quantile Regression framework to predict full conditional quantiles for a given set of housekeeping gene expressions. Conditional quantiles, in addition to being useful in providing rich interpretations of the predictions, are also robust to measurement noise. Our technique is particularly consequential in High-throughput Genomics, an area which is ushering a new era in personalized health care, and targeted drug design and delivery. However, check loss, used in quantile regression to drive the estimation process is not differentiable. We propose log-cosh as a smooth-alternative to the check loss. We apply our methods on GEO microarray dataset. We also extend the method to binary classification setting. Furthermore, we investigate other consequences of the smoothness of the loss in faster convergence. We further apply the classification framework to other healthcare inference tasks such as heart disease, breast cancer, diabetes etc. As a test of generalization ability of our framework, other non-healthcare related data sets for regression and classification tasks are also evaluated.
翻译:从大型数据集中挖掘信息并获取校准预测,对可靠的深度学习具有直接的相关性和实用性。本研究针对基因表达此类数据集,开发了基于深度神经网络的推断方法。与典型深度学习方法不同,我们提出的推断技术在实现最先进准确率的同时,还能提供可解释性及不确定性估计报告。采用分位数回归框架,针对给定的管家基因表达集合预测完整条件分位数。条件分位数除能为预测提供丰富解读外,还对测量噪声具有鲁棒性。该技术对推动个性化医疗新时代的高通量基因组学领域尤为关键,同时可应用于靶向药物设计与递送。然而,分位数回归中驱动估计过程的检查损失函数不可微。本文提出以log-cosh函数作为检查损失的平滑替代方案。该方法已应用于GEO微阵列数据集,并拓展至二分类场景。此外,我们探究了损失函数平滑性对加速收敛的影响,并将分类框架扩展至心脏病、乳腺癌、糖尿病等医疗推断任务。为验证框架泛化能力,还评估了其他非医疗领域(回归与分类任务)数据集的表现。