In a computer system, multiple indispensable components-such as the CPU, memory, and others-work together with other essential components to produce an overall effect, which can only be measured on an independently running system. Since the system operates as an integrated whole, isolating the effect of individual components is challenging. Accurately attributing the system's overall effect to its specific component is crucial for both computer design and evaluation. Taking CPU evaluation as a benchmark, our experiments reveal that the general-purpose rigorous methodologies, like DoE, RCTs, can not address this issue efficiently; A single-purpose empirical methodology, SPEC CPU2017, which is the industry-standard CPU benchmark, only reports the overall effect. Even more concerningly, for the identical CPU, the undefined configurations of other indispensable components introduce uncontrolled variability, with the SPEC scores fluctuating from 12.16\% to 436.80\%. We propose a rigorous methodology that can attribute the overall effect to its specific component, which can be utilized in computer component evaluations and design, as well as in other areas. Through theoretical analysis and pioneering controlled experiments, we systematically compare our methodology against three established methodologies: SPEC CPU2017, DoE, and RCTs. The results show our methodology can achieve its goal in a cost-efficient way, while others exhibit inherent limitations.
翻译:在计算机系统中,多个不可或缺的组件——例如CPU、内存等——与其他必要组件协同工作,共同产生整体效应,该效应仅能在独立运行的系统中被测量。由于系统作为一个整体运行,分离单个组件的影响颇具挑战性。将系统的整体效应准确归因于特定组件,对于计算机的设计与评估至关重要。以CPU评估为基准,我们的实验表明:通用的严谨方法论(如实验设计DoE、随机对照试验RCTs)无法有效解决此问题;单一用途的经验性方法论——行业标准CPU基准SPEC CPU2017——仅报告整体效应。更值得关注的是,对于同一CPU,其他不可或缺组件的未定义配置引入了不可控的变异性,导致SPEC评分波动范围达12.16%至436.80%。我们提出了一种严谨的方法论,能够将整体效应归因于特定组件,该方法可应用于计算机组件评估与设计及其他领域。通过理论分析与开创性控制实验,我们系统性地将本方法论与三种现有方法论(SPEC CPU2017、DoE、RCTs)进行了对比。结果表明:本方法论能以高成本效益实现目标,而其他方法论则存在固有局限性。