Engineers often measure many quantities-speed, pressure, temperature, length-expressed in different physical units. The Buckingham Pi-grec theorem states that these variables can always be combined into a smaller set of dimensionless numbers whose values fully determine the system's behaviour. Identifying the appropriate dimensionless groups has traditionally required expert knowledge and physical insight. This paper shows that they can instead be discovered automatically from data, without prior knowledge of the governing physics. The key observation is that, after logarithmic transformation, measurements collected under different scalings of the same system lie on a low-dimensional manifold whose geometry is determined by the underlying dimensionless groups. Singular value decomposition (SVD) identifies this manifold directly from data. A subsequent search over integer-exponent combinations recovers candidate dimensionless quantities, while a repeating-variable filter retains only those constructed from the machine's characteristic scales. This procedure recovers familiar engineering groups, including the flow coefficient, head coefficient, and Mach number, while excluding equivalent but less interpretable alternatives. The method is demonstrated on a synthetic compressor dataset containing 16,000 measurements. Starting from raw dimensional variables and no physics input, it recovers the correct dimensionless groups to numerical precision and reproduces the compressor performance map with an error below 0.01%. More broadly, the work reveals a close connection between classical dimensional analysis and modern data-driven learning. Both rely on the same underlying algebraic structure, suggesting new approaches for building physical models that are simultaneously interpretable, scalable, and data-efficient.
翻译:工程师常测量许多物理量——速度、压力、温度、长度——这些量以不同物理单位表示。白金汉Π定理指出,这些变量总可组合成一组更小的无量纲数,其值完全决定系统行为。传统上,识别恰当的无量纲群需要专业知识和物理直觉。本文表明,这些量纲群可从数据中自动发现,无需预先掌握物理论。关键观察在于:经对数变换后,同一系统在不同缩放比例下采集的测量值位于一个低维流形上,其几何结构由潜在的无量纲群决定。奇异值分解(SVD)直接从数据中识别该流形。随后的整数指数组合搜索可恢复候选无量纲量,而重复变量过滤器仅保留由机器特征尺度构建的组。该过程恢复了包括流量系数、压头系数和马赫数在内的常见工程无量纲群,同时排除了等价但可解释性较差的替代方案。该方法在包含16,000次测量的合成压缩机数据集上进行了验证。从原始量纲变量出发且无物理输入,该方法能以数值精度恢复正确的无量纲群,并再现压缩机性能图,误差低于0.01%。更广泛而言,本研究揭示了经典量纲分析与现代数据驱动学习之间的紧密联系。两者均依赖相同的底层代数结构,为构建同时具备可解释性、可扩展性和数据高效性的物理模型提供了新思路。