Ensuring the coherence of regional socio-economic statistics is a central task for national statistical institutes. Traditional validation tools, such as range edits, ratio checks, or univariate outlier detection, are effective for identifying extreme values in individual series but are less suited for detecting unusual combinations of indicators in high-dimensional settings. This paper proposes an unsupervised machine learning framework for identifying structurally atypical regional profiles within Europe using publicly available Eurostat data. We construct a cross-sectional dataset of NUTS2 regions (2022) covering four key indicators: GDP per capita in PPS, unemployment rate, tertiary educational attainment, and population density. We apply and compare five anomaly detection techniques, univariate z-scores, Mahalanobis distance, Isolation Forest, Local Outlier Factor, and One-Class SVM, and classify a region as a structural anomaly if it is flagged by at least three of the five methods. The findings show that machine learning methods identify a consistent set of regions whose multivariate profiles diverge substantially from the EU-wide pattern. These include both highly developed metropolitan economies (Brussels, Vienna, Berlin, Prague) and regions with persistent socio-economic disadvantages (Central and Western Slovakia, Northern Hungary, Castilla-La Mancha, Extremadura), as well as Istanbul, whose profile differs markedly from EU capital regions. Importantly, these anomalies do not necessarily signal data quality issues; rather, they reflect meaningful structural divergence that warrants analytical or policy attention. The proposed framework is fully reproducible, scalable, and compatible with existing validation workflows, offering a flexible tool for early detection of unusual regional configurations within the European Statistical System.
翻译:确保区域社会经济统计数据的连贯性是各国统计机构的核心任务。传统验证工具(如范围编辑、比率检验或单变量异常值检测)虽能有效识别单序列极端值,却难以应对高维数据中指标组合异常的问题。本文基于公开的欧盟统计局(Eurostat)数据,提出了一种用于识别欧洲区域结构性异常特征的无监督机器学习框架。我们构建了涵盖2022年NUTS2级区域的横截面数据集,包含四项核心指标:以购买力标准(PPS)计算的人均GDP、失业率、高等教育普及率及人口密度。通过对比五种异常检测技术(单变量Z分数、马氏距离、孤立森林、局部异常因子及单类支持向量机),我们将被至少三种方法标记的区域定义为结构性异常。研究发现,机器学习方法识别出一组多变量特征显著偏离欧盟整体模式的区域,包括高度发达的大都市经济体(布鲁塞尔、维也纳、柏林、布拉格)、持续面临社会经济劣势的区域(斯洛伐克中西部地区、匈牙利北部、卡斯蒂利亚-拉曼恰、埃斯特雷马杜拉),以及特征与欧盟首都区域存在显著差异的伊斯坦布尔。值得强调的是,这些异常并不必然反映数据质量问题,而是体现需要分析或政策关注的结构性分化。本框架具有完全可复现性、可扩展性,且与现有验证工作流兼容,为欧洲统计系统内早期识别异常区域配置提供了灵活工具。