Ensuring the coherence of regional socio-economic statistics is a central task for national statistical institutes. Traditional validation tools, such as range edits, ratio checks, or univariate outlier detection, are effective for identifying extreme values in individual series but are less suited for detecting unusual combinations of indicators in high-dimensional settings. This paper proposes an unsupervised machine learning framework for identifying structurally atypical regional profiles within Europe using publicly available Eurostat data. We construct a cross-sectional dataset of NUTS2 regions (2022) covering four key indicators: GDP per capita in PPS, unemployment rate, tertiary educational attainment, and population density. We apply and compare five anomaly detection techniques, univariate z-scores, Mahalanobis distance, Isolation Forest, Local Outlier Factor, and One-Class SVM, and classify a region as a structural anomaly if it is flagged by at least three of the five methods. The findings show that machine learning methods identify a consistent set of regions whose multivariate profiles diverge substantially from the EU-wide pattern. These include both highly developed metropolitan economies (Brussels, Vienna, Berlin, Prague) and regions with persistent socio-economic disadvantages (Central and Western Slovakia, Northern Hungary, Castilla-La Mancha, Extremadura), as well as Istanbul, whose profile differs markedly from EU capital regions. Importantly, these anomalies do not necessarily signal data quality issues; rather, they reflect meaningful structural divergence that warrants analytical or policy attention. The proposed framework is fully reproducible, scalable, and compatible with existing validation workflows, offering a flexible tool for early detection of unusual regional configurations within the European Statistical System.


翻译:确保区域社会经济统计数据的连贯性是各国统计机构的核心任务。传统验证工具(如范围编辑、比率检验或单变量异常值检测)虽能有效识别单序列极端值,却难以应对高维数据中指标组合异常的问题。本文基于公开的欧盟统计局(Eurostat)数据,提出了一种用于识别欧洲区域结构性异常特征的无监督机器学习框架。我们构建了涵盖2022年NUTS2级区域的横截面数据集,包含四项核心指标:以购买力标准(PPS)计算的人均GDP、失业率、高等教育普及率及人口密度。通过对比五种异常检测技术(单变量Z分数、马氏距离、孤立森林、局部异常因子及单类支持向量机),我们将被至少三种方法标记的区域定义为结构性异常。研究发现,机器学习方法识别出一组多变量特征显著偏离欧盟整体模式的区域,包括高度发达的大都市经济体(布鲁塞尔、维也纳、柏林、布拉格)、持续面临社会经济劣势的区域(斯洛伐克中西部地区、匈牙利北部、卡斯蒂利亚-拉曼恰、埃斯特雷马杜拉),以及特征与欧盟首都区域存在显著差异的伊斯坦布尔。值得强调的是,这些异常并不必然反映数据质量问题,而是体现需要分析或政策关注的结构性分化。本框架具有完全可复现性、可扩展性,且与现有验证工作流兼容,为欧洲统计系统内早期识别异常区域配置提供了灵活工具。

0
下载
关闭预览

相关内容

【WWW2024】知识数据对齐的弱监督异常检测
专知会员服务
23+阅读 · 2024年2月7日
索邦大学121页博士论文《时间序列中的无监督异常检测》
专知会员服务
104+阅读 · 2022年7月25日
基于图注意力机制和Transformer的异常检测
专知会员服务
62+阅读 · 2022年5月16日
异常检测(Anomaly Detection)综述
极市平台
20+阅读 · 2020年10月24日
使用 Canal 实现数据异构
性能与架构
20+阅读 · 2019年3月4日
边缘计算应用:传感数据异常实时检测算法
计算机研究与发展
11+阅读 · 2018年4月10日
无监督学习:决策树AI异常检测
AI前线
15+阅读 · 2018年1月14日
基于机器学习的KPI自动化异常检测系统
运维帮
13+阅读 · 2017年8月16日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
博士论文 | 用代码结构感知方法推进代码大模型
《决策模型比较研究》
专知会员服务
8+阅读 · 7月25日
《美军水下战与海床战概述及本地实施》
专知会员服务
6+阅读 · 7月25日
面向未来冲突推进陆军情报体制改革
专知会员服务
4+阅读 · 7月25日
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
3+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
相关VIP内容
【WWW2024】知识数据对齐的弱监督异常检测
专知会员服务
23+阅读 · 2024年2月7日
索邦大学121页博士论文《时间序列中的无监督异常检测》
专知会员服务
104+阅读 · 2022年7月25日
基于图注意力机制和Transformer的异常检测
专知会员服务
62+阅读 · 2022年5月16日
相关资讯
异常检测(Anomaly Detection)综述
极市平台
20+阅读 · 2020年10月24日
使用 Canal 实现数据异构
性能与架构
20+阅读 · 2019年3月4日
边缘计算应用:传感数据异常实时检测算法
计算机研究与发展
11+阅读 · 2018年4月10日
无监督学习:决策树AI异常检测
AI前线
15+阅读 · 2018年1月14日
基于机器学习的KPI自动化异常检测系统
运维帮
13+阅读 · 2017年8月16日
相关基金
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员