The growing volume of data in scientific domains has made spatial query processing increasingly challenging due to high data transfer costs across the memory hierarchy and limited memory bandwidth. To address these bottlenecks and reduce the energy consumed on data movement, this work explores Processing-in-Memory (PIM) systems by executing range queries directly inside memory chips. Unlike prior PIM studies centered on linear scans or hash-based queries, this work is the first to map R-tree range queries onto commercial PIM hardware. The proposed broadcast-based method constructs the R-tree bottom-up on the CPU, broadcasts top levels to UPMEM DPUs (DRAM Processing Units) for global filtering, and distributes lower levels for parallel batched queries in a CPU-DPU system. We evaluate our approach on two real spatial datasets, Sports (999K rectangles) and Lakes (8.4M rectangles), and assess scalability using a synthetic dataset with up to 16M rectangles and 3.9M queries on a commercial UPMEM PIM system with up to 2,540 DPUs. Across all datasets, broadcast-based execution consistently outperforms subtree partitioning by preventing communication from dominating execution. On the Lakes dataset, strong scaling from 512 to 2,540 DPUs reduces kernel time from 64.9 s to 17.6 s, yielding up to 3.66x kernel and 2.70x end-to-end speedup relative to the CPU R-tree search on the same system. The PIM kernel also consumes approximately 3.4x less energy than the corresponding CPU search (e.g., 59.6 kJ vs. 167.0 kJ on Lakes), demonstrating scalable and energy-efficient hierarchical spatial range queries.


翻译:科学领域数据量的持续增长使得空间查询处理面临严峻挑战,主要原因在于内存层级间高昂的数据传输成本和有限的内存带宽。为突破这些瓶颈并降低数据移动能耗,本文探索了内存计算(PIM)系统,通过直接在内存芯片内部执行范围查询来解决问题。不同于以往聚焦于线性扫描或哈希查询的PIM研究,本文首次将R树范围查询映射到商业PIM硬件上。所提出的基于广播的方法在CPU端自底向上构建R树,将顶层广播至UPMEM DPU(DRAM处理单元)进行全局过滤,并将底层分布至CPU-DPU系统中用于并行批量查询。我们使用两个真实空间数据集——Sports(99.9万个矩形)和Lakes(840万个矩形)——对方法进行评估,并通过包含最多1600万个矩形和390万个查询的合成数据集,在配备最多2540个DPU的商业UPMEM PIM系统上测试可扩展性。在所有数据集上,基于广播的执行方法通过避免通信主导执行过程,始终优于子树划分方法。以Lakes数据集为例,从512个DPU强扩展到2540个DPU时,内核时间从64.9秒降至17.6秒,相比同一系统上的CPU R树搜索,实现了最高3.66倍的内核加速比和2.70倍的端到端加速比。此外,PIM内核的能耗约为对应CPU搜索的1/3.4(例如在Lakes数据集上,59.6千焦对比167.0千焦),展示了可扩展且节能的分层空间范围查询能力。

0
下载
关闭预览

相关内容

专知会员服务
29+阅读 · 2021年2月26日
佐治亚理工2020《数据库系统实现》课程,不可错过!
专知会员服务
24+阅读 · 2020年10月14日
【Flink】基于 Flink 的流式数据实时去重
AINLP
14+阅读 · 2020年9月29日
基于MySQL Binlog的Elasticsearch数据同步实践
DBAplus社群
15+阅读 · 2019年9月3日
使用 Canal 实现数据异构
性能与架构
20+阅读 · 2019年3月4日
R语言数据挖掘利器:Rattle包
R语言中文社区
21+阅读 · 2018年11月17日
朴素贝叶斯和贝叶斯网络算法及其R语言实现
R语言中文社区
10+阅读 · 2017年10月2日
基于图片内容的深度学习图片检索(一)
七月在线实验室
20+阅读 · 2017年10月1日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
5+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关资讯
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员