Aggregating data in a database could also be called "integrating along fibers": given functions $\pi\colon E\to D$ and $s\colon E\to R$, where $(R,\circledast)$ is a commutative monoid, we want a new function $(\circledast s)_\pi$ that sends each $d\in D$ to the "sum" of all $s(e)$ for which $\pi(e)=d$. The operation lives alongside querying -- or more generally data migration -- in typical database usage: one wants to know how much Canadians spent on cell phones in 2021, for example, and such requests typically require both aggregation and querying. But whereas querying has an elegant category-theoretic treatment in terms of parametric right adjoints between copresheaf categories, a categorical formulation of aggregation -- especially one that lives alongside that for querying -- appears to be completely absent from the literature. In this paper we show how both querying and aggregation fit into the "polynomial ecosystem". Starting with the category $\mathbf{Poly}$ of polynomial functors in one variable, we review the relatively recent results of Ahman-Uustalu and Garner, which showed that the framed bicategory $\mathbb{C}\mathbf{at}^\sharp$ of comonads in $\mathbf{Poly}$ is precisely the right setting for data migration: its objects are categories and its bicomodules are parametric right adjoints between their copresheaf categories. We then develop a great deal of theory, compressed for space reasons, including local monoidal closed structures, a coclosure to bicomodule composition, and an understanding of adjoints in $\mathbb{C}\mathbf{at}^\sharp$. Doing so allows us to derive interesting mathematical results, e.g.\ that the ordinary operation of transposing a span can be decomposed into the composite of two more primitive operations, and then finally to explain how aggregation arises, alongside querying, in $\mathbb{C}\mathbf{at}^\sharp$.
翻译:数据库中的数据聚合也可称为“沿纤维积分”:给定函数 $\pi\colon E\to D$ 和 $s\colon E\to R$,其中 $(R,\circledast)$ 是交换幺半群,我们定义新函数 $(\circledast s)_\pi$,它将每个 $d\in D$ 映射为所有满足 $\pi(e)=d$ 的 $s(e)$ 之“和”。该操作在典型数据库使用中与查询(更一般地,数据迁移)并存:例如,用户需要了解加拿大人在2021年的手机消费金额,此类请求通常同时涉及聚合与查询。然而,尽管查询已存在优雅的范畴论处理——通过余预层范畴间的参数化右伴随函子——但关于聚合的范畴化表述(尤其是能与查询表述共存的框架)在文献中似乎完全缺失。本文展示了查询与聚合如何融入“多项式生态”。我们从单变量多项式函子范畴 $\mathbf{Poly}$ 出发,回顾Ahman-Uustalu和Garner近期的工作:他们证明$\mathbf{Poly}$中余单子的有框架双范畴 $\mathbb{C}\mathbf{at}^\sharp$ 正是数据迁移的恰当框架——其对象为范畴,双模为余预层范畴间的参数化右伴随函子。随后我们发展了因篇幅而压缩的大量理论,包括局部幺半闭结构、双模复合的余闭包性,以及 $\mathbb{C}\mathbf{at}^\sharp$ 中伴随函子的理解。由此推导出有趣的数学结论(例如,普通跨度转置操作可分解为两个更基本操作的复合),最终阐明在 $\mathbb{C}\mathbf{at}^\sharp$ 中聚合如何与查询并列产生。