Finding XPath Bugs in XML Document Processors via Differential Testing

Extensible Markup Language (XML) is a widely used file format for data storage and transmission. Many XML processors support XPath, a query language that enables the extraction of elements from XML documents. These systems can be affected by logic bugs, which are bugs that cause the processor to return incorrect results. In order to tackle such bugs, we propose a new approach, which we realized as a system called XPress. As a test oracle, XPress relies on differential testing, which compares the results of multiple systems on the same test input, and identifies bugs through discrepancies in their outputs. As test inputs, XPress generates both XML documents and XPath queries. Aiming to generate meaningful queries that compute non-empty results, XPress selects a so-called targeted node to guide the XPath expression generation process. Using the targeted node, XPress generates XPath expressions that reference existing context related to the targeted node, such as its tag name and attributes, while also guaranteeing that a predicate evaluates to true before further expanding the query. We tested our approach on six mature XML processors, BaseX, eXist-DB, Saxon, PostgreSQL, libXML2, and a commercial database system. In total, we have found 20 unique bugs in these systems, of which 25 have been verified by the developers, and 12 of which have been fixed. XPress is efficient, as it finds 12 unique bugs in BaseX in 24 hours, which is 2x as fast as naive random generation. We expect that the effectiveness and simplicity of our approach will help to improve the robustness of many XML processors.

翻译：可扩展标记语言(XML)是一种广泛用于数据存储和传输的文件格式。许多XML处理器支持XPath——一种能够从XML文档中提取元素的查询语言。这些系统可能受逻辑缺陷影响，即导致处理器返回错误结果的缺陷。为应对此类缺陷，我们提出了一种新方法，并将其实现为名为XPress的系统。作为测试预言，XPress采用差分测试技术，通过比较多个系统对同一测试输入的输出结果，借助输出差异识别缺陷。在测试输入方面，XPress同时生成XML文档和XPath查询。为生成能计算非空结果的有意义查询，XPress选择所谓的目标节点来引导XPath表达式生成过程。利用目标节点，XPress生成引用该节点相关上下文（如标签名和属性）的XPath表达式，同时确保谓词在进一步扩展查询前评估为真。我们对六款成熟的XML处理器（BaseX、eXist-DB、Saxon、PostgreSQL、libXML2及一款商业数据库系统）进行了测试。共发现20个独特缺陷，其中25个已获开发者确认，12个已被修复。XPress具有高效性，能在24小时内发现BaseX中的12个独特缺陷，速度是朴素随机生成的2倍。我们预期该方法简洁高效的特点将有助于提升众多XML处理器的鲁棒性。

相关内容

XPath

关注 1

XPath即为XML路径语言，它是一种用来确定XML（标准通用标记语言的子集）文档中某部分位置的语言。XPath基于XML的树状结构，提供在数据结构树中找寻节点的能力。起初 XPath 的提出的初衷是将其作为一个通用的、介于XPointer与XSLT间的语法模型。但是 XPath 很快的被开发者采用来当作小型查询语言。

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日