Most tasks in NLP require labeled data. Data labeling is often done on crowdsourcing platforms due to scalability reasons. However, publishing data on public platforms can only be done if no privacy-relevant information is included. Textual data often contains sensitive information like person names or locations. In this work, we investigate how removing personally identifiable information (PII) as well as applying differential privacy (DP) rewriting can enable text with privacy-relevant information to be used for crowdsourcing. We find that DP-rewriting before crowdsourcing can preserve privacy while still leading to good label quality for certain tasks and data. PII-removal led to good label quality in all examined tasks, however, there are no privacy guarantees given.
翻译:自然语言处理中的大多数任务都需要标注数据。由于可扩展性原因,数据标注通常在众包平台上进行。然而,只有当数据不包含隐私相关信息时,才能在公共平台上发布。文本数据通常包含如人名或地点等敏感信息。本研究探讨如何通过消除个人身份信息以及应用差分隐私重写,使包含隐私相关信息的文本能够用于众包。我们发现,在众包之前进行差分隐私重写可以在保护隐私的同时,仍能为某些任务和数据获得良好的标注质量。在所有受检任务中,消除个人身份信息都能带来良好的标注质量,但该方法无法提供隐私保证。