Releasing court decisions to the public relies on proper anonymization to protect all involved parties, where necessary. The Swiss Federal Supreme Court relies on an existing system that combines different traditional computational methods with human experts. In this work, we enhance the existing anonymization software using a large dataset annotated with entities to be anonymized. We compared BERT-based models with models pre-trained on in-domain data. Our results show that using in-domain data to pre-train the models further improves the F1-score by more than 5\% compared to existing models. Our work demonstrates that combining existing anonymization methods, such as regular expressions, with machine learning can further reduce manual labor and enhance automatic suggestions.
翻译:公开法院判决需要适当的匿名化处理以保护相关当事方,瑞士联邦最高法院当前采用一套结合传统计算方法与人工专家的现有系统。本研究通过使用标注需匿名化实体的大规模数据集,对现有匿名化软件进行增强。我们比较了基于BERT的模型与领域内数据预训练模型的性能。结果表明,相较于现有模型,采用领域内数据预训练模型可进一步提升F1分数超过5%。本研究证明,将现有匿名化方法(如正则表达式)与机器学习相结合,能进一步减少人工操作并优化自动建议功能。