Sentence-level representations are beneficial for various natural language processing tasks. It is commonly believed that vector representations can capture rich linguistic properties. Currently, large language models (LMs) achieve state-of-the-art performance on sentence embedding. However, some recent works suggest that vector representations from LMs can cause information leakage. In this work, we further investigate the information leakage issue and propose a generative embedding inversion attack (GEIA) that aims to reconstruct input sequences based only on their sentence embeddings. Given the black-box access to a language model, we treat sentence embeddings as initial tokens' representations and train or fine-tune a powerful decoder model to decode the whole sequences directly. We conduct extensive experiments to demonstrate that our generative inversion attack outperforms previous embedding inversion attacks in classification metrics and generates coherent and contextually similar sentences as the original inputs.
翻译:句子级表示有益于各类自然语言处理任务。人们普遍认为向量表示能够捕捉丰富的语言特性。目前,大型语言模型(LMs)在句子嵌入方面取得了最优性能。然而,近期研究表明,语言模型产生的向量表示可能导致信息泄露。在本文中,我们进一步探究了信息泄露问题,并提出了一种生成式嵌入逆置攻击(GEIA),旨在仅基于句子嵌入重构输入序列。在具备语言模型黑盒访问权限的前提下,我们将句子嵌入视为初始令牌的表示,并训练或微调一个强大的解码器模型,直接解码出完整的序列。通过广泛的实验,我们证明所提出的生成式逆置攻击在分类指标上优于以往的嵌入逆置攻击,并能生成与原始输入连贯且上下文相似的句子。