When using a public communication channel--whether formal or informal, such as commenting or posting on social media--end users have no expectation of privacy: they compose a message and broadcast it for the world to see. Even if an end user takes utmost precautions to anonymize their online presence--using an alias or pseudonym; masking their IP address; spoofing their geolocation; concealing their operating system and user agent; deploying encryption; registering with a disposable phone number or email; disabling non-essential settings; revoking permissions; and blocking cookies and fingerprinting--one obvious element still lingers: the message itself. Assuming they avoid lapses in judgment or accidental self-exposure, there should be little evidence to validate their actual identity, right? Wrong. The content of their message--necessarily open for public consumption--exposes an attack vector: stylometric analysis, or author profiling. In this paper, we dissect the technique of stylometry, discuss an antithetical counter-strategy in adversarial stylometry, and devise enhancements through Unicode steganography.
翻译:当使用公开通信渠道时——无论是正式还是非正式的,例如在社交媒体上发表评论或发帖——终端用户没有隐私预期:他们编写一条信息并广播给全世界。即便终端用户采取最极端的措施来匿名化其在线存在——使用别名或假名;隐藏其IP地址;伪造其地理位置;掩盖其操作系统和用户代理;部署加密技术;使用一次性电话号码或邮箱注册;禁用非必要设置;撤销权限;以及屏蔽Cookie和指纹识别——仍有一个明显的元素挥之不去:信息本身。假设他们避免判断失误或意外自我暴露,几乎不应有证据能证实其真实身份,对吗?错。他们的信息内容——必然公开以供公众消费——暴露了一个攻击向量:文体特征分析,即作者画像。在本文中,我们剖析了文体特征分析技术,讨论了对抗式文体特征分析中的一种反向策略,并通过Unicode隐写术设计了增强方案。