Online propaganda detection pipelines expose measurable privacy risks at multiple stages including data collection, feature extraction, and model inference. We conduct a structured analysis of $162$ peer-reviewed studies and formalize the problem using the Propaganda Risk Online Mitigation and Privacy-preserving Tactics (PROMPT) framework. PROMPT models risks $R$ and mitigation strategies $S$ through a mapping $M: R\to S$ guided by a utility function $α\cdot \mathrm{PrivacyGain}(s_j) - β\cdot \mathrm{PerfLoss}(s_j) - γ\cdot \mathrm{Cost}(s_j)$, with tunable $(α,β,γ)$ enabling stakeholders to balance privacy, accuracy, and deployment costs. To assess practical adoption, we introduce a compliance score that quantifies the alignment of existing methods with GDPR, CCPA etc. requirements. Our evaluation shows that many widely used pipelines remain non-compliant, particularly in metadata handling and user-level aggregation. We further present empirical fine-tuning experiments on transformer-based encoders and decoders under synthetic perturbation, demonstrating a monotonic privacy-utility trade-off: with $q = 0.05$ performance decreased by 1-2% F$_1$, while at $q = 0.20$ the reduction reached 13-14%. These results establish quantitative baselines for privacy costs in propaganda detection. Our contributions include a formal risk-to-defense mapping, a compliance-oriented auditing metric, and experimental evidence of privacy-performance trade-offs, providing a technical foundation for building regulation-compliant and privacy-aware detection systems.
翻译:在线宣传检测流水线在数据收集、特征提取和模型推理等多个阶段暴露了可量化的隐私风险。我们对162篇同行评审研究进行了结构化分析,并利用宣传风险在线缓解与隐私保护策略(PROMPT)框架对该问题进行了形式化定义。PROMPT通过映射函数$M: R\to S$对风险$R$和缓解策略$S$进行建模,该函数由效用函数$\alpha\cdot \mathrm{PrivacyGain}(s_j) - \beta\cdot \mathrm{PerfLoss}(s_j) - \gamma\cdot \mathrm{Cost}(s_j)$指导,其中可调参数$(\alpha,\beta,\gamma)$使得利益相关者能够平衡隐私、准确性和部署成本。为了评估实际应用情况,我们引入了一个合规性评分,用于量化现有方法符合GDPR、CCPA等法规要求的程度。我们的评估表明,许多广泛使用的流水线仍然不合规,特别是在元数据处理和用户级聚合方面。我们进一步在合成扰动条件下对基于Transformer的编码器和解码器进行了经验性微调实验,展示了单调的隐私-效用权衡:当$q = 0.05$时,性能下降了1-2%的F$_1$值,而当$q = 0.20$时,下降幅度达到了13-14%。这些结果为宣传检测中的隐私成本建立了定量基准。我们的贡献包括形式化的风险到防御的映射、面向合规的审计指标以及隐私-性能权衡的实验证据,为构建符合法规和隐私感知的检测系统提供了技术基础。