Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model builders and data providers are often unwilling to reveal their proprietary assets. We present PrivaDE, a privacy-preserving protocol that allows a model owner and a data owner to jointly compute a utility score for a candidate dataset without fully exposing model parameters, raw features, or labels. PrivaDE provides strong security against malicious behavior and can be integrated into blockchain-based marketplaces, where smart contracts enforce fair execution and payment. To make the protocol practical, we propose optimizations to enable efficient secure model inference, and a model-agnostic scoring method that uses only a small, representative subset of the data while still reflecting its impact on downstream training. Evaluation shows that PrivaDE performs data evaluation effectively, achieving online runtimes within 15 minutes even for models with millions of parameters. Our work lays the foundation for fair and automated data marketplaces in decentralized machine learning ecosystems.
翻译:在获取高质量机器学习模型所需数据时,预先评估数据的实用性至关重要,但模型构建者与数据提供者往往不愿透露其专有资产。我们提出PrivaDE——一种隐私保护协议,允许模型拥有者和数据拥有者在无需完全暴露模型参数、原始特征或标签的情况下,联合计算候选数据集的效用分数。PrivaDE可抵御恶意行为带来的强安全威胁,并能集成至基于区块链的市场中,由智能合约确保执行公平性与支付安全性。为提升协议实用性,我们提出了优化方案以实现高效的安全模型推理,并设计了一种模型无关的评分方法——该方法仅使用少量代表性数据子集,仍能反映其对下游训练的影响。评估表明,PrivaDE能有效执行数据评估,即使针对参数规模达数百万的模型,其在线运行时间也保持在15分钟以内。我们的工作为去中心化机器学习生态中公平、自动化的数据市场奠定了基础。