The recent boom in crowdsourcing has opened up a new avenue for utilizing human intelligence in the realm of data analysis. This innovative approach provides a powerful means for connecting online workers to tasks that cannot effectively be done solely by machines or conducted by professional experts due to cost constraints. Within the field of social science, four elements are required to construct a sound crowd - Diversity of Opinion, Independence, Decentralization and Aggregation. However, while the other three components have already been investigated and implemented in existing crowdsourcing platforms, 'Diversity of Opinion' has not been functionally enabled yet. From a computational point of view, constructing a wise crowd necessitates quantitatively modeling and taking diversity into account. There are usually two paradigms in a crowdsourcing marketplace for worker selection: building a crowd to wait for tasks to come and selecting workers for a given task. We propose similarity-driven and task-driven models for both paradigms. Also, we develop efficient and effective algorithms for recruiting a limited number of workers with optimal diversity in both models. To validate our solutions, we conduct extensive experiments using both synthetic datasets and real data sets.
翻译:众包领域的近期繁荣为利用人类智能进行数据分析开辟了新途径。这种创新方法为连接在线工作者与那些因成本限制而无法有效由机器单独完成或由专业专家执行的任务提供了有力手段。在社会学领域,构建一个合理的群体需要四个要素:观点多样性、独立性、分散化和聚合。然而,尽管其他三个组成部分已在现有众包平台中得到研究和实现,“观点多样性”尚未实现功能化。从计算角度来看,构建智慧群体需要对多样性进行量化建模并加以考虑。在众包市场中,工作者选择通常存在两种范式:构建群体以等待任务到来,以及为特定任务选择工作者。我们针对这两种范式分别提出了相似驱动模型和任务驱动模型。同时,我们开发了高效且有效的算法,用于在这两种模型中招募具有最优多样性的有限数量工作者。为验证我们的解决方案,我们利用合成数据集和真实数据集进行了大量实验。