Dataset licensing is currently an issue in the development of machine learning systems. And in the development of machine learning systems, the most widely used are publicly available datasets. However, since the images in the publicly available dataset are mainly obtained from the Internet, some images are not commercially available. Furthermore, developers of machine learning systems do not often care about the license of the dataset when training machine learning models with it. In summary, the licensing of datasets for machine learning systems is in a state of incompleteness in all aspects at this stage. Our investigation of two collection datasets revealed that most of the current datasets lacked licenses, and the lack of licenses made it impossible to determine the commercial availability of the datasets. Therefore, we decided to take a more scientific and systematic approach to investigate the licensing of datasets and the licensing of machine learning systems that use the dataset to make it easier and more compliant for future developers of machine learning systems.
翻译:数据集许可是当前机器学习系统开发中的一个问题。在机器学习系统开发中,应用最广泛的是公开数据集。然而,由于公开数据集中的图像主要来源于互联网,部分图像不具备商业可用性。此外,机器学习系统开发者在利用数据集训练模型时,往往不关心数据集的许可协议。总之,现阶段机器学习系统数据集的许可在各个方面都处于不完善的状态。我们对两个收集数据集的研究表明,当前大多数数据集缺乏许可协议,而许可协议的缺失使得无法判断数据集的商业可用性。因此,我们决定采用更科学、更系统的方法来研究数据集许可以及使用该数据集的机器学习系统的许可,以便为未来的机器学习系统开发者提供更便捷且合规的参考。