In strategic classification, agents modify their features, at a cost, to ideally obtain a positive classification from the learner's classifier. The typical response of the learner is to carefully modify their classifier to be robust to such strategic behavior. When reasoning about agent manipulations, most papers that study strategic classification rely on the following strong assumption: agents fully know the exact parameters of the deployed classifier by the learner. This often is an unrealistic assumption when using complex or proprietary machine learning techniques in real-world prediction tasks. We initiate the study of partial information release by the learner in strategic classification. We move away from the traditional assumption that agents have full knowledge of the classifier. Instead, we consider agents that have a common distributional prior on which classifier the learner is using. The learner in our model can reveal truthful, yet not necessarily complete, information about the deployed classifier to the agents. The learner's goal is to release just enough information about the classifier to maximize accuracy. We show how such partial information release can, counter-intuitively, benefit the learner's accuracy, despite increasing agents' abilities to manipulate. We show that while it is intractable to compute the best response of an agent in the general case, there exist oracle-efficient algorithms that can solve the best response of the agents when the learner's hypothesis class is the class of linear classifiers, or when the agents' cost function satisfies a natural notion of submodularity as we define. We then turn our attention to the learner's optimization problem and provide both positive and negative results on the algorithmic problem of how much information the learner should release about the classifier to maximize their expected accuracy.
翻译:在策略分类中,智能体以一定成本修改自身特征,理想情况下能从学习器的分类器获得正面分类。学习器的典型应对方式是谨慎调整其分类器,以对此类策略行为具备鲁棒性。在分析智能体操纵行为时,大多数研究策略分类的论文依赖以下强假设:智能体完全知晓学习器所部署分类器的确切参数。当现实预测任务中使用复杂或专有机器学习技术时,这往往是不切实际的假设。我们开创性地研究了学习者在策略分类中部分信息发布的问题。我们摒弃了智能体完全知晓分类器的传统假设,转而考虑智能体对学习器所使用的分类器持有共同分布先验。在我们的模型中,学习器可以向智能体发布关于已部署分类器的真实但不一定完整的信息。学习器的目标是仅发布足够多的分类器信息以最大化准确率。我们展示了这种部分信息发布如何反直觉地提升学习器的准确率,尽管这增加了智能体的操纵能力。我们证明,虽然一般情况下计算智能体的最优响应是不可行的,但当学习器的假设类别为线性分类器类,或当智能体的成本函数满足我们定义的自然子模性概念时,存在预言机高效算法可求解智能体的最优响应。随后我们关注学习器的优化问题,并就学习器应发布多少分类器信息以最大化其期望准确率这一算法问题提供了正反两方面结果。