Modelling agent preferences has applications in a range of fields including economics and increasingly, artificial intelligence. These preferences are not always known and thus may need to be estimated from observed behavior, in which case a model is required to map agent preferences to behavior, also known as structural estimation. Traditional models are based on the assumption that agents are perfectly rational: that is, they perfectly optimize and behave in accordance with their own interests. Work in the field of behavioral game theory has shown, however, that human agents often make decisions that are imperfectly rational, and the field has developed models that relax the perfect rationality assumption. We apply models developed for predicting behavior towards estimating preferences and show that they outperform both traditional and commonly used benchmark models on data collected from human subjects. In fact, Nash equilibrium and its relaxation, quantal response equilibrium (QRE), can induce an inaccurate estimate of agent preferences when compared against ground truth. A key finding is that modelling non-strategic behavior, conventionally considered uniform noise, is important for estimating preferences. To this end, we introduce quantal-linear4, a rich non-strategic model. We also propose an augmentation to the popular quantal response equilibrium with a non-strategic component. We call this augmented model QRE+L0 and find an improvement in estimating values over the standard QRE. QRE+L0 allows for alternative models of non-strategic behavior in addition to quantal-linear4.
翻译:对代理人偏好的建模在包括经济学及日益增长的人工智能等多个领域具有应用价值。这些偏好并非总为人所知,因此可能需要通过观察到的行为进行估计,此时需要建立从代理人偏好到行为的映射模型(即结构估计)。传统模型基于代理人完全理性的假设:即他们能完美优化自身行为并遵循自身利益。然而,行为博弈理论领域的研究表明,人类代理人做出的决策往往并非完全理性,该领域已发展出放松完全理性假设的模型。我们将用于预测行为的模型应用于偏好估计,并证明在人类受试者数据上,这些模型优于传统及常用基准模型。事实上,与真实基准相比,纳什均衡及其松弛形式——量子响应均衡(QRE)会引发对代理人偏好的不准确估计。关键发现是,对通常被视为均匀噪声的非策略行为进行建模,对于估计偏好至关重要。为此,我们引入了丰富的非策略模型quantal-linear4。我们还提出对流行的量子响应均衡进行增强,加入非策略成分,将增强模型命名为QRE+L0,并发现其对价值估计优于标准QRE。QRE+L0允许除quantal-linear4之外的其他非策略行为模型。