Geo-tagged photo based tourist attraction recommendation can discover users' travel preferences from their taken photos, so as to recommend suitable tourist attractions to them. However, existing visual content based methods cannot fully exploit the user and tourist attraction information of photos to extract visual features, and do not differentiate the significances of different photos. In this paper, we propose multi-level visual similarity based personalized tourist attraction recommendation using geo-tagged photos (MEAL). MEAL utilizes the visual contents of photos and interaction behavior data to obtain the final embeddings of users and tourist attractions, which are then used to predict the visit probabilities. Specifically, by crossing the user and tourist attraction information of photos, we define four visual similarity levels and introduce a corresponding quintuplet loss to embed the visual contents of photos. In addition, to capture the significances of different photos, we exploit the self-attention mechanism to obtain the visual representations of users and tourist attractions. We conducted experiments on a dataset crawled from Flickr, and the experimental results proved the advantage of this method.
翻译:基于地理标记照片的旅游景点推荐能够从用户拍摄的照片中挖掘其旅行偏好,从而为其推荐合适的旅游景点。然而,现有基于视觉内容的方法未能充分利用照片中的用户与景点信息来提取视觉特征,且未区分不同照片的重要性。本文提出了一种基于多层级视觉相似性的地理标记照片个性化旅游景点推荐方法(MEAL)。MEAL利用照片的视觉内容与交互行为数据,获取用户与旅游景点的最终嵌入表示,并基于此预测访问概率。具体而言,通过交叉照片中的用户与景点信息,我们定义了四种视觉相似性层级,并引入相应的五元组损失函数来嵌入照片的视觉内容。此外,为捕获不同照片的重要性,我们利用自注意力机制获取用户与旅游景点的视觉表征。在从Flickr爬取的数据集上进行了实验,结果证明了该方法的优势。