Unstructured data are a promising new source of information that insurance companies may use to understand their risk portfolio better and improve the customer experience. However, these novel data sources are difficult to incorporate into existing ratemaking frameworks due to the size and format of the unstructured data. In this paper, we propose a framework to use street view imagery within a generalized linear model. To do so, we use representation learning to extract an embedding vector containing useful information from the image. This embedding is dense and low-dimensional, making it appropriate to use within existing ratemaking models. We find that there is useful information included in street view imagery to predict the frequency of claims for certain types of perils. This model can be used as-is in a ratemaking framework but also opens the door to future empirical research on attempting to extract the causal effect from images that lead to increased or decreased predicted claim frequencies. Throughout, we discuss the practical difficulties (technical and social) of using this type of data for insurance pricing.
翻译:非结构化数据是保险公司可能用于更好理解风险组合及改善客户体验的前景广阔的新型信息源。然而,由于非结构化数据的规模和格式,这些新型数据源难以融入现有费率厘定框架。本文提出一种将街景图像纳入广义线性模型的框架。具体而言,我们采用表示学习从图像中提取包含有用信息的嵌入向量。该嵌入向量具有稠密且低维的特性,适用于现有费率厘定模型。研究发现,街景图像包含预测特定风险类型索赔频率的有效信息。该模型可直接应用于费率厘定框架,同时也为后续实证研究探索如何从图像中提取导致预测索赔频率增减的因果效应奠定了基础。本文自始至终讨论了将此类数据用于保险定价时面临的实际困难(技术层面与社会层面)。