Catastrophic events create uncertain situations for humanitarian organizations locating and providing aid to affected people. Many people turn to social media during disasters for requesting help and/or providing relief to others. However, the majority of social media posts seeking help could not properly be detected and remained concealed because often they are noisy and ill-formed. Existing systems lack in planning an effective strategy for tweet preprocessing and grasping the contexts of tweets. This research, first of all, formally defines request tweets in the context of social networking sites, hereafter rweets, along with their different primary types and sub-types. Our main contributions are the identification and categorization of rweets. For rweet identification, we employ two approaches, namely a rule-based and logistic regression, and show their high precision and F1 scores. The rweets classification into sub-types such as medical, food, and shelter, using logistic regression shows promising results and outperforms existing works. Finally, we introduce an architecture to store intermediate data to accelerate the development process of the machine learning classifiers.
翻译:灾难事件给人道主义组织定位和援助受灾人员造成不确定性。许多人在灾害期间转向社交媒体寻求帮助或为他人提供救济。然而,大多数寻求帮助的社交媒体帖子往往因嘈杂且格式不规范而无法被有效检测并隐藏起来。现有系统缺乏对推文预处理的有效策略规划以及语境理解能力。本研究首先在社交网络语境下正式定义求助推文(以下简称rweet)及其不同主要类型与子类型。我们的主要贡献在于对rweet的识别与分类。在rweet识别方面,我们采用了基于规则和逻辑回归两种方法,并证明其具有高精确率和F1得分。使用逻辑回归将rweet分类为医疗、食品、住所等子类型展现出有前景的结果,并优于现有研究。最后,我们引入了一种中间数据存储架构,以加速机器学习分类器的开发进程。