There is a growing tradition in the joint field of network studies and drama history that produces interpretations from the character networks of the plays.The potential of such an interpretation is that the diagrams provide a different representation of the relationships between characters as compared to reading the text or watching the performance. Our aim is to create a method that is able to cluster texts with similar structures on the basis of the play's well-interpretable and simple properties, independent from the number of characters in the drama, or in other words, the size of the network. Finding these features is the most important part of our research, as well as establishing the appropriate statistical procedure to calculate the similarities between the texts. Our data was downloaded from the DraCor database and analyzed in R (we use the GerDracor and the ShakeDraCor sub-collection). We want to propose a robust method based on the distribution of words among characters; distribution of characters in scenes, average length of speech acts, or character-specific and macro-level network properties such as clusterization coefficient and network density. Based on these metrics a supervised classification procedure is applied to the sub-collections to classify comedies and tragedies using the Support Vector Machine (SVM) method. Our research shows that this approach can also produce reliable results on a small sample size.
翻译:网络研究与戏剧史的交叉领域正日益形成一种传统,即从剧本的角色网络中提炼解读。这种解读的潜力在于,与阅读文本或观看表演相比,图表提供了角色之间关系的另一种表征方式。我们的目标是创建一种方法,能够根据剧本易于解释且简洁的性质(独立于剧中角色数量,即网络规模)对具有相似结构的文本进行聚类。寻找这些特征是我们研究的最重要部分,同时还需建立适当的统计程序来计算文本之间的相似度。我们的数据来自DraCor数据库,并使用R语言进行分析(我们使用了GerDracor和ShakeDraCor子集)。我们提出一种基于以下指标的稳健方法:角色间词语分布、场景中角色分布、平均话语长度,以及角色特定和宏观层面的网络属性(如聚类系数和网络密度)。基于这些指标,我们采用监督分类程序对子集进行喜剧与悲剧的分类,应用支持向量机(SVM)方法。我们的研究表明,该方法在小样本量下也能产生可靠的结果。