In Stack Overflow (SO), the quality of posts (i.e., questions and answers) is subjectively evaluated by users through a voting mechanism. The net votes (upvotes - downvotes) obtained by a post are often considered an approximation of its quality. However, about half of the questions that received working solutions got more downvotes than upvotes. Furthermore, about 18% of the accepted answers (i.e., verified solutions) also do not score the maximum votes. All these counter-intuitive findings cast doubts on the reliability of the evaluation mechanism employed at SO. Moreover, many users raise concerns against the evaluation, especially downvotes to their posts. Therefore, rigorous verification of the subjective evaluation is highly warranted to ensure a non-biased and reliable quality assessment mechanism. In this paper, we compare the subjective assessment of questions with their objective assessment using 2.5 million questions and ten text analysis metrics. According to our investigation, four objective metrics agree with the subjective evaluation, two do not agree, one either agrees or disagrees, and the remaining three neither agree nor disagree with the subjective evaluation. We then develop machine learning models to classify the promoted and discouraged questions. Our models outperform the state-of-the-art models with a maximum of about 76% - 87% accuracy.
翻译:在Stack Overflow(SO)中,帖子的质量(即问题和答案)通过用户投票机制进行主观评估。帖子获得的净票数(赞成票-反对票)常被视作其质量的近似指标。然而,约半数获得有效解决方案的问题的反对票数高于赞成票数。此外,约18%的被采纳答案(即已验证的解决方案)也未获得最高票数。这些与直觉相悖的发现对SO采用的评估机制的可靠性提出了质疑。同时,许多用户对评估结果表示不满,尤其是针对其帖子收到的反对票。因此,亟需对主观评估进行严谨验证,以确保建立无偏且可靠的质量评估机制。本文利用250万个问题与十项文本分析指标,将问题的主观评估与客观评估进行对比。研究发现:四项客观指标与主观评估结果一致,两项不一致,一项存在部分一致,其余三项与主观评估结果既不一致也不矛盾。我们进一步构建机器学习模型对受鼓励与受抑制的问题进行分类。该模型在准确率上最高达到约76%-87%,优于现有最优模型。