Streaming feature selection techniques have become essential in processing real-time data streams, as they facilitate the identification of the most relevant attributes from continuously updating information. Despite their performance, current algorithms to streaming feature selection frequently fall short in managing biases and avoiding discrimination that could be perpetuated by sensitive attributes, potentially leading to unfair outcomes in the resulting models. To address this issue, we propose FairSFS, a novel algorithm for Fair Streaming Feature Selection, to uphold fairness in the feature selection process without compromising the ability to handle data in an online manner. FairSFS adapts to incoming feature vectors by dynamically adjusting the feature set and discerns the correlations between classification attributes and sensitive attributes from this revised set, thereby forestalling the propagation of sensitive data. Empirical evaluations show that FairSFS not only maintains accuracy that is on par with leading streaming feature selection methods and existing fair feature techniques but also significantly improves fairness metrics.
翻译:流式特征选择技术已成为处理实时数据流的关键手段,其能够从持续更新的信息中识别最相关的属性。尽管现有方法表现优异,但当前流式特征选择算法在管理偏见和避免由敏感属性导致的歧视方面常显不足,这可能使最终模型产生不公平的结果。为解决此问题,我们提出FairSFS,一种新颖的公平流式特征选择算法,旨在维护特征选择过程的公平性,同时不损害在线处理数据的能力。FairSFS通过动态调整特征集以适应输入的特征向量,并从调整后的集合中识别分类属性与敏感属性之间的相关性,从而阻止敏感数据的传播。实证评估表明,FairSFS不仅保持了与主流流式特征选择方法及现有公平特征技术相当的准确性,还显著提升了公平性指标。