We introduce the concept of "universal password model" -- a password model that, once pre-trained, can automatically change its guessing strategy based on the target system. To achieve this, the model does not need to access any plaintext passwords from the target credentials. Instead, it exploits users' auxiliary information, such as email addresses, as a proxy signal to predict the underlying password distribution. Specifically, the model uses deep learning to capture the correlation between the auxiliary data of a group of users (e.g., users of a web application) and their passwords. It then exploits those patterns to create a tailored password model for the target system at inference time. No further training steps, targeted data collection, or prior knowledge of the community's password distribution is required. Besides improving over current password strength estimation techniques and attacks, the model enables any end-user (e.g., system administrators) to autonomously generate tailored password models for their systems without the often unworkable requirements of collecting suitable training data and fitting the underlying machine learning model. Ultimately, our framework enables the democratization of well-calibrated password models to the community, addressing a major challenge in the deployment of password security solutions at scale.
翻译:我们提出“通用密码模型”概念——该模型经过预训练后,可根据目标系统自动调整其猜测策略。这一过程无需访问目标系统的明文密码,而是利用用户的辅助信息(如电子邮件地址)作为代理信号来预测底层密码分布。具体而言,模型通过深度学习捕捉一组用户(例如某Web应用的用户群体)的辅助数据与其密码之间的关联性,进而在推理阶段利用这些关联模式为目标系统生成定制化密码模型。这一过程无需额外训练、针对性数据采集或对群体密码分布的预先了解。该模型不仅改进了现有密码强度评估与攻击技术,更使任何终端用户(如系统管理员)能够自主生成适配其系统的密码模型,而无需满足收集合适训练数据并拟合底层机器学习模型这类往往难以实现的条件。最终,我们的框架推动了可校准密码模型的民主化,解决了大规模部署密码安全解决方案的核心挑战。