A deployed question answering (QA) model can easily fail when the test data has a distribution shift compared to the training data. Robustness tuning (RT) methods have been widely studied to enhance model robustness against distribution shifts before model deployment. However, can we improve a model after deployment? To answer this question, we evaluate test-time adaptation (TTA) to improve a model after deployment. We first introduce COLDQA, a unified evaluation benchmark for robust QA against text corruption and changes in language and domain. We then evaluate previous TTA methods on COLDQA and compare them to RT methods. We also propose a novel TTA method called online imitation learning (OIL). Through extensive experiments, we find that TTA is comparable to RT methods, and applying TTA after RT can significantly boost the performance on COLDQA. Our proposed OIL improves TTA to be more robust to variation in hyper-parameters and test distributions over time.
翻译:部署后的问答(QA)模型在测试数据相较于训练数据存在分布偏移时容易失效。鲁棒性调优(RT)方法已被广泛研究以在模型部署前增强其对分布偏移的鲁棒性。然而,我们能否在模型部署后进一步提升其性能?为回答这一问题,本研究评估了测试时自适应(TTA)方法在模型部署后的改进效果。我们首先提出COLDQA,一个针对文本损坏、语言变化及领域迁移的鲁棒问答统一评估基准。随后,我们在COLDQA上评估现有TTA方法,并将其与RT方法进行对比。此外,我们提出一种名为在线模仿学习(OIL)的新型TTA方法。通过大量实验发现:TTA与RT方法性能相当,且将TTA应用于RT后可显著提升COLDQA上的表现。我们提出的OIL方法能增强TTA对超参数变化及测试分布时变性的鲁棒性。