In recent years, large language models (LLMs) have achieved strong performance on benchmark tasks, especially in zero or few-shot settings. However, these benchmarks often do not adequately address the challenges posed in the real-world, such as that of hierarchical classification. In order to address this challenge, we propose refactoring conventional tasks on hierarchical datasets into a more indicative long-tail prediction task. We observe LLMs are more prone to failure in these cases. To address these limitations, we propose the use of entailment-contradiction prediction in conjunction with LLMs, which allows for strong performance in a strict zero-shot setting. Importantly, our method does not require any parameter updates, a resource-intensive process and achieves strong performance across multiple datasets.
翻译:近年来,大型语言模型在基准任务中展现出强劲性能,尤其是在零样本或小样本场景下。然而,这些基准通常未能充分应对现实世界中的挑战,例如层级分类任务。为解决这一难题,我们提出将层级数据集上的传统任务重构为更具指示性的长尾预测任务。我们观察到大型语言模型在此类情况下更易出现错误。为应对这些局限,我们提出将蕴含-矛盾预测与大型语言模型相结合,该方法能在严格的零样本设定下实现优异性能。重要的是,我们的方法无需任何参数更新这一资源密集型过程,并在多个数据集上均取得了强劲表现。