We propose the use of conversational GPT models for easy and quick few-shot text classification in the financial domain using the Banking77 dataset. Our approach involves in-context learning with GPT-3.5 and GPT-4, which minimizes the technical expertise required and eliminates the need for expensive GPU computing while yielding quick and accurate results. Additionally, we fine-tune other pre-trained, masked language models with SetFit, a recent contrastive learning technique, to achieve state-of-the-art results both in full-data and few-shot settings. Our findings show that querying GPT-3.5 and GPT-4 can outperform fine-tuned, non-generative models even with fewer examples. However, subscription fees associated with these solutions may be considered costly for small organizations. Lastly, we find that generative models perform better on the given task when shown representative samples selected by a human expert rather than when shown random ones. We conclude that a) our proposed methods offer a practical solution for few-shot tasks in datasets with limited label availability, and b) our state-of-the-art results can inspire future work in the area.
翻译:我们提出利用对话式GPT模型,基于Banking77数据集实现金融领域快速便捷的少样本文本分类方法。该方法采用GPT-3.5和GPT-4的上下文学习技术,在降低技术门槛、免去昂贵GPU计算需求的同时,能快速获得精准分类结果。此外,我们通过SetFit(一种最新的对比学习技术)微调其他预训练的掩码语言模型,在完整数据和少样本场景下均取得最优性能。实验表明,即使使用更少的示例,调用GPT-3.5和GPT-4仍能超越微调后的非生成式模型,但这类解决方案的订阅费用对小型机构而言可能成本过高。最后我们发现,当由领域专家选取代表性样本而非随机样本时,生成式模型在给定任务中表现更优。我们得出结论:(a)所提方法为标签稀缺数据集中的少样本任务提供了实用解决方案;(b)本研究的先进成果可为该领域未来研究提供参考。