Our paper presents team MasonTigers submission to the SemEval-2024 Task 9 - which provides a dataset of puzzles for testing natural language understanding. We employ large language models (LLMs) to solve this task through several prompting techniques. Zero-shot and few-shot prompting generate reasonably good results when tested with proprietary LLMs, compared to the open-source models. We obtain further improved results with chain-of-thought prompting, an iterative prompting method that breaks down the reasoning process step-by-step. We obtain our best results by utilizing an ensemble of chain-of-thought prompts, placing 2nd in the word puzzle subtask and 13th in the sentence puzzle subtask. The strong performance of prompted LLMs demonstrates their capability for complex reasoning when provided with a decomposition of the thought process. Our work sheds light on how step-wise explanatory prompts can unlock more of the knowledge encoded in the parameters of large models.
翻译:本文介绍了MasonTigers团队在SemEval-2024任务9中的提交成果——该任务提供了一套用于测试自然语言理解的谜题数据集。我们采用大型语言模型(LLMs),通过多种提示技术解决该任务。与开源模型相比,零样本和少样本提示在使用专有LLMs测试时生成的结果较为理想。通过思维链提示(一种逐步分解推理过程的迭代式提示方法),我们进一步提升了结果质量。利用思维链提示的集成方法,我们取得了最佳成绩,在单词谜题子任务中位列第二,在句子谜题子任务中位列第十三。所提示的LLMs展现出的强性能,证明了当提供分解的思维过程时,其具备进行复杂推理的能力。本研究揭示了逐步解释性提示如何能够解锁大型模型参数中编码的更多知识。