In our opinion the exuberance surrounding the relative success of data-driven large language models (LLMs) is slightly misguided and for several reasons (i) LLMs cannot be relied upon for factual information since for LLMs all ingested text (factual or non-factual) was created equal; (ii) due to their subsymbolic na-ture, whatever 'knowledge' these models acquire about language will always be buried in billions of microfeatures (weights), none of which is meaningful on its own; and (iii) LLMs will often fail to make the correct inferences in several linguistic contexts (e.g., nominal compounds, copredication, quantifier scope ambi-guities, intensional contexts. Since we believe the relative success of data-driven large language models (LLMs) is not a reflection on the symbolic vs. subsymbol-ic debate but a reflection on applying the successful strategy of a bottom-up reverse engineering of language at scale, we suggest in this paper applying the effective bottom-up strategy in a symbolic setting resulting in symbolic, explainable, and ontologically grounded language models.
翻译:我们认为,当前对数据驱动的大语言模型相对成功的狂热情绪略有误导,原因如下:(i)大语言模型不能作为事实信息的可靠来源,因为对所有输入文本(无论事实与否)而言,其处理方式毫无区别;(ii)由于其亚符号特性,这些模型获取的任何语言“知识”都将深埋在数十亿个微观特征(权重)中,每个特征本身毫无意义;(iii)大语言模型在多种语言上下文中往往无法做出正确推理(例如名词复合、共谓述、量词辖域歧义、内涵语境)。由于我们认为数据驱动大语言模型的相对成功并非反映符号主义与亚符号主义之争的胜负,而反映了大规模自底向上逆向工程这一成功策略的应用,本文建议在符号化框架内采用有效的自底向上策略,从而构建符号化、可解释且基于本体论的语言模型。