Large language models (LLMs) have recently gained much popularity due to their surprising ability at generating human-like English sentences. LLMs are essentially predictors, estimating the probability of a sequence of words given the past. Therefore, it is natural to evaluate their performance from a universal prediction perspective. In order to do that fairly, we introduce the notion of batch regret as a modification of the classical average regret, and we study its asymptotical value for add-constant predictors, in the case of memoryless sources and first-order Markov sources.
翻译:大语言模型(LLMs)近期因其生成类似人类英文句子的惊人能力而广受欢迎。LLMs本质上是预测器,能够根据过去的信息估计单词序列的概率。因此,从通用预测的角度评估其性能是自然而然的。为公平地进行评估,我们引入了批量遗憾(batch regret)这一概念,作为经典平均遗憾的改进形式,并研究了在无记忆信源和一阶马尔可夫信源的情况下,加常数预测器(add-constant predictors)的渐近值。