With large language models surpassing human performance on an increasing number of benchmarks, we must take a principled approach for targeted evaluation of model capabilities. Inspired by pseudorandomness, we propose pseudointelligence, which captures the maxim that "(perceived) intelligence lies in the eye of the beholder". That is, that claims of intelligence are meaningful only when their evaluator is taken into account. Concretely, we propose a complexity-theoretic framework of model evaluation cast as a dynamic interaction between a model and a learned evaluator. We demonstrate that this framework can be used to reason about two case studies in language model evaluation, as well as analyze existing evaluation methods.
翻译:随着大型语言模型在越来越多的基准测试中超越人类表现,我们必须采取一种原则性方法来对模型能力进行针对性评估。受伪随机性启发,我们提出"伪智能"概念,它体现了"(感知到的)智能存在于观察者眼中"这一准则。即,关于智能的断言仅在其评估者被纳入考量时才有意义。具体而言,我们提出一个复杂度理论框架,将模型评估建模为模型与经过训练的评估器之间的动态交互过程。我们证明该框架可用于推演语言模型评估中的两个案例研究,并对现有评估方法进行分析。