Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-benchmarks, however, rank models as consistently as the full benchmarks they replace? And can they rank models more consistently than selecting a random subset of data points? In many scenarios, we find that the answer is no. We introduce a meta-evaluation measure for micro-benchmarking which investigates how well a micro-benchmark can rank two models as a function of their performance difference on the full benchmark. This approach can determine which model pairs can be ranked correctly by a micro-benchmark, allowing for a finer-grained analysis of the trade-off between micro-benchmark size and reliability. Prior work has suggested selecting as few as 10 examples; we find that no micro-benchmarking method can consistently rank model pairs 3.5 points of accuracy apart on MMLU-Pro or 4 points apart on BIG-bench Hard. In order to consistently rank model pairs with relatively similar performances, we show that often as many as 250 examples must be selected, at which point random sampling is competitive with existing micro-benchmarking methods. When comparing only 8B instruction-tuned models on MMLU-Pro micro-benchmarks with 25 examples, we find that more than half of pairwise comparisons are not likely to be preserved. Our work provides actionable guidance for both micro-benchmark users and developers in navigating the trade-off between evaluation efficiency and reliability.
翻译:微基准测试为语言模型开发中通常难以承受的时间与成本提供了一种解决方案:在现有基准测试的极小子集上进行评估。然而,这些微基准测试能否像它们所替代的完整基准测试一样,对模型进行一致的排序?它们能否比随机选择数据点子集更一致地对模型进行排序?在许多场景中,我们发现答案是否定的。我们提出了一种用于微基准测试的元评估度量,该度量通过研究微基准测试对两个模型的排序能力,作为它们在完整基准测试上性能差异的函数。这种方法可以确定微基准测试能够正确排序哪些模型对,从而允许对微基准测试规模与可靠性之间的权衡进行更细粒度的分析。先前的研究建议选择少至10个示例;我们发现,在MMLU-Pro上准确率相差3.5个点或在BIG-bench Hard上相差4个点的模型对,没有任何微基准测试方法能够一致地对其进行排序。为了能够一致地对性能相对接近的模型对进行排序,我们表明通常需要选择多达250个示例,而在此规模下,随机采样与现有的微基准测试方法相比已具有竞争力。当在MMLU-Pro微基准测试(包含25个示例)上仅比较8B指令微调模型时,我们发现超过一半的成对比较结果很可能无法保持一致。我们的工作为微基准测试的用户和开发者在评估效率与可靠性之间进行权衡时,提供了可操作的指导。