The dominance of proprietary LLMs has led to restricted access and raised information privacy concerns. High-performing open-source alternatives are crucial for information-sensitive and high-volume applications but often lag behind in performance. To address this gap, we propose (1) A untargeted variant of iterative self-critique and self-refinement devoid of external influence. (2) A novel ranking metric - Performance, Refinement, and Inference Cost Score (PeRFICS) - to find the optimal model for a given task considering refined performance and cost. Our experiments show that SoTA open source models of varying sizes from 7B - 65B, on average, improve 8.2% from their baseline performance. Strikingly, even models with extremely small memory footprints, such as Vicuna-7B, show a 11.74% improvement overall and up to a 25.39% improvement in high-creativity, open ended tasks on the Vicuna benchmark. Vicuna-13B takes it a step further and outperforms ChatGPT post-refinement. This work has profound implications for resource-constrained and information-sensitive environments seeking to leverage LLMs without incurring prohibitive costs, compromising on performance and privacy. The domain-agnostic self-refinement process coupled with our novel ranking metric facilitates informed decision-making in model selection, thereby reducing costs and democratizing access to high-performing language models, as evidenced by case studies.
翻译:专有LLMs的主导地位导致访问受限并引发信息隐私担忧。高性能的开源替代方案对于信息敏感及高容量应用至关重要,但往往在性能上落后于专有模型。为弥合这一差距,我们提出:(1)一种无需外部干预的无目标迭代自我批判与自我精炼变体;(2)一种新型排序指标——性能、精炼与推理成本评分(PeRFICS)——用于在考虑精炼性能与成本的前提下,为给定任务寻找最优模型。实验表明,规模从7B至65B不等的最先进开源模型,其基线性能平均提升8.2%。值得注意的是,即便是内存占用极小的模型(如Vicuna-7B),在Vicuna基准测试中整体性能提升11.74%,而在高创造性开放式任务中提升幅度高达25.39%。Vicuna-13B更进一步,精炼后性能超越ChatGPT。本研究对希望在无需承担过高成本、不牺牲性能与隐私的前提下使用LLMs的资源受限及信息敏感环境具有深远意义。领域无关的自我精炼过程结合新型排序指标,有助于在模型选择中做出明智决策,从而降低成本并推动高性能语言模型的民主化访问——案例研究亦证实了这一点。