In recent years, advancements in large language models have been remarkable, with models such as ChatGPT demonstrating exceptional proficiency in diverse linguistic tasks. The pre-training of large models with billions of parameters, poses a formidable challenge, primarily due to the scarcity of datasets of a commensurate scale for effective training. Nevertheless, innovative strategies have emerged, including methods to fine-tune these pre-trained models using fewer parameters set, as evidenced by models like MiniGPT-4 and LLaVA. Despite their potential in various domains, these models remain limited in their understanding of artistic imagery. They have yet to fully grasp the intricate nuances of art images or to provide an objective articulation of the emotions they evoke, in a manner akin to human perception. This work introduces ArtGPT-4, a pioneering large vision-language model tailored to address the deficiencies of contemporary models in artistic comprehension. ArtGPT-4 underwent training on image-text pairs utilizing a Tesla A100 device in a mere 2 hours, with a dataset comprising approximately 0.52M entries. Impressively, the model can render images with an artistic-understanding and convey the emotions they inspire, mirroring human interpretation. Additionally, this work presents a unique dataset designed to evaluate the efficacy of vision-language models. In subsequent evaluations, ArtGPT-4 not only achieved state-of-the-art performance on the ArtEmis and ArtEmis-v2.0 datasets but also exceeded the established benchmarks introduced in This study, lagging behind professional artists' descriptions by a negligible 0.15 points on a 6-point scale. The code and the pre-trained model are accessible in https://huggingface.co/Tyrannosaurus/ArtGPT-4.
翻译:近年来,大型语言模型取得了显著进展,诸如ChatGPT等模型在多样化语言任务中展现了卓越能力。然而,针对拥有数十亿参数的大型模型进行预训练,因其所需的大规模数据集匮乏而面临严峻挑战。尽管如此,创新策略层出不穷,包括利用少量参数集对预训练模型进行微调的方法,如MiniGPT-4和LLaVA等模型所展示的那样。尽管这些模型在多个领域具有潜力,但它们对艺术图像的理解仍存在局限,尚未能完全把握艺术图像的细微之处,也无法像人类感知那样客观阐述其所引发的情绪。本研究提出ArtGPT-4,这是一种开创性的大型视觉语言模型,旨在解决当前模型在艺术理解方面的不足。ArtGPT-4仅需在特斯拉A100设备上训练2小时,使用的数据集包含约52万个图文对。令人印象深刻的是,该模型能够以艺术理解方式渲染图像,并传达其所激发的情感,映射人类解读。此外,本研究还提出一个独特的数据集,用于评估视觉语言模型的有效性。在后续评估中,ArtGPT-4不仅在ArtEmis和ArtEmis-v2.0数据集上取得了最先进性能,还超越了本研究设立的新基准,仅在6分量表上落后专业艺术家描述0.15分。代码与预训练模型可在https://huggingface.co/Tyrannosaurus/ArtGPT-4获取。