Language models (LMs) are no longer restricted to ML community, and instruction-tuned LMs have led to a rise in autonomous AI agents. As the accessibility of LMs grows, it is imperative that an understanding of their capabilities, intended usage, and development cycle also improves. Model cards are a popular practice for documenting detailed information about an ML model. To automate model card generation, we introduce a dataset of 500 question-answer pairs for 25 ML models that cover crucial aspects of the model, such as its training configurations, datasets, biases, architecture details, and training resources. We employ annotators to extract the answers from the original paper. Further, we explore the capabilities of LMs in generating model cards by answering questions. Our initial experiments with ChatGPT-3.5, LLaMa, and Galactica showcase a significant gap in the understanding of research papers by these aforementioned LMs as well as generating factual textual responses. We posit that our dataset can be used to train models to automate the generation of model cards from paper text and reduce human effort in the model card curation process. The complete dataset is available on https://osf.io/hqt7p/?view_only=3b9114e3904c4443bcd9f5c270158d37
翻译:语言模型(LMs)已不再局限于机器学习社区,经过指令微调的语言模型更推动了自主AI智能体的兴起。随着语言模型可及性的提升,对其能力、预期用途及开发周期的理解也必须同步加强。模型卡片是记录机器学习模型详细信息的常用实践。为实现模型卡片自动生成,我们针对25个机器学习模型构建了一个包含500个问答对的数据集,内容涵盖模型的关键方面,如训练配置、数据集、偏差、架构细节及训练资源。我们聘请标注员从原始论文中提取答案。此外,我们探索了语言模型通过回答问题生成模型卡片的能力。我们使用ChatGPT-3.5、LLaMa和Galactica进行的初步实验显示,上述语言模型在理解研究论文及生成事实性文本响应方面存在显著差距。我们提出,本数据集可用于训练模型,使其能够从论文文本中自动生成模型卡片,从而减少人工整理模型卡片的工作量。完整数据集可在https://osf.io/hqt7p/?view_only=3b9114e3904c4443bcd9f5c270158d37获取。