This chapter explores the evolution, classification, and health implications of food processing, while emphasizing the transformative role of machine learning, artificial intelligence (AI), and data science in advancing food informatics. It begins with a historical overview and a critical review of traditional classification frameworks such as NOVA, Nutri-Score, and SIGA, highlighting their strengths and limitations, particularly the subjectivity and reproducibility challenges that hinder epidemiological research and public policy. To address these issues, the chapter presents novel computational approaches, including FoodProX, a random forest model trained on nutrient composition data to infer processing levels and generate a continuous FPro score. It also explores how large language models like BERT and BioBERT can semantically embed food descriptions and ingredient lists for predictive tasks, even in the presence of missing data. A key contribution of the chapter is a novel case study using the Open Food Facts database, showcasing how multimodal AI models can integrate structured and unstructured data to classify foods at scale, offering a new paradigm for food processing assessment in public health and research.
翻译:本章探讨了食品加工的演变、分类及其对健康的影响,同时强调了机器学习、人工智能(AI)和数据科学在推动食品信息学变革中的关键作用。首先回顾了食品加工的历史,并对NOVA、Nutri-Score和SIGA等传统分类框架进行了批判性分析,指出了它们的优势与局限,尤其是主观性和可重复性挑战如何制约了流行病学研究与公共政策制定。针对这些问题,本章提出了新颖的计算方法,包括FoodProX——一种基于营养成分数据训练的随机森林模型,用于推断加工水平并生成连续的FPro评分。此外,还探讨了BERT和BioBERT等大型语言模型如何在存在缺失数据的情况下,对食品描述和配料表进行语义嵌入以完成预测任务。本章的核心贡献在于一项基于Open Food Facts数据库的案例研究,展示了多模态AI模型如何整合结构化与非结构化数据,大规模分类食品,为公共卫生与研究中的食品加工评估提供了全新范式。