Accurate dietary assessment is critical for precision nutrition, yet most image-based methods rely on a single pre-consumption image and provide only coarse, meal-level estimates. These approaches cannot determine what was actually consumed and often require restrictive inputs such as depth sensing, multi-view imagery, or explicit segmentation. In this paper, we propose a simple vision-language framework for food-item-level nutritional analysis using paired before-and-after eating images. Instead of relying on rigid segmentation masks, our method leverages natural language prompts to localize specific food items and estimate their weight directly from a single RGB image. We further estimate food consumption by predicting weight differences between paired images using a two-stage training strategy. We evaluate our method on three publicly available datasets and demonstrate consistent improvements over existing approaches, establishing a strong baseline for before-and-after dietary image analysis.
翻译:精确的膳食评估对于精准营养至关重要,然而大多数基于图像的方法仅依赖单张餐前图像,且只能提供粗略的餐级估计。这些方法无法判断实际摄入量,且通常需要深度传感、多视角成像或显式分割等限制性输入。本文提出了一种简单的视觉-语言框架,利用成对的餐前餐后图像实现食品级的营养分析。该方法无需依赖刚性分割掩码,而是通过自然语言提示定位特定食品,并直接从单张RGB图像中估算其重量。我们进一步采用两阶段训练策略,通过预测成对图像间的重量差异来估算食品摄入量。在三个公开数据集上的评估表明,本方法相较于现有方法取得了一致性改进,为餐前餐后膳食图像分析建立了强基线。