Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture. As AI systems increasingly target manga understanding, OCR, and translation, Manga109 has become a foundational dataset for manga-related AI research. However, the current Manga109 dataset contains inaccurate transcriptions and coarse annotations, which do not align well with modern OCR and multimodal manga understanding tasks. In this work, we revisit the dialogue text annotations of Manga109 and identify five categories of annotation issues, including inaccurate transcriptions, missing text regions, overlapping dialogue and onomatopoeia, and under-segmented speech balloons. To address these issues, we combine OCR-based issue detection and manual revision to construct Manga109-v2026, revising approximately 29,000 dialogue annotations. Our revisions better align Manga109 with modern OCR and multimodal manga understanding systems while preserving expressive structures characteristic of manga.
翻译:漫画是一种具有文化特色的多模态媒介,也是日本流行文化中最具影响力的形式之一。随着人工智能系统日益聚焦于漫画理解、OCR(光学字符识别)与翻译任务,Manga109已成为漫画相关AI研究的基础数据集。然而,当前的Manga109数据集存在转录不准确与标注粗糙的问题,难以与现代OCR及多模态漫画理解任务良好适配。本文重新审视了Manga109中的对话文本标注,识别出五类标注问题,包括转录不准确、文本区域缺失、对话与拟声词重叠,以及语音气泡分割不足。为解决这些问题,我们结合基于OCR的缺陷检测与人工修正,构建了Manga109-v2026,修订了约29,000条对话标注。我们的修订使Manga109与现代OCR及多模态漫画理解系统更好地对齐,同时保留了漫画特有的表达结构。