The growth of online video platforms drives the need for effective, semantically grounded event retrieval. We present MERVIN, a unified multimodal framework for Vietnamese news videos that integrates keyframes, transcripts, and video summaries. Transcript quality is enhanced via Gemini 1.5 Flash, reducing noise from accents, background sounds, and recognition errors. Visual features are extracted with Perception Encoder, while a Vietnamese language model produces textual embeddings; both are indexed in Milvus for efficient similarity-based retrieval. In addition, a React-based interface enables iterative query refinement across modalities, improving semantic alignment. Experimental results on Vietnamese news videos demonstrate the effectiveness of the proposed system, with MERVIN achieving 79 out of 88 points in AI Challenge HCMC 2025 qualification phase and successfully retrieved all results for every query in the final round.
翻译:在线视频平台的增长推动了对于高效、基于语义的事件检索的需求。我们提出了MERVIN,一个面向越南新闻视频的统一多模态框架,该框架整合了关键帧、转录文本和视频摘要。通过Gemini 1.5 Flash提升转录质量,减少了由口音、背景音和识别错误带来的噪声。视觉特征由感知编码器(Perception Encoder)提取,同时一个越南语语言模型生成文本嵌入;两者均在Milvus中建立索引,以支持基于相似度的高效检索。此外,一个基于React的界面实现了跨模态的迭代式查询优化,改善了语义对齐。在越南新闻视频上的实验结果表明了所提系统的有效性:MERVIN在2025年胡志明市人工智能挑战赛(AI Challenge HCMC 2025)资格赛中获得88分中的79分,并在决赛轮对每个查询均成功检索出全部结果。