We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
翻译:我们提出Omni,一种统一的多模态模型,原生训练于多种模态,包括文本、图像、视频、3D几何和隐式表示。我们发现此类训练能够实现上下文展开,即模型在生成预测前显式地跨多种模态表示进行推理。该过程使模型能够整合异构模态中的互补信息,促进对共享多模态知识流形的更忠实逼近,并提升下游推理的准确性。因此,Omni在多模态生成与理解基准测试中均取得了强劲性能,同时展现出先进的多模态推理能力,包括文本、图像、视频和3D几何的上下文生成。