Multimodal learning with next-token prediction for large multimodal models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41606344.
- Also identified by DOI 10.1038/s41586-025-10041-x and PMC identifier 12893917.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Developing a unified algorithm that can learn from and generate across modalities such as text, images and video has been a fundamental challenge in artificial intelligence. Although next-token prediction has driven major advances in large language models<sup>1</sup>, its extension to multimodal domains has remained limited, and diffusion models for image and video synthesis<sup>2,3</sup> and compositional frameworks that integrate vision encoders with language models<sup>4</sup> still dominate. Here we introduce Emu3, a family of multimodal models trained solely with next-token prediction. Emu3 equals the performance of well-established task-specific models across both perception and generation, matching flagship systems while removing the need for diffusion or compositional architectures. It further demonstrates coherent, high-fidelity video generation, interleaved vision-language generation and vision-language-action modelling for robotic manipulation. By reducing multimodal learning to unified token prediction, Emu3 establishes a robust foundation for large-scale multimodal modelling and offers a promising route towards unified multimodal intelligence.
Medical subject headings
- Algorithms
- Artificial Intelligence