The Synergy Between Data and Multi-Modal Large Language Models: A Survey From Co-Development Perspective.

Qin, Zhen; Chen, Daoyuan; Zhang, Wenhao; Yao, Liuyi; Huang, Yilun; Ding, Bolin; Li, Yaliang; Deng, Shuiguang · IEEE Trans Pattern Anal Mach Intell · 2025

review · Level V

Where this comes from

Abstract

Recent years have witnessed the rapid development of large language models (LLMs). Multi-modal LLMs (MLLMs) extend modality from text to various domains, attracting widespread attention due to their diverse application scenarios. As LLMs and MLLMs rely on vast amounts of model parameters and data to achieve emergent capabilities, the importance of data is gaining increasing recognition. Reviewing recent data-driven works for MLLMs, we find that the development of models and data is not two separate paths but rather interconnected. Vaster and higher-quality data improve MLLM performance, while MLLMs, in turn, facilitate the development of data. The co-development of multi-modal data and MLLMs requires a clear view of 1) at which development stages of MLLMs specific data-centric approaches can be employed to enhance certain MLLM capabilities, and 2) how MLLMs, using these capabilities, can contribute to multi-modal data in specific roles. To promote data-model co-development for MLLM communities, we systematically review existing works on MLLMs from the data-model co-development perspective.