CNER-Omni: A unified dynamic modality learning framework for Chinese named entity recognition across text and speech.

Ning, Jinzhong; Mu, Wenxuan; Li, Songtao; Zhang, Yijia; Luo, Ling; Sun, Yuanyuan; Lu, Mingyu; Lin, Hongfei · Neural Netw · 2026

Where this comes from

Abstract

With the proliferation of multimodal data in real-world applications, Chinese Named Entity Recognition (CNER) has extended beyond textual inputs to incorporate speech and speech-text multimodal scenarios. However, existing methods often treat text-based, speech-based, and multimodal NER as independent tasks, leading to redundant architectures and limited cross-modal generalization. In this paper, we propose CNER-Omni, a unified framework for Integrated Multimodal Named Entity Recognition (IMNER) that consolidates the three subtasks into a single model. To this end, we design a unified input representation schema and introduce IMAGE, an Integrated Multimodal Generation framework that formulates NER as an entity-aware sequence generation task. IMAGE leverages pseudo-modal inputs, where one modality is synthetically generated to complement the other, enabling effective cross-modal representation learning. In addition, a modality-composition-aware mixture-of-experts (MoE) module is introduced to dynamically adapt the model to various input configurations. Extensive experiments on AISHELL-NER, CNERTA, and MSRA benchmarks demonstrate that IMAGE achieves state-of-the-art performance across all three modalities while significantly reducing model complexity. Notably, our approach excels in both flat and nested NER settings and exhibits strong robustness in low-resource and cross-modal scenarios. Code is publicly available at: https://github.com/NingJinzhong/IMAGE4IMNER.

Medical subject headings