Single-Stage Instance Segmentation Survey: A "1+2+3" Technical Framework.
Where this comes from
- Record sourced from PubMed, PMID 42262955.
- Also identified by DOI 10.1109/TPAMI.2026.3701328.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Single-stage instance segmentation is a prominent research direction in computer vision. However, in existing single-stage instance segmentation methods, there is a lack of summary on the 6 key questions: "How to select the learning paradigm?", "How to extract features?", "How to fuse features?", "How to classify objects?", "How to generate object masks?", and "How to localize objects?". These questions restrict the development and application to some degree. To address these questions, this paper proposes a "1+2+3" technical framework for summarizing single-stage instance segmentation. "1" represents one learning paradigm. Four mainstream learning paradigms and selection strategies are systematically summarized. "2" represents two local structures. The three characteristics of feature extraction are summarized, and the two feature fusion strategies are summarized. "3" represents three aspects of instance prediction. The object classification, object mask generation methods and object localization are discussed. Furthermore, the paper discusses new technologies such as instance segmentation based on promptable foundation models and instance segmentation based on vision-language models, as well as their typical applications in medical, video, and remote sensing image segmentation. The "1+2+3" technology framework proposed in the paper provides a problem-oriented technical map for instance segmentation, promoting the further development of single-stage instance segmentation.