Efficient Video Polyp Segmentation by Deformable Alignment and Local Attention.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40711901.
- Also identified by DOI 10.1109/JBHI.2025.3592897.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurate and efficient Video Polyp Segmentation (VPS) is vital for the early detection of colorectal cancer and the effectivetreatment of polyps. However, achieving this remains highly challenging due to the inherent difficulty in modeling the spatial-temporal relationships within colonoscopy videos. Existing methods that directly associate video frames frequently fail to account for variations in polyp or background motion, leading to excessive noise and reduced segmentation accuracy. Conversely, approaches that rely on optical flow models to estimate motion and align frames incur significant computational overhead. To address these limitations, we propose a novel VPS framework, termed Deformable Alignment and Local Attention (DALA). In this framework, we first construct a shared encoder to jointly encode the feature representations of paired video frames. Subsequently, we introduce a Multi-Scale Frame Alignment (MSFA) module based on deformable convolution to estimate the motion between reference and anchor frames. The multi-scale architecture is designed to accommodate the scale variations of polyps arising from differing viewing angles and speeds during colonoscopy. Furthermore, Local Attention (LA) is employed to selectively aggregate the aligned features, yielding more precise spatial-temporal feature representations. Extensive experiments conducted on the challenging SUN-SEG dataset and PolypGen dataset demonstrate that DALA achieves superior performance compared to state-of-the-art models.
Medical subject headings
- Colonic Polyps
- Video Recording
- Image Interpretation, Computer-Assisted