SurgflowNet: Leveraging unannotated video for consistent endoscopic pituitary surgery workflow recognition.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41330256.
- Also identified by DOI 10.1016/j.artmed.2025.103309.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Surgical workflow recognition has the potential to accelerate training initiatives through the analysis of surgical videos, improve intraoperative efficiency, and support preemptive postoperative care. Unlike well-explored minimally invasive surgeries, where surgical workflows are consistent across patients, automating endoscopic pituitary surgery workflow recognition is challenging. Pituitary surgery involves a large number of steps, diverse sequences, optional steps, and frequent transitions, making it challenging for current state-of-the-art (SOTA) methods, which struggle with transferability. Progress is largely limited by the lack of annotated data that captures the complexity of pituitary surgery, and obtaining such annotations is both time-consuming and resource-intensive. This paper presents SurgflowNet, a novel spatio-temporal model for consistent pituitary workflow recognition leveraging unannotated data. We utilise a limited yet fully annotated dataset to infer quasi-labels for unannotated videos and curate a balanced dataset to train a robust frame encoder using the student-teacher framework. A spatio-temporal network that combines the resulting frame encoder and an LSTM network is trained with a consistency loss to ensure stability in step predictions. With a 5% improvement in macro F<sub>1</sub>-score and 13.4% in Edit Score over the SOTA, SurgflowNetdemonstrates a significant improvement in workflow recognition for endoscopic pituitary surgery.
Medical subject headings
- Workflow
- Video Recording
- Pituitary Gland
- Endoscopy