End to end AI system for surgical gesture sequence recognition and clinical outcome prediction.
other
Where this comes from
- Record sourced from PubMed, PMID 42337001.
- Also identified by DOI 10.1038/s41746-026-02927-5.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remains a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (˜2 s) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features-gesture frequency, duration, and transitions-predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (∆d<sub>avg</sub> ≈ 0.07), and strong correlation (r = 0.96, p < 1 × 10<sup>-14</sup>). F2O also captured key patterns linked to erectile function recovery, including prolonged tissue peeling and reduced energy use. By enabling automatic interpretable assessment, F2O establishes a foundation for data-driven surgical feedback and prospective clinical decision support.