Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound: a consensus framework.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42680190.
- Also identified by DOI 10.1111/anae.70364.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon. We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance. The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p < 0.001). The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.