Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42757812.
- Also identified by DOI 10.1097/XCS.0000000000002211.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
National Surgical Quality Improvement Program (NSQIP) data collection depends on labor-intensive manual chart abstraction, limiting efficiency, increasing cost, and necessitating patient sampling. This study evaluated whether a large language model (LLM) could accurately abstract unstructured NSQIP breast reconstruction variables compared with conventional human abstraction. Clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024-February 28, 2025) were manually de-identified and processed using a customized ChatGPT 4.1 workflow targeting individual variables. A faculty plastic surgeon established the reference standard. Overall accuracy of LLM and human abstraction was compared using McNemar's and Chi-square tests. Among 105 patients (73 bilateral, 32 unilateral), 9,048 data points were evaluated. Overall abstraction accuracy was 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction (McNemar p<0.001; Chi-square p<0.001). LLM performance exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables. The most frequent LLM error involved prior breast surgical history (29/61 errors), followed by prepectoral versus subpectoral implant or expander placement, a variable frequently requiring inference from documentation. In this proof-of-concept validation study, a customized LLM achieved significantly higher abstraction accuracy than conventional human review for general and breast reconstruction NSQIP variables. These findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, although multicenter validation remains necessary.