Large Language Models Improve Operative Note Coding Accuracy and Financial Outcomes in Neurotology.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42532681.
- Also identified by DOI 10.1002/lary.70777.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To compare coding accuracy and financial impact between an institution's large language model (LLM) and centralized human coders for neurotology operative notes. This retrospective cohort study reviewed 124 consecutive operative notes performed between July 1, 2024, and June 30, 2025, from neurotology attendings at a tertiary academic medical center. Each note was independently coded by the institution's LLM and the centralized coding team. A surgeon-adjudicated reference standard was established through blinded surgeon review and validated by a second blinded neurotologist (Cohen's κ = 0.88). Primary outcomes were coding accuracy and financial variance measured in relative value units (RVUs). The LLM achieved significantly higher coding accuracy than human coders (86.3% vs. 49.2%; p = 1.04 × 10<sup>-11</sup>, McNemar exact test). Human coders demonstrated a mean negative RVU variance of -5.04 (SD 10.79), indicating systematic under-coding, compared with a mean LLM variance of +0.93 (SD 4.85; p < 0.01). Relative human under-coding averaged 17.3% of reference standard RVUs and scaled with procedural complexity. All 61 human errors involved missing or incorrect codes, whereas LLM errors were split between missing and extraneous codes. Extrapolated to the 5-surgeon division, human under-coding projected an annual loss of 1950 RVUs (-$145,342). The LLM demonstrated significantly higher concordance with the surgeon-adjudicated reference standard than centralized human coders by 37.1 percentage points. Human errors were driven by under-coding of complex procedures, resulting in substantial projected revenue loss. A hybrid model using LLM-generated drafts verified by specialty-trained coders may optimize coding accuracy and revenue integrity for subspecialty surgical practices. N/A.