LLM-assisted clinical coding audit through an interpretable coding pipeline.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42570566.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106605.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Clinical coding is vital yet complex, often hindered by imperfect training data. This study addresses the overlooked issue of undercoding and coding errors in standard datasets and investigates their impact on automated coding algorithms. Furthermore, it explores the potential of AI-assisted tools to facilitate the challenging process of clinical coding audit. We developed a novel and effective interpretable coding pipeline that integrates Large Language Models (LLMs) for evidence extraction and code verification with a multiclass classifier trained on a large-scale silver-standard evidence dataset for code prediction. Using this pipeline as an audit tool, three professional coders systematically identified, categorised, and corrected errors in two widely used datasets containing human-annotated evidence (MDACE and CodiEsp). The audit uncovered significant data quality issues, including an estimated 76.3% undercoding rate in MDACE and a 29.7% error rate in CodiEsp. Re-evaluating existing models on the corrected datasets improved performance, consistently across all metrics on MDACE, and in recall on CodiEsp. Additionally, the proposed pipeline achieved superior or comparable performance to state-of-the-art LLM-based methods, demonstrating its viability as a high-potential dual-use coding framework. Models and results that are publishable in compliance with the data privacy requirements of the datasets have been made publicly accessible in the acocmpanied Github repository: https://github.com/Supriya090/LLM-Clinical-Coding-Audit. Our findings demonstrate that substantial errors in benchmark datasets significantly impact model evaluation. The results underscore a critical need to shift the research focus from purely model-centric approaches to data-centric solutions in clinical AI, highlighting the effectiveness of AI-assisted audits in improving data integrity.