Extraction of Treatments and Responses From Non-Small Cell Lung Cancer Clinical Notes Using Natural Language Processing.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41505664.
- Also identified by DOI 10.1200/CCI-25-00138 and PMC identifier 12788794.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Manual extraction of treatment outcomes from unstructured oncology clinical notes is a significant challenge for real-world evidence (RWE) generation. This study aimed to develop and evaluate a robust natural language processing (NLP) system to automatically extract cancer treatments and their associated RECIST-based response categories (complete response, partial response, stable disease, and progressive disease) from non-small cell lung cancer (NSCLC) clinical notes. This retrospective NLP development and validation study used a corpus of 250 NSCLC oncology notes from University of Pittsburgh Medical Center (UPMC) Hillman Cancer Center, annotated by physician experts. An end-to-end NLP pipeline was designed, integrating a rule-based module for entity extraction (treatments and responses) and a machine learning module using biomedical clinical bidirectional encoder representations from transformers for relation classification. The system's performance was evaluated on a held-out test set, with partial external validation for relation extraction on a Mayo Clinic data set. The NLP system achieved high overall accuracy. On the UPMC test set (64 notes), the relation classification model attained an area under the receiver operating characteristic curve of 0.94 and an F1 score of 0.92 for linking treatments with documented responses. The rule-based entity extraction demonstrated a macro-averaged F1 score of 0.87 (precision 0.98, recall 0.81). Although precision was high for chemotherapy and most response types (1.00), recall for cancer surgery was 0.45. External validation at Mayo Clinic showed moderate relation extraction F1 scores (range: 0.51-0.64). The proposed NLP system can reliably extract structured treatment and response information from unstructured NSCLC oncology notes with high accuracy. This automated approach can assist in abstracting critical cancer treatment outcomes from clinical narrative text, thereby streamlining real-world data analysis and supporting the generation of RWE in oncology.
Medical subject headings
- Natural Language Processing
- Carcinoma, Non-Small-Cell Lung
- Lung Neoplasms
- Electronic Health Records
- Data Mining