Advancing the science of qualitative patient preference assessment using large language models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41818287.
- Also identified by DOI 10.1371/journal.pdig.0001263 and PMC identifier 12981457.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Patient experiences and perspectives are essential for shaping patient-centered healthcare. While large language models (LLMs) in healthcare are typically applied to specific clinical or patient-facing tasks, they have not been used for qualitative patient preference assessment, which often relies on thematic analysis to understand patient views expressed in interviews or focus groups. LLMs show initial promise for performing inductive thematic analysis of healthcare interview or focus group transcripts, yet no empirical studies have investigated LLMs to facilitate qualitative patient preference assessment. We employed the open-source Hermes-3-Llama-3.1-70B LLM to perform inductive thematic analysis on focus group transcripts from a previously published qualitative patient preference assessment study using three optimized prompt frameworks, and evaluated semantic similarity of LLM generated themes against human-analyzed themes using the Sentence-T5-XXL language embedding model. Sentence-level theme similarity was assessed using Jaccard similarity coefficients (0-1 range), computing coefficient scores across a broad range of discrete cosine similarity thresholds. We further evaluated LLM themes for similarity in lexical diversity and reading grade-level metrics and benchmarked semantic similarity results with published similarity thresholds previously used with qualitative healthcare data. All prompt frameworks generated themes with median Jaccard similarity coefficients with human-analyzed themes between 0.46-0.64, indicating moderate semantic overlap. Our best-performing framework instructed to pursue thematic saturation scored closest to human-analyzed themes on all reading grade-level metrics, and demonstrated 12% higher semantic overlap with human-analyzed themes compared to published benchmarks. Our worst-performing framework produced themes with moderate semantic overlap and hallucinated findings unidentified in human-analyzed themes. We demonstrate that LLMs can perform inductive thematic analysis of qualitative patient preference data, producing themes substantively similar in content and style to human-analyzed themes when augmented with sufficient domain-specific context. While LLMs may augment thematic analysis, the contextual nature of qualitative analysis remains a challenge requiring collaborative LLM frameworks integrating human expertise.