Validity evidence for patient-centered communication behavior descriptions developed through an LLM-assisted workflow: SME and lay patient perspectives.
Where this comes from
- Record sourced from PubMed, PMID 42715626.
- Also identified by DOI 10.1016/j.pec.2026.109841.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Effective physician-patient communication is essential for patient outcomes, yet translating communication frameworks into actionable feedback remains challenging. LLM-assisted workflows may support scalable development of patient-centered communication behavior descriptions (PCBs), but their alignment with expert and patient perspectives is unclear. This study gathered validity evidence for PCBs developed using an LLM-assisted workflow through review by subject matter experts (SMEs) and lay patients. We examined PCB appropriateness for feedback and importance for physician-patient communication, identified additional behaviors not initially included in the AI-assisted set, compared SME and lay patient evaluations, and explored sample size considerations for human-in-the-loop review. We recruited 33 SMEs and 60 lay patients to evaluate 25 PCBs developed through an LLM-assisted workflow across six vignettes. Participants rated each PCB for appropriateness and importance and suggested additional behaviors. We analyzed appropriateness, importance and open-ended responses. Chi-squared tests compared appropriateness across groups, Pearson correlation assessed alignment in importance ratings, and bootstrapping estimated sampling precision by rater size. SMEs and lay patients endorsed the PCBs as appropriate at high rates, with mean PCB-level percentages of 90.18% and 93.40%, respectively. Open-ended responses revealed few additional behaviors not initially included in the AI-assisted set: most suggestions were low frequency, but seven behaviors met a conservative 10% threshold, representing potential areas for review. Chi-squared tests indicated group alignment in appropriateness, and importance ratings showed strong SME-lay patient concordance (r = .85). Bootstrapping indicated higher SEMs for lay patients, although larger lay patient samples may offset this difference. An LLM-assisted workflow generated PCBs broadly endorsed by SMEs and lay patients as appropriate for learner feedback. Reviewing additional behaviors may enhance comprehensiveness. Findings provide guidance for sample size decisions and scalable human-in-the-loop review strategies balancing precision and feasibility. Results support AI-assisted content development that reduces repeated human content generation while preserving review.