Assessment of LLMs for Generating Synthetic Health-Related Quality-of-Life Survey Responses Among U.S. Adults.
Where this comes from
- Record sourced from PubMed, PMID 42225193.
- Also identified by DOI 10.1016/j.amepre.2026.108454.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large Language Models (LLMs) present a potential alternative to traditional surveys by generating synthetic health-related quality-of-life (HRQoL) survey responses, offering a faster and more cost-effective approach to population health monitoring. This study evaluated the accuracy of LLM-generated predictions of population-level HRQoL survey response. The primary outcomes were the Physical Component Summary (PCS) and Mental Component Score (MCS) scores, calculated from responses to the 12-item Short Form Survey (SF-12). LLaMA 4 was used to generate synthetic item-level SF-12 responses; PCS and MCS scores were then derived using the standard SF-12 scoring algorithm. Prompts were constructed in four scenarios: no individual-level characteristics, demographic characteristics, demographic and socioeconomic characteristics, and demographic, socioeconomic, and health-related characteristics. Data from the 2022 Medical Expenditure Panel Survey were analyzed in 2025. LLM performance in estimating PCS was poor when no individual-level data were provided (mean absolute percentage error [MAPE]: 11.10; root mean square error [RMSE]: 156.53; R²: 0.00; prediction-to-observation ratio: 0.97; Pearson correlation: 0.02). Accuracy improved substantially with the addition of individual-level information, achieving the best performance when demographic, socioeconomic, and health-related characteristics were all incorporated (MAPE: 7.91; RMSE: 109.17; R²: 0.26; prediction-to-observation ratio: 1.00; Pearson correlation: 0.51). However, limitations persisted, particularly in accurately capturing values at the distributional extremes. For MCS, inclusion of individual characteristics led to only modest improvements. LLMs showed potential to generate synthetic SF-12 response patterns, particularly for physical HRQoL; however, limitations in individual-level accuracy and performance at the distributional extremes underscore the need for further methodological refinement.