Efficient information extraction using LLMs and knowledge distillation: A study on HPV health communication.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 41805771.
- Also identified by DOI 10.1371/journal.pdig.0001275 and PMC identifier 12974803.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
State Department of Health (DOH) websites serve as authoritative sources of HPV-related health communications, presenting state-specific content that influences public awareness and vaccination decisions. We develop a computationally efficient framework to systematically evaluate these information repositories based on their content quality, completeness, and their motivational impact on vaccination behavior. We propose a dataset consolidating 48 different DOH websites' data targeted towards HPV and HPV vaccination. By developing an annotated dataset (n = 400), efficient prompting techniques and a Knowledge Distillation framework, we develop and evaluate efficient student models based on the Llama family of Large Language Models (LLMs) and the RoBERTa Large encoder architecture. We finally deploy the best-performing student model for a computationally feasible evaluation of the content of DOH websites. We show that fine-tuned RoBERTa Large model achieves an F1 score of 0.74 on the test set, outperforming all other student models and approaching the teacher model's performance (F1 = 0.77). The fine-tuned RoBERTa-Large model is subsequently applied to data from various state DOH websites to evaluate the information presented. We also discuss the broader implications, limitations, and ethical and legal considerations of the proposed approach.