Compliance and factuality of large language models for clinical research document generation.

Wang, Zifeng; Gao, Junyi; Danek, Benjamin; Theodorou, Brandon; Shaik, Ruba; Thati, Shivashankar; Won, Seunghyun; Sun, Jimeng · J Am Med Inform Assoc · 2026

other · Level V

Where this comes from

Abstract

Large language models' (LLMs') performance in high-stakes, compliance-driven settings such as drafting clinical research documents remains underexplored. This study aims to build a benchmark and an evaluation framework for assessing LLMs' compliance and factuality in generating informed consent forms (ICFs) from clinical trial protocols. We introduce InformBench, a benchmark comprising 900 clinical trial documents, and propose an evaluation framework grounded in regulatory guidelines and site-specific consent templates. We assess LLM performance on transforming trial protocols, often hundreds of pages, into concise, patient-facing ICFs. Additionally, we design InformGen, a retrieval-augmented, human-in-the-loop pipeline aimed at improving generation quality. Baseline LLMs such as GPT-4o achieved only 70%-80% compliance and exhibited factual errors in 18%-43% of cases. In contrast, InformGen substantially improved outputs, achieving nearly 100% regulatory compliance and over 90% factual accuracy, as validated by 5 domain-expert annotators. The study reveals critical limitations in current LLMs for clinical research document drafting, particularly in regulatory sensitivity and factual grounding. Our results highlight the need for domain-specific benchmarks and structured evaluations to support safe deployment in real-world clinical research workflows. LLMs offer value in clinical research document generation but must be adapted and rigorously evaluated for high-stakes applications. Our benchmark and framework provide a foundation for improving and assessing LLM-generated outputs in compliance-critical domains.

Medical subject headings