Characterizing LLM scientific concept generation: A multi-dimensional measurement study.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42743252.
- Also identified by DOI 10.1371/journal.pone.0357892.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large Language Models (LLMs) can generate text describing scientific concepts, but the characteristics of these outputs remain poorly understood. We present a multi-dimensional characterization framework analyzing 9,666 outputs from seven models (GPT-4.1, GPT-5.2, GPT-5.5, o4-mini, Claude Sonnet 4.5, Claude Opus 4.5, and the open-weight Gemma-3-27B) across five scientific domains. Rather than making claims about creativity or novelty, we measure five independent dimensions: coherence, domain relevance, lexical profile, structural properties, and semantic position using four sentence embedding models spanning 2020-2024. After quality filtering (99.2% coherence, 99.9% domain relevance pass rates), we find that outputs exhibit graduate-level readability (median Flesch-Kincaid grade 16.3) and occupy semantic positions at the 83rd percentile of calibration distributions. All 39 metrics differ significantly across models (Kruskal-Wallis, p < 0.05), with structural properties showing the largest effects (ϵ2 = 0.35-0.54) and semantic position showing small-to-medium effects (ϵ2 ≈ 0.01-0.14). Dunn's post-hoc tests with Bonferroni correction confirm that all model pairs differ on the top structural metrics (21/21 pairs for paragraph count, 20/21 for the next three). Centroid positions are robust to calibration sampling (bootstrap cosine similarity ≥ 0.99) and exceed a shuffled-domain baseline in 96.8-98.7% of cases; cross-model embedding consistency is moderate (Spearman ρ = 0.38-0.67) and is not explained by embedding dimensionality. Temperature effects replicate across two independent full-range models (GPT-4.1 and Gemma). This work provides calibrated measurements and validated methodology for future research without making interpretive claims about novelty.
Medical subject headings
- Large Language Models