AI Writing Vocabulary in Ophthalmology: Faster Non-Native Adoption, Unchanged Outcomes.

Berni, Alessandro; Ting, Daniel Shu Wei; Caputo, Greta; Russo, Alessandro; Sinisterra, Laura Gutierrez; Avitabile, Alessandro; Virgili, Gianni; Reibaldi, Michele et al. · Am J Ophthalmol · 2026

cross_sectional · Level IV

Where this comes from

Abstract

To quantify the change in large language model (LLM)-associated writing vocabulary in ophthalmology after ChatGPT, to test whether it differed by first-author affiliation-country language group, and to determine whether publishing outcomes shifted. Retrospective, cross-sectional bibliometric study, with an interrupted time-series analysis of a 10-year publication census. Published articles, not human subjects. Primary corpus, 15,683 PubMed abstracts from the 40 highest-impact ophthalmology journals (top 10 per 2025 Journal Citation Reports quartile), pre-ChatGPT (2018-2019) versus post-ChatGPT (2023-2024); supportive corpus, 6,139 open-access full texts; and a 48,468-article census (2015-2024) for outcomes. Articles were grouped by whether the first author's affiliation country was native-English-speaking. We counted 41 curated LLM-associated excess words per document, with the word count as a Poisson offset and frequency-common control words for specificity. The prespecified primary analysis was the period-by-language-group interaction in a Poisson generalized estimating equation clustered on first author. The instrument was construct-validated against 30 LLM-generated abstracts and against pre-ChatGPT, definitionally LLM-free, human abstracts. The period-by-language-group interaction incidence rate ratio (IRR) for the excess-word rate; and the non-native share of publications and of top-quartile journal placements. The excess-word rate rose 2.1-fold, from 320 to 668 per million words. The rise was steeper for non-native-English-affiliated first authors (interaction IRR, 0.61; 95% CI, 0.48-0.78; P < .001), robust to adjustment for journal quartile and country income, to excluding China (IRR, 0.65), and in an independent full-text corpus (IRR, 0.64). The counter separated LLM-generated from human abstracts (area under the receiver operating characteristic curve [AUC], 0.88), whereas per-article discrimination was near chance (AUC, 0.53), confirming a population-level rate rather than a classifier. An interrupted time series showed no post-ChatGPT change in the non-native share of publications or of top-quartile journals. After ChatGPT, non-native-English-affiliated authors in ophthalmology adopted AI-associated writing vocabulary faster than native-affiliated authors, without any accompanying gain in publication frequency or journal placement. The measure reflects population-level AI-associated style, not fluency, quality, or confirmed AI use, and should not be used to classify individual articles.