Evaluating the Potential of Reasoning Large Language Models to Perpetuate Racial and Gender Disease Stereotypes in Health Care.

Docking, Joshua J; Li, Lee X; Menz, Bradley D; Bacchi, Stephen; Hopkins, Ashley M; Sorich, Michael J · J Med Internet Res · 2026

Where this comes from

Abstract

This evaluation of 36,000 clinical vignettes found that next-generation reasoning large language models, o3-mini and DeepSeek-R1, frequently perpetuate racial and gender stereotypes for common medical conditions, indicating that advancements in reasoning do not inherently improve representational fairness.

Medical subject headings