Critique of impure reason: Unveiling the reasoning behaviour of medical large language models.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 41150728.
- Also identified by DOI 10.7554/eLife.106187 and PMC identifier 12563542.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Despite the current ubiquity of large language models (LLMs) across the medical domain, there is a surprising lack of studies which address their <i>reasoning behaviour</i>. We emphasise the importance of understanding <i>reasoning behaviour</i> as opposed to high-level prediction accuracies, since it is equivalent to explainable AI (XAI) in this context. In particular, achieving XAI in medical LLMs used in the clinical domain will have a significant impact across the healthcare sector. Therefore, in this work, we adapt the existing concept of <i>reasoning behaviour</i> and articulate its interpretation within the specific context of medical LLMs. We survey and categorise current state-of-the-art approaches for modelling and evaluating <i>reasoning</i> in medical LLMs. Additionally, we propose theoretical frameworks which can empower medical professionals or machine learning engineers to gain insight into the low-level reasoning operations of these previously obscure models. We also outline key open challenges facing the development of <i>large reasoning models</i>. The subsequent increased transparency and trust in medical machine learning models by clinicians as well as patients will accelerate the integration, application as well as further development of medical AI for the healthcare system as a whole.
Medical subject headings
- Machine Learning
- Language
- Models, Theoretical