Large language models require a new form of oversight: capability-based monitoring.
Where this comes from
- Record sourced from PubMed, PMID 42141031.
- Also identified by DOI 10.1038/s41746-026-02740-0 and PMC identifier 13179377.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Large language models (LLMs) have been rapidly adopted in healthcare, but oversight strategies are lacking. We propose capability-based monitoring, motivated by the fact that LLMs are generalist systems whose overlapping internal capabilities are reused across numerous downstream tasks. This approach organizes monitoring around shared capabilities to enable cross-task detection of systemic weaknesses, long-tail errors, and emergent behaviors. We describe considerations for developers, organizational leaders, and professional societies. and policymakers.