Exploring efficiency frontiers of thinking budget in medical reasoning: Scaling laws between computational resources and reasoning quality.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41862148.
- Also identified by DOI 10.1016/j.jbi.2026.105025.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To establish a generalizable methodological framework for thinking budget optimization in medical reasoning, deriving predictive scaling laws and operationalizable efficiency regimes that enable principled computational resource allocation in clinical AI deployment. A thinking budget is the number of tokens allocated to intermediate reasoning before answer generation. We evaluated two model families (Qwen3: 1.7B-235B parameters; DeepSeek-R1: 1.5B-70B) on 15 datasets spanning diverse medical specialties and difficulty levels. Cross-architecture validation was achieved through methodological triangulation: Qwen3 used its native API while DeepSeek-R1 employed a truncation method, enabling assessment of framework generalizability across implementation mechanisms. Accuracy was measured across budgets and model scales, with empirical scaling relationships formalized into predictive models. Accuracy exhibited logarithmic scaling with thinking budget (R<sup>2</sup>=0.94), yielding a predictive model for a priori performance estimation. Three operationalizable regimes emerged: high-efficiency (0-256 tokens) for real-time use, balanced (256-512) for routine clinical support, and high-accuracy (>512) for critical diagnostic tasks. The complementarity principle was identified: smaller models (Qwen3:1.7B +22.6%, DeepSeek-R1:1.5B +43.8%) achieved disproportionately larger gains from extended thinking than larger models (Qwen3:235B +3.5%, DeepSeek-R1:70B +13.9%), revealing that reasoning depth can partially compensate for capacity constraints. Domain patterns showed neurology and gastroenterology requiring deeper reasoning than cardiovascular or respiratory medicine. Concordant findings between architecturally distinct implementations (native API vs. truncation control) validate cross-architecture generalizability. This study advances beyond benchmarking to establish thinking budget control as a methodological framework for medical AI optimization, providing predictive models, operationalizable guidelines, and architecture-agnostic principles for dynamic resource allocation aligned with clinical need.