Generalized post-training quantization for medical image segmentation foundation model.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42705143.
- Also identified by DOI 10.1016/j.media.2026.104287.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Medical image segmentation foundation models (MedFMs) perform strongly across diverse imaging modalities, but their large size and computational demands hinder deployment in resource-limited clinical settings. Lightweight fine-tuning is impractical given the high training cost of MedFMs, and the efficiency benefits of integer inference remain underused. Post-training quantization (PTQ) is a promising alternative, yet existing PTQ methods fail on MedFMs because their heterogeneous weight distributions lead to severe accuracy degradation at low bit-widths. To address this challenge, we propose GPTQ-MedFM, a generalized post-training quantization framework tailored for medical foundation models. GPTQ-MedFM standardizes complex weight distributions under an ℓ<sub>∞</sub>-constrained normalization to produce quantization-friendly matrices, and then applies an efficient coordinate-descent solver to obtain high-fidelity low-bit representations. The method adds no extra computation or memory overhead at inference, enabling seamless deployment on medical edge devices. By explicitly modeling and mitigating quantization-induced errors, GPTQ-MedFM achieves state-of-the-art low-bit compression across six medical foundation models and nine imaging modalities - spanning nearly the full range of clinical imaging scenarios - and remains robust even with a single calibration sample. Its broad generalization and minimal calibration cost make GPTQ-MedFM a practical pathway for real-time AI-assisted diagnostics in resource-limited healthcare settings.