OsteoFusionFormer: dual-stage transformer fusion framework for knee osteoporosis diagnosis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41865617.
- Also identified by DOI 10.1016/j.knee.2026.104428.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The aim of the study is to introduce a new model, OsteoFusionFormer, namely a dual transformer model for automatic classification of knee osteoporosis into three groups: Normal, Osteopenia, and Osteoporosis. The objective was to overcome single-branch transformer limitations by incorporating anatomical global context and fine-grained bone features to increase diagnostic accuracy. OsteoFusionFormer combines two parallel arms, a Vision Transformer (ViT) for global anatomical representation and a Bone-Aware Transformer (BAT) for localised bone-specific features. These are combined with a hierarchical dual fusion strategy. First, cross-attention enables feature-level fusion between ViT and BAT embeddings. Next, a confidence-weighted decision-level fusion employs an auxiliary gating network to compute adaptive weights (α<sub>1</sub>, α<sub>2</sub>), yielding a soft-ensemble prediction. Ablation studies systematically remove modules (ViT, BAT, gating, CLAHE) to assess contributions. Interpretability is assessed via attention maps and Grad-CAM++. OsteoFusionFormer had 96.8% overall accuracy, which exceeds the accuracy of ViT-only (91.3%), BAT-only (90.1%), and late fusion averaging (93.6%). Ablation verified a drop in performance without BAT (-5.5%), ViT (-6.7%), gating (-3.2%), and CLAHE (-4.4%). Performance was verified with 15 new kinds of bone-specific indicators: Bone-Aware Accuracy: 96.8%, Trabecular Sensitivity Index: 95.2%, Cortical Degeneration Detection Rate: 97.6%, Joint Space Narrowing Recall: 96.1%, Bone Class Specificity: 97.2%, Osteopenia Detection Precision: 92.4%, Bone Focus Ratio: 91.8%, Bone Entropy Index: 0.26 bits, Visual interpretability showed good expert agreement (BIAS: 87.3%). By combining global and local bone features, OsteoFusionFormer provides better accuracy, diagnosis sensitivity, and structure focus with an explainability guarantee.
Anatomy
- knee