TDFormer: Top-Down Token Generation for 3D Medical Image Segmentation.

Du, Hao; Dong, Qihua; Xu, Yan; Liao, Jing · IEEE J Biomed Health Inform · 2025

basic_science · Level V

Where this comes from

Abstract

Accurate medical image segmentation is critical to effective treatment strategies. Existing transformer-based methods for image segmentation mostly split the input image into a fixed and regular grid and regard cells in the grid as the vision tokens. However, not all tokens are of equal importance in the medical segmentation tasks, e.g., the tokens in tumor areas must be processed in a higher resolution than the background tokens which can be easily predicted with fewer transformer layers. In this paper, we propose a simple yet efficient segmentation framework called Top-Down Transformer (TDFormer), which incorporates a spatially adaptive token generation scheme into the transformer. The proposed top-down token generation comprises the following three components: attentiveness calculation, token splitting, and token fusion, where the collaboration of these components gradually fuses redundant background tokens and focuses only on the most critical areas. This allows for allocating more computation to process tokens containing delicate details in a finer resolution. Extensive experiments are conducted to demonstrate the robustness and effectiveness of the proposed TDFormer, that our method are superior to other state-of-the-art methods on the following publicly accessible datasets: BTCV Challenge, LiTS and BraTS 2020. We also dissect our method and evaluate the performance of each component.

Medical subject headings