MSB-2.5DUNet++: A Multi-Scale Bottleneck Enhanced 2.5D UNet++ Framework with EfficientNet Encoder for Explainable Brain Tumor Segmentation
DOI:
https://doi.org/10.47981/j.mijst.14(01)2026.593(85-97)Keywords:
Brain Tumor Segmentation, UNET++, Explainable AI (XAI), GradcamAbstract
The segmentation of brain tumors in MRI images still presents challenges because of the heterogeneous appearance of brain tumors, unclear tumor boundaries, and significant class imbalance, which can lead to less accurate and less interpretable automated brain tumor segmentation methods. To overcome these challenges, this study presents an accurate and reproducible framework for brain tumor segmentation in the BRISC 2025 dataset using a modified 2.5D UNet++ model with a multi-scale bottleneck. This study used an EfficientNet-B4 model for brain tumor segmentation in five adjacent axial brain MRI slices at 512 × 512 resolution. This approach enables the model to leverage more contextual information across five slices rather than a single slice. To improve the model's robustness, Kornia-based data augmentation, mixed precision, and exponential moving averages are used. To mitigate the effects of severe class imbalance in brain tumor segmentation, the Dice and Binary Cross-Entropy loss functions are combined. To improve the reliability of brain tumor segmentation, the model uses flip-based augmentation during testing, validation-based threshold sweeping to determine the optimal threshold, and minimum-area filtering to eliminate false positives. In addition, the study used LayerCAM for brain tumor segmentation to improve model interpretability and to develop a CAM-rescue approach for handling empty or very small prediction problems. This study demonstrates the effectiveness of the proposed framework for brain tumor segmentation, achieving Dice scores of 0.87 and 0.862 in validation and testing, respectively, and high AP and AUC values in validation (AP: 0.946, AUC: 0.996) and testing (AP: 0.947, AUC: 0.995). This demonstrates the 2.5D UNet++ model's competitiveness with post-processing for brain tumor segmentation.
Downloads
References
Ahmed, M. M., Hossain, M. M., Islam, M. R., Ali, M. S., Nafi, A. A. N., Ahmed, M. F., Ahmed, K. M., Miah, M. S., Rahman, M. M., Niu, M., & Islam, M. K. (2024). Brain tumor detection and classification in MRI using hybrid ViT and GRU model with explainable AI in Southern Bangladesh. Scientific Reports, 14, Article 22797.
https://doi.org/10.1038/s41598-024-71893-3
BRISC 2025 Challenge Dataset. (2025). BRISC 2025: Brain tumor MRI dataset for segmentation and classification [Data set]. Kaggle.
https://doi.org/10.34740/kaggle/ds/7632487
Cai, Y., Long, Y., Han, Z., Liu, M., Zheng, Y., Yang, W., & Chen, L. (2023). Swin Unet3D: A three-dimensional medical image segmentation network combining vision transformer and convolution. BMC Medical Informatics and Decision Making, 23, Article 33. https://doi.org/10.1186/s12911-023-02129-z
Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A. L., & Zhou, Y. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv.
https://doi.org/10.48550/arXiv.2102.04306
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer Vision – ECCV 2018 (pp. 833–851). Springer.
https://doi.org/10.1007/978-3-030-01234-2_49
Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H. R., & Xu, D. (2022). Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images. In A. Crimi & S. Bakas (Eds.), Brainlesion: Glioma, multiple sclerosis, stroke and traumatic brain injuries (pp. 272–284). Springer.
https://doi.org/10.1007/978-3-031-08999-2_22
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778).
https://doi.org/10.1109/CVPR.2016.90
Huang, H., Lin, L., Tong, R., Hu, H., Zhang, Q., Iwamoto, Y., Han, X., Chen, Y.-W., & Wu, J. (2020). UNet 3+: A full-scale connected UNet for medical image segmentation. In ICASSP 2020 - IEEE International Conference on Acoustics, Speech and Signal Processing (pp. 1055–1059).
https://doi.org/10.1109/ICASSP40776.2020.9053405
Jiang, P.-T., Zhang, C.-B., Hou, Q., Cheng, M.-M., & Wei, Y. (2021). LayerCAM: Exploring hierarchical class activation maps for localization. IEEE Transactions on Image Processing, 30, 5875–5888.
https://doi.org/10.1109/TIP.2021.3089943
Milletari, F., Navab, N., & Ahmadi, S.-A. (2016). V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 Fourth International Conference on 3D Vision (3DV) (pp. 565–571).
https://doi.org/10.1109/3DV.2016.79
Petsiuk, V., Das, A., & Saenko, K. (2018). RISE: Randomized input sampling for explanation of black-box models. In Proceedings of the British Machine Vision Conference (BMVC).
Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In N. Navab, J. Hornegger, W. Wells, & A. Frangi (Eds.), Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015 (pp. 234–241). Springer.
https://doi.org/10.1007/978-3-319-24574-4_28
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2020). Grad-CAM: Visual explanations from deep networks via gradient-based localization. International Journal of Computer Vision, 128, 336–359.
https://doi.org/10.1007/s11263-019-01228-7
Smilkov, D., Thorat, N., Kim, B., Viégas, F., & Wattenberg, M. (2017). SmoothGrad: Removing noise by adding noise. arXiv.
https://doi.org/10.48550/arXiv.1706.03825
Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (pp. 6105–6114). PMLR. https://proceedings.mlr.press/v97/tan19a.html
Wang, H., Wang, Z., Du, M., Yang, F., Zhang, Z., Ding, S., Mardziel, P., & Hu, X. (2020). Score-CAM: Score-weighted visual explanations for convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (pp. 111–119).
https://doi.org/10.1109/CVPRW50498.2020.00020
Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6230–6239).
https://doi.org/10.1109/CVPR.2017.660
Zhou, H.-Y., Guo, J., Zhang, Y., Yu, L., Wang, L., & Yu, Y. (2021). nnFormer: Interleaved transformer for volumetric segmentation. arXiv.
https://doi.org/10.48550/arXiv.2109.03201
Zhou, Z., Siddiquee, M. M. R., Tajbakhsh, N., & Liang, J. (2018). UNet++: A nested U-Net architecture for medical image segmentation. In D. Stoyanov, Z. Taylor, B. Bernhardstetter, R. Zeghal, F. S. Cohen, E. Hosseinzadeh, Z. Yaniv, P. Abolmaesumi, D. R. Haynor, P. de Jong, M. A. Viergever, T. S. S. Schouw, & S. R. van de Weijer (Eds.), Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support (pp. 3–11). Springer. https://doi.org/10.1007/978-3-030-00889-5_1
Downloads
Published
Data Availability Statement
Datasets generated during the current study are available from the corresponding author upon reasonable request.
Issue
Section
License
Copyright (c) 2026 Mohammad Shahjahan Majib, Md Raiyan Bhuiyan Loreen, Anika Tasnim , Md. Arif Abdullah , Farzana Mozammel Samia , Faysal Ahmed

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
This journal provides immediate open access to its content on the principle that making research freely available to the public supports a greater global exchange of knowledge. Users are permitted to read, download, copy, distribute, print, search, or link to the full texts of the articles, provided that appropriate credit is given to the original authors.