Diabetic foot ulcers are among the most quietly devastating complications of diabetes, and clinicians have long relied on manual tracing and visual estimation to track whether a wound is healing or worsening. A new study published in BMC Medical Imaging offers a carefully validated artificial intelligence system that can outline these ulcers in clinical photographs automatically, and, crucially, explain where it was looking when it made its decision. The work, led by Akwasi Asare and colleagues at Ghana Communication Technology University together with Stephen E. Moore of the University of Cape Coast, presents a hybrid deep learning model that pairs the local precision of convolutional networks with the global awareness of vision transformers, wrapped in a layer of quantitative explainability that is still rare in medical imaging research.
The technical heart of the system is an architecture the authors call TransUNet-GradCAM, which combines two complementary traditions in computer vision. Convolutional neural networks of the U-Net family excel at localising fine structures in images because their stacked filters capture edges, textures, and small-scale patterns, and their encoder-decoder design with skip connections preserves spatial detail through the network. However, convolutions only see a small neighbourhood of pixels at a time, which makes it hard for them to relate distant parts of an image. Vision transformers solve this by using self-attention, a mechanism that lets every patch of an image weigh the relevance of every other patch, capturing long-range dependencies across the whole photograph. The new model embeds a Vision Transformer bottleneck in the middle of a convolutional U-Net, so that local detail from the encoder and global context from the transformer are fused before the decoder reconstructs a pixel-precise wound mask.
Two further refinements sharpen the design. Attention-gated skip connections act as filters on the pathways that carry fine spatial information from encoder to decoder, suppressing irrelevant background features so that the decoder concentrates on the wound region. The model was trained on the public Foot Ulcer Segmentation Challenge dataset using a hybrid loss function that blends Dice loss, which directly optimises the overlap between predicted and true wound masks, with binary cross-entropy, which supervises each pixel independently. This combination matters because ulcer segmentation is extremely imbalanced: the wound often occupies a small fraction of the image, and a naive pixel-wise loss can be dominated by easy background pixels. The authors are refreshingly candid about their contribution, emphasising that the novelty lies not in inventing a new architecture but in rigorously validating and explaining an application of one.
The performance figures are reported with a statistical discipline that deserves attention in its own right. Rather than quoting a single best run, the team trained the model across five random seeds and reported all results as means with 95 percent confidence intervals at a single fixed decision threshold. On the internal validation set the model achieved a Dice score of 0.8035 plus or minus 0.0053, meaning its wound outlines overlapped expert annotations by roughly eighty percent, and an intersection-over-union of 0.7149 plus or minus 0.0073. Boundary accuracy metrics were similarly solid, with a 95th-percentile Hausdorff distance of 19.74 pixels and an average symmetric surface distance of 6.12 pixels, indicating that even the worst-case boundary errors remained within clinically tolerable ranges.
Perhaps the most scientifically interesting result is what the ablation study did and did not find. When the authors systematically removed components one at a time, only the hybrid loss produced a statistically significant change in Dice score, a drop of 0.038 with a p-value below 0.001. The transformer bottleneck, the attention gates, and the data augmentation pipeline each produced small effects that did not reach statistical significance on the internal validation data. In an era when deep learning papers often attribute large gains to architectural novelties, this negative result is a valuable corrective: the training objective, not the fancy modules, drove the measurable improvement in-domain. It also illustrates why multi-seed reporting with confidence intervals should become standard practice in medical imaging.
Generalisation to unseen data was tested honestly rather than optimistically. Without any retraining, the model retained about 92 percent of its internal Dice performance on an external cohort of 278 images from the AZH Wound Care Center, scoring 0.7460. A much smaller Medetec subset of eight images was used only as a qualitative check. The authors interpret this as partial rather than robust generalisation under domain shift, a phrase that reflects the reality that clinical photographs vary enormously with camera type, lighting, skin tone, and imaging protocol. For a model intended to support wound monitoring across hospitals, that gap between internal and external performance is exactly the kind of information clinicians and regulators need to see reported transparently.
The explainability component distinguishes this work from most segmentation studies. The team compared two ways of visualising what the network attends to: Grad-CAM, which uses gradients flowing back through the convolutional layers to produce heatmaps of class-relevant regions, and attention rollout, which aggregates the transformer’s self-attention weights across layers. A quantitative evaluation on 200 images found that Grad-CAM heatmaps were far more tightly localised to the wound, concentrating 87.1 percent of their energy inside the annotated mask compared with only 10.2 percent for attention rollout. Yet attention rollout was significantly more faithful to the model’s actual decision process, with the difference reaching statistical significance at p equals 0.038. The conclusion is that the two methods are complementary: Grad-CAM tells you where the model looked, while attention rollout better reflects how the decision was formed.
For clinical translation, the practical numbers matter as much as the benchmarks. Predicted wound areas correlated strongly with expert measurements, with a Pearson correlation coefficient of 0.944, suggesting the system could reliably support longitudinal tracking of ulcer size, one of the key indicators clinicians use to judge healing progress. Equally important, the model is lightweight by modern standards, containing only 8.79 million parameters. That compact footprint opens the door to deployment on modest hospital hardware or even edge devices, rather than requiring data-centre GPUs, which is a meaningful consideration for healthcare settings in low-resource regions where diabetic foot complications are often most severe.
The study’s methodology also models good research ethics and reproducibility. All experiments used publicly available, fully de-identified datasets, including the Foot Ulcer Segmentation Challenge data and the publicly released AZH and Medetec wound collections, with images de-identified at source in accordance with HIPAA requirements. Because no new human data were collected, no additional institutional ethical approval was required, and the work was conducted under the principles of the Declaration of Helsinki. The authors declare no competing interests and received no external funding, conducting the research with institutional resources. The open access publication means that clinicians and researchers anywhere can examine the full methods, results, and limitations without a subscription.
The broader significance of this study lies less in its headline numbers than in its template for trustworthy medical AI. It shows a hybrid architecture combining convolutional and transformer strengths, an ablation that honestly separates real contributions from negligible ones, external validation that quantifies rather than hides the domain-shift penalty, and an explainability analysis that measures whether visual explanations actually reflect model behaviour. Diabetic foot ulcers affect millions of people worldwide and can lead to infection, amputation, and death when healing is not monitored closely. A tool that segments wounds accurately, agrees with expert area measurements, runs on lightweight hardware, and shows clinicians exactly where it looked could become a genuine aid to diagnosis, treatment planning, and longitudinal wound care, provided future work continues to close the gap between laboratory benchmarks and the messy variability of real-world clinical photography.
Subject of Research: Explainable deep learning for automated diabetic foot ulcer segmentation in clinical photographs
Article Title: TransUNet-GradCAM: a hybrid transformer-U-Net with self-attention and explainable visualizations for foot ulcer segmentation
Article References: Asare, A., Sagoe, M., Asare, J. W., & Moore, S. E. (2026). TransUNet-GradCAM: a hybrid transformer-U-Net with self-attention and explainable visualizations for foot ulcer segmentation. BMC Medical Imaging, 26(1), Article 474. https://doi.org/10.1186/s12880-026-02793-3
Image Credits: AI Generated
DOI: 10.1186/s12880-026-02793-3
Keywords: diabetic foot ulcer, image segmentation, U-Net, vision transformer, Grad-CAM, explainable AI, medical imaging, wound monitoring, deep learning, domain shift, Dice score, BMC Medical Imaging
News Source: Ophelia Keating. (October 11, 2026). AI Learns to See Diabetic Foot Ulcers Clearly and Shows Its Work. Scienmag.



