Research
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
arXiv:2604.10912v1 Announce Type: new Abstract: Medical image segmentation remains challenging due to limited fine-grained annotations, complex anatomical structures, and image degradation from noise,
arXiv:2604.10912v1 Announce Type: new Abstract: Medical image segmentation remains challenging due to limited fine-grained annotations, complex anatomical structures, and image degradation from noise, low contrast, or illumination variation. We propose TAMISeg, a text-guided segmentation framework that incorporates clinical language prompts and semantic distillation as auxiliary semantic cues to enhance visual understanding and reduce reliance on pixel-level fine-grained annotations. TAMISeg integrates three core components: a consistency-aware encoder pretrained with strong perturbations for robust feature extraction, a semantic encoder distillation module with supervision from a frozen DINOv3 teacher to enhance semantic discriminability, and a scale-adaptive decoder that segments anatomical structures across different spatial scales. Experiments on the Kvasir-SEG, MosMedData+, and QaTa-COV19 datasets demonstrate that TAMISeg consistently outperforms existing uni-modal and multi-modal methods in both qualitative and quantitative evaluations. Code will be made publicly available at https://github.com/qczggaoqiang/TAMISeg.
Related
- M^{2}SNet: Multi-scale in Multi-scale Subtraction Network for Medical Image Segmentation
- SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation
- RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation
- Weakly-Supervised Lung Nodule Segmentation via Training-Free Guidance of 3D Rectified Flow
- Hide-and-Seek Attribution: Weakly Supervised Segmentation of Vertebral Metastases in CT
Source: arXiv cs.CV | 2026-04-14