CAP-UNet: frequency-decoupled and geometry-aware attention for medical image segmentation across multiple imaging modalities
Abstract
Accurate medical image segmentation is frequently hampered by pervasive challenges such as severe visual camouflage, extreme scale variations, and highly irregular anatomical boundaries across diverse imaging modalities. To address these challenges, we propose CAP-UNet, a geometry-aware encoder-decoder framework evaluated across four independent single-modality medical image segmentation benchmarks. Specifically, we introduce a decoupled skip-connection pathway featuring a Camouflage-Decoupled Frequency Attention (CDFA) module, which leverages parallel spatial and frequency-domain transformations to explicitly enhance target structures while suppressing low-contrast background interference. This pathway is further refined by a Boundary-Aligned Multi-scale Masking (BAMM) module that dynamically aggregates hierarchical contextual features and improves boundary consistency. Furthermore, we design a Curvature & Recalibration Attention (CRA) module to replace conventional decoder convolutions. By jointly modeling local geometric deformation and channel-wise feature recalibration, the CRA module effectively adapts to irregular anatomical structures while reducing modality-specific imaging interference. Extensive experiments on four representative public datasets, including colonoscopy, dermoscopy, ultrasound, and panoramic X-ray images, demonstrate that CAP-UNet consistently achieves superior segmentation performance compared with recent CNN-based, Transformer-based, and hybrid methods. These results indicate that CAP-UNet provides a robust and effective segmentation framework across diverse single-modality clinical scenarios.
// Source
Authors: Yuting Zheng, Qianping Zhu, 牛殿礼
Institutions: Zhejiang Hospital