Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity
| Source: arXiv AI
Tags: medical imaging, domain adaptation, ultrasound, image segmentation, GAN, speech science
Starting from just 5 labeled source images, a source-free domain adaptation framework for ultrasound tongue segmentation uses pseudo-label refinement and a conditional GAN to outperform supervised baselines across 12 cross-dataset transfer pairs spanning 8 datasets.
Details
Ultrasound tongue imaging supports speech therapy and phonetics research, but cross-dataset generalization is poor due to probe variability, acquisition noise, and limited annotations. This paper presents a source-free domain adaptation framework operating under extreme data scarcity. The method starts from a checkpoint pretrained on only 5 labeled source images (intentionally underfitted), then adapts to a fully unlabeled target domain through three mechanisms: iterative pseudo-label refinement, a contour-based quality-control module filtering unreliable masks, and a segmentation-guided conditional GAN generating synthetic target-style image-mask pairs. A student model trains on a mixture of clean pseudo-labels, noisy pseudo-labels with consistency regularization, and synthetic samples. Evaluated on 12 source-target transfer pairs across 8 ultrasound tongue imaging datasets, the framework outperforms baselines including supervised methods. Ablation studies and source-size scaling experiments validate each component's contribution. For medical imaging practitioners, the source-free design addresses a common clinical constraint: source training data often cannot be shared for privacy reasons. The ability to adapt from 5 labeled examples reduces annotation burden substantially.