To do($x$) or not to do($x$): Medical Image Counterfactuals for Dataset Augmentation
| Source: arXiv AI
Tags: medical imaging, counterfactual generation, data augmentation, causal AI, fairness, dataset bias
A causal approach to counterfactual image generation outperforms standard conditional methods for medical data augmentation — causally-generated synthetic training data reduces model sensitivity to dataset biases and improves fairness across patient subgroups.
Details
Medical AI models frequently inherit biases from skewed training datasets, limiting clinical reliability across patient populations. This paper from Ibrahim, Evans, and Kamnitsas compares three strategies for synthetic data augmentation: Deterministic (selected variables changed, others fixed), Undirected (variables updated via learned statistical associations without causal direction), and Causal (interventions propagated along a directed causal graph).\n\nThe key finding is that the choice between true counterfactual generation and conditional generation has concrete downstream effects. Models trained on causally-generated synthetic images show measurably better performance and fairness across sensitive subgroups — particularly relevant where anatomy, pathology, and demographic factors have known structural relationships.\n\nPractical implication: when the goal is reducing model sensitivity to correlated dataset biases, causally-grounded augmentation is worth the added complexity of constructing a structural causal model. The paper provides concrete guidance on when this pays off versus when simpler methods suffice.\n\nThis is an arXiv preprint (cs.LG, cs.AI) and has not yet undergone peer review.