Transfer learning for genomic prediction in underrepresented populations

| Source: Google Research Blog

Tags: polygenic risk scores, transfer learning, genomics, UK Biobank, Biobank Japan, Google Research, health equity

Google Research finds that transfer learning from European genome databases (UK Biobank) improves polygenic risk score accuracy in underrepresented Japanese populations only when target cohort size is small—once local samples grow, European-derived models degrade performance, especially for population-specific traits.

Details

Polygenic risk scores (PRS) predict disease risk from genetic variants, but existing models are heavily biased toward European populations. Most genome-wide association studies have used European cohorts, causing significant accuracy drops when applied to non-European groups—a long-standing equity problem in precision medicine. Google Research evaluated PRS transferability by training on UK Biobank data (hundreds of thousands of European individuals) and testing on Biobank Japan (BBJ, ~200,000 Japanese individuals) across eight clinical traits: BMI, systolic and diastolic blood pressure, red and white blood cell counts, HDL cholesterol, LDL cholesterol, and blood glucose. The key finding is conditional: transfer learning from European cohorts improves predictions when the target population's sample size is small, but degrades accuracy as the target cohort scales up. The degradation is especially severe for traits with population-specific genetic architectures, where the underlying genetic signals differ meaningfully between populations. For health systems considering PRS tools for non-European patients, the practical takeaway is clear: target-population-specific training data becomes more valuable than European-derived transfer learning at scale. Mixed transfer approaches should be evaluated trait by trait rather than assumed to generalize.