💥Join UPSC 2027,2028 Mentorship (August Batch) + XFactor Notes & Microthemes PDF

Foreign Policy Watch: India-SAARC Nations

Why are South Asians missing from global health databases

Why in the News

Genome wide association studies between 2005 and 2025 drew more than 86 per cent of their participants from European ancestry populations, while South Asians accounted for less than 1 per cent. That skew is now being carried into the reference atlases used to train artificial intelligence models in medicine, which converts a historical sampling gap into a bias that reproduces itself at clinical scale across South Asia and the wider low and middle income world.

What is an integrated biobank?

  1. Definition: An integrated biobank is a large repository that stores biological samples from consenting participants alongside linked data about them, and makes both available to researchers.
  2. What it integrates: It combines participants’ genomic information with electronic health records, environmental exposures and lifestyle data, so that genetic variation can be read against real health outcomes.

What is a genome wide association study?

  1. Definition: A genome wide association study (GWAS) scans the genomes of many individuals to find genetic variants that occur more often in people with a particular disease than in people without it.
  2. What it produces: It yields a list of variants statistically associated with a trait or disease, which is the raw material for downstream risk prediction tools.

What is a polygenic risk score?

  1. Definition: A polygenic risk score combines the effects of many genetic variants associated with a disease to estimate a person’s overall genetic risk for it.
  2. Why ancestry matters to it: The score’s weights are derived from the population it was built in, so applying it to a population with a different variant frequency structure changes its accuracy.

What is a single cell atlas?

  1. Definition: A single cell atlas is a reference map that catalogues the gene activity of individual cells across tissues and organs, rather than of a tissue sample as a whole.

What is a low and middle income country?

  1. Definition: Low and middle income countries are the economies classified by the World Bank below the high income threshold on gross national income per capita, a grouping used in global health to identify where disease burden and research funding diverge.
  2. Why the category is used here: The under representation problem is stated at the level of this group, with India, Pakistan, Bangladesh and Sri Lanka as instances inside it rather than as separate cases.

What are potential years of life lost?

  1. Definition: Potential years of life lost is a measure of premature mortality that counts the years a person would have lived had they reached a reference life expectancy.
  2. What it captures that a death count does not: It weights a death at a young age more heavily than a death in old age, which is why it shifts burden sharply towards countries with high early mortality.

What is G6PD deficiency?

  1. Definition: Glucose-6-phosphate dehydrogenase (G6PD) deficiency is an inherited enzyme disorder that can cause a form of anaemia when red blood cells break down under oxidative stress from certain drugs, infections or foods.

What is metabolic syndrome?

  1. Definition: Metabolic syndrome is a clustering of obesity, raised blood sugar, abnormal cholesterol and high blood pressure that together raise the risk of cardiovascular disease and type 2 diabetes.

How large is the ancestry gap in global genomic databases?

  1. The genome wide association study record: The GWAS Catalogue is maintained by the National Human Genome Research Institute (NHGRI) and the European Bioinformatics Institute (EBI). It records that more than 86 per cent of participants in these studies between 2005 and 2025 were of European ancestry.
  2. The South Asian share: South Asians accounted for less than 1 per cent of participants over that same twenty year period.
  3. The gap at the country income level: Over 90 per cent of the world’s potential years of life lost occurred in low and middle income countries. About 10 per cent of global health research funding addressed the health needs of those countries.
  4. The share of humanity excluded: More than 20 per cent of the world is being neglected in multi modal data integration, and the exclusion denies those populations the opportunity to attain the maximal possible health.
  5. The pattern repeats in newer tools: A study published in Cell Genomics reviewed more than 13,500 samples across three major single cell resources and found a striking and pervasive European over representation alongside under representation of Asian and Latino individuals.
  6. The three resources reviewed: The study covered the Human Cell Atlas, the Human Tumour Atlas Network and the PsychAD Consortium.
  7. South Asians absent from the biobanks too: South Asians remain largely absent from integrated biobanks such as the U.K. Biobank, which are the repositories that transformed biomedical research.

Why does a European skewed dataset produce worse clinical tools for South Asians?

  1. The burden runs the other way: South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than people of European ancestry, so the tools built on European heavy data are least accurate for the population that needs them most.
  2. The diabetes case: More than one in ten adults globally now live with diabetes, the risk is higher for people of South Asian ancestry and it appears earlier than in many other populations.
  3. India’s projected burden: The number of people with diabetes in India alone is projected to reach 125 million by 2045.
  4. Risk scores lose accuracy across ancestry: A 2023 study found that polygenic risk scores for multiple sclerosis were less accurate when applied to South Asian populations.
  5. Functional predictions are untested: Most predictions about how variants affect gene expression or cell function are inferred from European datasets, and it is not known which of those predictions hold in South Asians.
  6. The consequence for drug discovery: This limits the ability to understand disease mechanisms and to identify drug targets relevant to South Asian populations.
  7. Thresholds themselves need recalibration: Diagnostic thresholds, risk scores and prediction models developed predominantly from European populations require validation and, where necessary, recalibration using South Asian data.

Why can South Asia not be treated as a single genetic block?

  1. One of the most diverse populations on earth: South Asia constitutes one of the most diverse human populations in the world, shaped by thousands of years of migration, cultural diversity, endogamy and consanguineous marriages.
  2. Lumping erases the differences: Much existing research groups South Asians, Southeast Asians, West Asians and other Asian populations together, obscuring important differences between them.
  3. Variation within the region: G6PD deficiency varies considerably across South Asia, with some ethnic groups in Pakistan and Afghanistan carrying the trait at much higher rates than others.
  4. Variation within a single population: A study from Sri Lanka found that cardiometabolic risk did not fit into a single metabolic syndrome profile, and within the same population men and women showed distinct patterns of obesity, blood sugar, cholesterol and blood pressure.
  5. The scale of Indian variation: The GenomeIndia Project has already identified more than 40 million genetic variants unique to the Indian population.
  6. Who must be sampled: India cannot realistically be treated as one genetic block, and inclusion must extend to distinct endogamous and tribal groups rather than a few urban cohorts, since many of the harmful variants found there are not seen anywhere else.

Why is the data missing in the first place?

  1. Infrastructure followed the money: Research funding, institutions, registries, biobanks and large population cohorts have historically been built and sustained where the money already was.
  2. What that left behind: Low and middle income countries were left with inadequate laboratory infrastructure, inadequate biobanking facilities and too few trained personnel to run comparable studies at scale.
  3. The imbalance is not only financial: It shapes whose problems are studied, whose questions are prioritised and whose evidence informs health policy and practice.
  4. Ancestry classification practice: Where non European participants are recruited, they are frequently pooled into broad continental categories, which means the data collected does not resolve the differences it was collected to capture.

Why is genomic research hard for South Asian countries to prioritise?

  1. Competing immediate needs: For most South Asian countries genomic research is difficult to prioritise against more immediate and pressing public health demands.
  2. Infectious disease: Communicable disease control absorbs public health budgets and personnel that a genomics programme would otherwise draw on.
  3. Maternal and child health: Maternal and child health programmes command prior claim because their outcomes are measurable within a single planning cycle.
  4. Non communicable diseases: Treatment and screening for non communicable diseases compete for the same budget line that genomic infrastructure would need.
  5. The mismatch in horizons: Genomic infrastructure returns value over a decade or more, while the health systems being asked to fund it are assessed on annual outcome indicators.
  6. Why deferring is costly: Every year the region defers, the reference atlases and the models trained on them are built further without it, which raises the cost of correction later.

What genomic cohorts already exist in South Asia and why do they not add up?

  1. GenomeIndia: India’s national population reference cohort.
  2. Phenome India: An Indian longitudinal cohort linking health, lifestyle and clinical measurements across participants.
  3. Longevity India: An Indian cohort focused on ageing and the biological determinants of long life.
  4. Sri Lankan Twin Registry Biobank: A Sri Lankan registry and biobank built around twin pairs, which permits separation of genetic and environmental effects.
  5. Pakistan Genome Resource: A Pakistani national genomic resource built on population sampling.
  6. Why they do not combine: These independent cohorts and biobanks are mostly focused on individual diseases or specific populations, and often use different systems for collecting and storing data, which makes it difficult to bring them together for large genetic studies.
  7. The Indian case specifically: India has several sizeable cohorts, but no harmonised system yet exists that lets researchers within and across borders work across them easily.

What does the U.K. Biobank model demonstrate that South Asian cohorts currently cannot?

  1. United Kingdom, the integrated design: The U.K. Biobank links each participant’s genomic information to electronic health records, environmental exposure data and lifestyle data in a single resource, which is the feature that allows genotype to be read against outcome.
  2. What that integration produced: Repositories of this design accelerated drug development, informed clinical guidelines and shaped public health policy across multiple countries, not only in the country that built them.
  3. The contrast with South Asia: South Asian cohorts are disease specific or population specific and are stored on divergent systems, so no equivalent linkage across genomics, clinical records and exposure exists in the region.
  4. The limit of this comparison: The U.K. Biobank is the single substantive institutional model in the evidence here, so it establishes what an integrated design makes possible, not a ranked set of alternative national models to choose between.

What does the regional proposal recommend?

  1. The authorship: A perspective in the Lancet Regional Health – Southeast Asia, written by scientists across India, Pakistan, Bangladesh and Sri Lanka, sets out the regional response.
  2. The core warning: The region risks being excluded from the genomic revolution unless it builds the infrastructure itself, rather than waiting for inclusion in datasets built elsewhere.
  3. Regional collaboration between existing assets: The proposal is to build greater collaboration between existing biobanks and cohorts, rather than to construct a new central repository from scratch.
  4. Interoperability: The aim is a system in which existing datasets can speak to each other, which is the specific technical gap that keeps Indian cohorts from being analysed together.
  5. Inclusion of overlooked populations: Populations that have historically been overlooked, including distinct endogamous and tribal groups, are to be brought into the sampling frame.
  6. Retained control over data use: South Asian researchers and institutions are to retain a meaningful role in how their data are used.
  7. Benefit sharing: The researchers generating the data are to share in the scientific benefits, which addresses the extraction pattern rather than only the data gap.

Why does the gap compound rather than stay constant?

  1. The atlases became reference maps: Single cell atlases are now the reference maps for biology and medicine, so an error in the map propagates into everything read against it.
  2. They are now training data: Those same atlases are increasingly used to train the artificial intelligence models that will shape future research and care.
  3. Scale changes the nature of the problem: If the underlying data continues to be skewed, the artificial intelligence models and clinical tools built on top of it will reproduce and repeat those biases at a much larger scale.
  4. From a research gap to a clinical one: A skewed research dataset produced inaccurate studies, a skewed training dataset produces inaccurate bedside tools deployed on populations that were never in the data.
  5. The window is closing but not shut: It is late for the region to build its own infrastructure, and it is still not too late.

Challenges to building a South Asian genomic data infrastructure

  1. Non interoperable data standards: Existing cohorts use different collection, phenotyping and storage systems, so pooling requires retrospective harmonisation that the original consent may not permit. Eg. India’s several sizeable cohorts have no harmonised system that lets researchers work across them.
  2. Consent and benefit sharing for community level data: Genomic data from an endogamous or tribal group carries group level implications that individual consent does not cover. Eg. The Biological Diversity Act, 2002 governs access and benefit sharing for biological resources, and its application to human genomic data drawn from identified communities is unsettled.
  3. Sustained financing beyond donor cycles: Climate and health workforce experience across the region shows that capacity built on project funding disappears when the project ends. Eg. Genomic surveillance capacity expanded rapidly during the pandemic and contracted once the emergency funding lapsed.
  4. Cross border data transfer rules: Regional pooling requires moving identifiable health data across national jurisdictions with differing data protection regimes. Eg. The Digital Personal Data Protection Act, 2023 permits the Central Government to restrict transfer of personal data to notified countries.
  5. Shortage of trained personnel: Bioinformatics, genetic counselling and biobank management skills are scarce relative to the sequencing capacity being installed. Eg. Genetic counsellors in India number in the low hundreds against a population carrying a large inherited disease burden.
  6. Risk of genetic discrimination: Widening genomic data collection without a statutory bar exposes participants to insurance and employment consequences. Eg. The Delhi High Court in United India Insurance vs Jai Parkash Tayal, 2018 held the exclusion of genetic disorders from health insurance cover unconstitutional, in the absence of any general anti discrimination statute.
  7. Sampling reaching only urban cohorts: Recruitment gravitates to tertiary hospitals and metropolitan volunteers, reproducing inside India the same skew the region objects to globally. Eg. Inclusion of distinct endogamous and tribal groups has been identified as the specific gap in Indian sampling, not the overall sample size.

Conclusion

The under representation of South Asians in global genomic databases is no longer only an equity problem in research, it is becoming an engineering problem in clinical artificial intelligence. With more than 86 per cent of genome wide association study participants of European ancestry and South Asians below 1 per cent, the reference atlases now being used as training data carry that skew forward at scale. The response has shifted from asking for inclusion in datasets built elsewhere to building interoperable regional infrastructure that keeps control and benefit with the researchers generating the data. What remains unresolved is financing, since the region must fund a decade long investment against infectious disease, maternal and child health and non communicable disease needs that compete for the same budget.

“[2026] Which of the following statements with regard to Genome India Project is/are correct?

1. It is a part of the Human Genome Project.

2. The project is funded by the Department of Biotechnology (DBT), Government of India.

3. Its primary aim is to build a catalogue of genetic diversity of the Indian population.

(a) 1 only

(b) 2 and 3 only

(c) 1 and 2 only

(d) 1, 2 and 3


Join the Community

Join us across Social Media platforms.