What gaps in my own work changed
My research on populations of Jammu and Kashmir began with phylogenetic questions: how lineages are distributed, what they may reveal about population history, and how regional genetic variation fits within wider South Asian patterns. Work across phylogenetics, healthcare genetics and rare diseases also exposed important limitations in my own studies—gaps in sampling, incomplete reference panels, uncertain resolution of lineage markers, and too little translation from population observations to clinical benefit.
Those gaps changed my working ideology. Stronger claims about origin cannot compensate for weak evidence, and claims of exceptional ancestry, purity or supremacy create controversy without improving anyone’s health. The more important need is to build genetics that works for people: representative local frequency resources, more reliable rare-disease variant interpretation, pharmacogenomic evidence, clinically relevant screening, and careful validation before a result reaches a patient.
I therefore shifted the centre of my work from ancestry narratives to healthcare applications of genomics. Phylogenetics remains useful when its uncertainty is stated honestly, but it should support—not overshadow—the larger purpose: reducing diagnostic delay, preventing avoidable adverse drug reactions, recognizing founder variants, and making genomic medicine more useful for communities in Jammu and Kashmir.
The labels are real. The stories attached to them are not.
When a paper assigns a haplogroup such as M65a, R1a, or U2b2, the classification itself is sound science. It is a genealogical clade defined by shared mutations, as valid as any phylogenetic label in biology.
The problem begins the moment meaning is draped over it. “Steppe ancestry.” “Iranian farmer component.” “Aryan lineage.” These names quietly convert statistical constructs into ancestral tribes, as if an ADMIXTURE component were a real people whose essence modern groups carry. Technically, these components are mathematical decompositions that depend entirely on which reference samples the analyst chose. Change the reference panel and the “ancestries” change. Every methods section admits this. Every abstract, press release, and YouTube video forgets it.
I treat ancestry components the way we treat principal components anywhere else in biology: useful axes of variation, not identities.
What the raw data actually shows: a braided stream, not arrows on a map
Strip away the imposed population labels and plot human variation directly, and you do not see discrete peoples. You see clines. Smooth gradients tracking geography, classic isolation by distance, exactly as in any other widespread mammal species. Humans are not exceptional here.
“Populations” appear as crisp entities largely because of how we sample. Collect fifty individuals from one location, assign them an ethnonym, and the label hardens into a lineage. Add recent endogamy, which in South Asia largely crystallized only in the last two thousand years or so, and some boundaries sharpen further. But the deep structure underneath remains continuous.
The same applies to “migration events.” The textbook picture of Steppe pastoralists “arriving” around 2000 BCE compresses centuries of movement in both directions into a single arrow. The ancient DNA record itself shows South Asian ancestry flowing into Bronze Age Central Asia, not only the reverse. That direction almost never makes a headline. The honest model of our species is a braided stream: continuous exchange to and fro, with lineages diverging only where time and geography thinned the flow, and even then rarely completely.
Gene pools are shaped mainly by demography, not destiny
One refinement I insist on, because it closes a door that bad actors love to walk through: what differentiates regional gene pools is overwhelmingly neutral demographic process. Drift, founder effects, bottlenecks, marriage customs, and migration. Not adaptation.
Selection is real but touches a small fraction of the genome: EPAS1 for altitude, LCT for lactase persistence, a handful of pigmentation and immune loci. Most of what statistically distinguishes a Kashmiri from a Tamil from a Tajik is drift accumulated across the exchange network. Historical accident, not adaptive superiority. This matters, because the phrasing “group X adapted to be Y” is precisely how hierarchy sneaks back into the science. The drift dominant view is both more accurate and harder to abuse.
The supremacy problem is structural, not incidental
Population genetics was born inside race science, and the field has never fully paid off that debt. Today the bias is subtler but systematic.
Reference panels remain Eurocentric, so everyone else’s variation is measured against a European yardstick. “West Eurasian” ancestry is narrated as an active input that “shaped” other populations, while gene flow in the opposite direction becomes a footnote in the supplement. The superlative vocabulary, words like “purest,” “most preserved,” “highest ancestry of any community,” maps directly onto old hierarchy of lineages thinking.
Nor is this only a Western problem. Within South Asia the same machinery now services caste and regional pride. Steppe percentage has become a status currency on ancestry forums, communities compete over R1a frequencies, and viral videos manufacture genetic exceptionalism for whichever audience they target. The recent Kashmir video is a textbook case. It debunked one origin myth, the Lost Tribes and Alexander’s soldiers, only to install a new one: uniquely ancient, uniquely pure, uniquely preserved. It inflated shared regional lineages into exclusive ones and deleted the one ancient sample that showed medieval demographic turnover. Old supremacy logic, new lab coat.
Fragmentation dressed up as narrative
Ancient DNA studies work with brutally sparse material. In the Kashmir case, twelve skeletons yielded four usable mitogenomes, and an entire “4,000 years of unbroken continuity” narrative was hung on a single Neolithic sample. Missing links are not neutral gaps. They are degrees of freedom that let an author draw whichever arc the data does not forbid. A haplogroup shared between one ancient individual and a modern population is consistent with continuity. It does not demonstrate it. I now read every ancestry narrative with that gap versus claim ratio in mind.
A note on AI tools, which many of us use for literature work. A language model reads the literature the field produced, so the field’s framing biases become its priors. It can help flag bias when explicitly asked, but it cannot fully step outside the corpus it learned from. The correction still has to come from us, at the level of raw frequency tables and study design.
Where this stops being academic: medicine pays the bill
Here is the part I most want colleagues to sit with. The prime objective of studying human genetics is to reduce the burden of disease. Every false association and every inflated ancestry label works directly against that objective, and the damage is measurable.
The discovery base is skewed, so the tools do not transfer. The large majority of genome wide association study participants to date are of European ancestry, while South Asians, nearly a quarter of humanity, contribute only a few percent. Polygenic risk scores trained on those cohorts lose substantial accuracy when applied to our populations. A cardiac or diabetes risk score that performs well in a UK cohort can misclassify a patient in Srinagar or Jammu, and South Asians carry some of the highest cardiometabolic burden in the world. The people who most need the prediction are the ones the prediction serves worst.
Ethnic labels are poor proxies for the genetics that actually matters in the clinic. Pharmacogenomics runs on specific variants: CYP2D6 and CYP2C19 status for antidepressants and clopidogrel, HLA-B*15:02 for carbamazepine hypersensitivity, G6PD deficiency for primaquine and a long list of oxidant drugs, VKORC1 and CYP2C9 for warfarin dosing. The frequencies of these variants vary enormously across South Asian communities precisely because of the endogamy and founder effects I described above. A single “South Asian” or “Indian” box on a trial form averages over that structure and erases it. When medicine borrows the mythologized labels instead of measuring the variants, patients get the wrong dose.
Race based corrections institutionalize the error. For years, kidney function equations carried a race coefficient, and drugs like BiDil were approved for a racial category rather than a mechanism. These are cases where a social label was mistaken for a biological variable. The lesson generalizes: whenever an ancestry story is treated as a physiological fact, someone eventually writes it into a clinical algorithm, and unwinding it takes a decade.
Variant interpretation fails without representative references. Clinical genetics decides whether a variant is pathogenic partly by asking how common it is in reference databases. Because our populations are underrepresented, benign variants common in a Pahari or Gujjar or Kashmiri community can look alarmingly rare against a European database and get flagged as disease causing, while true founder pathogenic variants go uncatalogued. Families receive wrong answers in both directions. Endogamous South Asian groups are exactly the setting, like Finland or the Ashkenazi population, where founder variant mapping delivers the biggest clinical wins. We are leaving those wins on the table while the field spends its energy ranking lineages.
Drug development inherits all of it. Trials recruit on labels, dose on averages, and generalize from unrepresentative cohorts. Efficacy and adverse event profiles that differ across genetic backgrounds surface only after deployment, in the populations that were never enrolled. The cost of the origin story genre is not only cultural. It is paid in misdosed patients, missed diagnoses, and drugs that work less well for most of humanity.
So my position is simple. If the purpose of genetics is to reduce disease burden, then every unit of effort spent constructing lineage prestige is worse than wasted. It actively trains clinicians, regulators, and the public to treat labels as biology, which is the exact error precision medicine exists to eliminate. Precision medicine done right is the opposite of ancestry romance: measure the individual genome, know the local variant landscape, and let the ethnic label carry no clinical weight at all.
What an unbiased analysis would look like
If we drop the origin story genre altogether, the alternative is straightforward and, oddly, almost nobody publishes it.
Model regions as connectivity networks, not lineage trees. Report who shares what with whom, in both directions, with dates and confidence intervals. A gene flow map rather than a descent pedigree. Present haplogroup data as frequency surfaces across geography instead of ethnic possessions. Name ancestral components after archaeological sites or coordinates, never after modern peoples. State the sample sizes and the missing links in the abstract, not the supplement. Retire the superlatives entirely: “enriched” is a measurement, “purest” is an ideology. And pair every population survey with its clinical deliverable, the local allele frequencies for actionable pharmacogenomic and disease variants, so the work feeds medicine rather than mythology.
Read this way, our own region stops being a battleground of competing origin claims and becomes what the data has shown all along: Kashmir, Jammu, the Pahari belt, Ladakh, Swat, and Central Asia as one continuously exchanging system. Deep local roots and repeated incoming and outgoing flow, carried by women as much as men, layered over millennia. No population in this network is anyone’s ancestor in chief or anyone’s derivative. And every one of them deserves a reference panel that lets a doctor dose a drug correctly. That is not a smaller story than the myths. It is the only one the evidence supports, and the only one that saves lives.
