Population Genomic Interpretationv2.1.3
Re-evaluated aggregate J&K data · research/learning only

Evidence based Interpretation of SNPs effect at Population Level

694,576SNP frequency rows
24,825K16 SNPs harmonized

00 — Regional context

Jammu & Kashmir: geography, community and genomic variation

A population-frequency view of the Indian Union Territory—not a definition of any person, community or identity.

Place

A connected Himalayan region

Jammu & Kashmir spans subtropical plains and foothills, the Pir Panjal, and the temperate Kashmir Valley. Elevation, valleys, passes and seasonal accessibility have shaped where people settled and how communities connected, while long-distance trade and migration linked the region with the wider Indian subcontinent and Central and West Asia.

Government health-plan geography · regional migration study

People & culture

Diversity is the starting point

The region includes multiple languages, faiths, livelihoods and cultural traditions. Geography can reduce gene flow, while migration increases it; marriage within communities can amplify founder effects and genetic drift. These processes alter some allele frequencies, but they do not create discrete biological “races,” and culture or identity cannot be read from a SNP.

J&K cultural overview · review of endogamy and rare disease research

Population genetics

Frequencies reflect history and sampling

South Asian genomes lie on overlapping ancestry clines shaped by ancient mixture, later mobility and population-specific drift. ANI and ASI are historical statistical constructs—not present-day peoples. A regional average can therefore be informative for research while concealing substantial variation among localities, families and individuals.

Reich et al., Nature (2009) · J&K mitogenome study

What JKDNA offers here

A privacy-preserving community frequency atlas

This page analyzes aggregate SNP frequencies: rsID/name and GRCh37 position matching, haplogroup-marker evidence, reference-component projections, pharmacogenomic markers, disease-associated markers and a ClinVar quality audit. The published output contains no names, sample identifiers, pedigrees or individual genotypes.

With explicit informed consent, a participant’s own genotype could be run through the same auditable marker pipeline and compared with the community reference to create a personal educational report. Population frequencies are useful priors, not a substitute for that person’s genotype, clinical history, ancestry model or confirmatory testing.

Community-drivenBroader voluntary participation can improve representativeness and expose sampling gaps.
Data-minimizedPublic summaries should remain aggregate; consent, withdrawal and controlled raw-data access belong upstream.
Careful interpretationNo identity assignment, diagnosis, prescribing or reproductive decision should be made from this atlas.

01 — Ancestry & lineage

Expanded haplogroup marker evidence

rsID/name matching first, GRCh37 position fallback second; alleles are strand-checked before use.

Y-DNA reference coverage

marker count — not population frequency

mtDNA candidate lineage signals

median diagnostic-marker frequency %
L3
N
R
U
23.79%
U7
8.80%
Other candidate direct-node signals: M4 · M18 · R5 · R8. U7 is nested within U; values must not be summed.
Y haplogroup labelmatched supplied markers
BT173
CT88
I79
G77
C1b1a2b51
E1b1a1~49
E46
F46
J44
E1a43
R1b1a1b43
H341
P1 or K2b2a37
D136
mtDNA nodesignalusable/coveredspreadevidence
U23.79%3/30.39%strong marker concordance
U78.80%3/40.76%strong marker concordance
M44.12%2/21.81%supportive marker concordance
M181.87%2/21.96%supportive marker concordance
R51.62%3/30.52%strong marker concordance
R80.89%2/40.00%supportive marker concordance
Y-DNA frequency distribution remains unavailable. The supplied reference mapped 2,643/2,643 marker rows to the database, but every matched Y allele has frequency zero. The chart is platform/tree coverage only. Recompute chrY frequencies with verified male sex/ploidy coding.
mtDNA expansion: all 1,138 supplied array rows matched (987 unique database sites; 252 polymorphic). U and nested U7 each have three concordant direct-node markers. Other candidate values require extra caution because parent-marker consistency is incomplete. These are population marker signals—not individual calls or a distribution that sums to 100. Source: HaploGrep PhyloTree 17.

02 — Indian population comparison

J&K within modern South Asian frequency space

Common, non-palindromic autosomal SNP frequencies harmonized by rsID, allele and genome build.

3,935complete 1000G SAS markers
3,671GenomeIndia overlaps
2.36%J&K vs pooled GenomeIndia mean |ΔAF|
3.67%five-reference fit RMSE

Best-fitting modern-reference combination

descriptive coefficient %

Aggregate population-frequency PCA

centroids, not individuals
Code1000 Genomes populationfit coefficientchromosome sensitivitymean |ΔAF|RMS ΔAF
PJLPunjabi in Lahore, Pakistan43.5%42.8%–45.4%3.21%4.44%
GIHGujarati Indian in Houston, USA27.5%22.5%–28.5%3.78%5.05%
BEBBengali in Bangladesh20.0%19.3%–20.8%3.86%5.25%
ITUIndian Telugu in the United Kingdom5.2%4.5%–6.0%3.97%5.40%
STUSri Lankan Tamil in the United Kingdom3.8%3.1%–5.4%4.03%5.48%
GenomeIndia overall comparison: 3,671 harmonized markers; mean absolute frequency difference 2.36%, RMS difference 3.16%. The public 9,768-genome archive is pooled across India, so it is used as a national frequency centroid—not as an ancestry source or an 83-group map. Source: GenomeIndia/IBDC.
Affinity, not ancestry fractions. The five coefficients are the constrained combination of PJL, GIH, BEB, ITU and STU population-average frequencies that best reconstructs the aggregate J&K vector. These modern populations are themselves related and admixed; coefficients do not identify genealogy, community, ANI, ASI or AASI. The PCA contains population centroids only and cannot describe variation among J&K individuals. Sources: Ensembl population-frequency API · 1000 Genomes Phase 3.

03 — Global reference projection

Continental-scale and ANI-related components

J&K frequencies projected onto a real 16-component SNPweights panel.

Grouped ancestry-related components

model proportion

K16 components

model proportion
Grouped componentEstimatechromosome sensitivity*
ANI-specific51.0%49.5%–52.5%
West Eurasian-related30.6%29.2%–31.9%
East Eurasian-related13.2%12.2%–14.3%
African-related3.3%2.4%–4.2%
Oceania/Americas-related1.8%1.1%–2.5%
Unresolved ancestor0.0%0.0%–0.2%
K16 componentEstimatesensitivity*
ANI51.0%49.5%–52.5%
Caucasian-HG19.1%17.8%–20.3%
SouthEastAsian8.8%7.9%–9.7%
EHG6.9%5.8%–8.1%
Siberian4.2%3.1%–5.2%
WHG-UHG3.9%3.3%–4.5%
Subsaharian3.1%2.5%–3.7%
Australian1.0%0.4%–1.7%
Neolithic0.7%0.0%–1.4%
Amerindian0.6%0.0%–1.3%
Arctic0.3%0.0%–1.2%
Oceanic0.2%0.0%–0.8%
EastAfrican0.2%0.0%–1.0%
NorthAfrcian0.1%0.0%–1.1%
NearEast0.0%0.0%–0.0%
Ancestor0.0%0.0%–0.2%
Model coverage: 24,825/116,446 panel SNPs harmonized (13,024 direct; 11,801 complement). Projection uses SNPweights normalization and M/M′ correction. *Delete-one-autosome sensitivity intervals are not cohort sampling CIs. Method.
Do not call the remainder “ASI.” K16 has an ANI-labelled component but no AASI/ASI reference. Modern models separate AASI-, Iranian-related and Steppe-related ancestry. These labels are model similarities, not literal genealogy.

04 — Pharmacogenomics

Drug-wise marker distribution and metabolism signal

Genotype frequencies expanded from allele frequency under Hardy-Weinberg; phenotype categories are a single-marker simplification (see caveats) — not full diplotype/CNV-based clinical calling.

Share of population with a non-normal phenotype, by marker

%
abacavir
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
HLA-B (rs2395029)response risk marker6.9%Tag negative 86.7%, Tag positive 13.3%13.3%
azathioprine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.1%Normal Metabolizer 97.9%, Intermediate Metabolizer 2.1%, Poor Metabolizer 0.0%2.1%
TPMT (rs1800460)no function0.2%Normal Metabolizer 99.7%, Intermediate Metabolizer 0.4%, Poor Metabolizer 0.0%0.4%
NUDT15 (rs116855232)no function6.1%Normal Metabolizer 88.2%, Intermediate Metabolizer 11.5%, Poor Metabolizer 0.4%11.8%
capecitabine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
DPYD (rs3918290)no function0.4%Normal Metabolizer 99.1%, Intermediate Metabolizer 0.9%, Poor Metabolizer 0.0%0.9%
clopidogrel
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function1.2%Normal Metabolizer 97.5%, Intermediate Metabolizer 2.4%, Poor Metabolizer 0.0%2.5%
CYP2C19 (rs12248560)increased function16.9%Normal Metabolizer 69.1%, Rapid Metabolizer 28.0%, Ultrarapid Metabolizer 2.8%30.9%
fluorouracil
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
DPYD (rs3918290)no function0.4%Normal Metabolizer 99.1%, Intermediate Metabolizer 0.9%, Poor Metabolizer 0.0%0.9%
mercaptopurine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.1%Normal Metabolizer 97.9%, Intermediate Metabolizer 2.1%, Poor Metabolizer 0.0%2.1%
TPMT (rs1800460)no function0.2%Normal Metabolizer 99.7%, Intermediate Metabolizer 0.4%, Poor Metabolizer 0.0%0.4%
NUDT15 (rs116855232)no function6.1%Normal Metabolizer 88.2%, Intermediate Metabolizer 11.5%, Poor Metabolizer 0.4%11.8%
omeprazole
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function1.2%Normal Metabolizer 97.5%, Intermediate Metabolizer 2.4%, Poor Metabolizer 0.0%2.5%
CYP2C19 (rs12248560)increased function16.9%Normal Metabolizer 69.1%, Rapid Metabolizer 28.0%, Ultrarapid Metabolizer 2.8%30.9%
phenytoin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C9 (rs1057910)decreased function13.5%Normal Metabolizer 74.8%, Normal-Intermediate Metabolizer 23.4%, Intermediate Metabolizer 1.8%25.2%
CYP2C9 (rs1799853)decreased function5.0%Normal Metabolizer 90.3%, Normal-Intermediate Metabolizer 9.4%, Intermediate Metabolizer 0.2%9.7%
simvastatin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
SLCO1B1 (rs4149056)decreased function3.3%Normal Metabolizer 93.6%, Normal-Intermediate Metabolizer 6.3%, Intermediate Metabolizer 0.1%6.4%
tacrolimus
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP3A5 (rs776746)no function76.5%Poor Metabolizer 58.5%, Intermediate Metabolizer 36.0%, Normal Metabolizer 5.5%94.5%
thioguanine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.1%Normal Metabolizer 97.9%, Intermediate Metabolizer 2.1%, Poor Metabolizer 0.0%2.1%
TPMT (rs1800460)no function0.2%Normal Metabolizer 99.7%, Intermediate Metabolizer 0.4%, Poor Metabolizer 0.0%0.4%
NUDT15 (rs116855232)no function6.1%Normal Metabolizer 88.2%, Intermediate Metabolizer 11.5%, Poor Metabolizer 0.4%11.8%
voriconazole
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function1.2%Normal Metabolizer 97.5%, Intermediate Metabolizer 2.4%, Poor Metabolizer 0.0%2.5%
warfarin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C9 (rs1057910)decreased function13.5%Normal Metabolizer 74.8%, Normal-Intermediate Metabolizer 23.4%, Intermediate Metabolizer 1.8%25.2%
CYP2C9 (rs1799853)decreased function5.0%Normal Metabolizer 90.3%, Normal-Intermediate Metabolizer 9.4%, Intermediate Metabolizer 0.2%9.7%
Simplification notice. Phenotype categories here are derived from a single diagnostic SNP per gene under a simplified function model (no-function / decreased-function / increased-function relative to reference). Real clinical PGx phenotyping (esp. for CYP2D6, which needs copy-number detection) combines multiple star alleles per haplotype. Treat this as population-level screening signal, not individual diplotype calls.

05 — Disease Risk

Population-attributable risk summary

Levin's formula combines each lead SNP's population risk-allele frequency with its published odds ratio into a share of disease burden statistically attributable to that variant, at the population level.

Combined PAR by disease

%
Prostate Cancer Moderate population-level contribution
16.4% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs6983267POU5F1B52.1%1.209.4%
rs10993994MSMB55.4%1.157.7%
Type 2 Diabetes Moderate population-level contribution
15.6% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs7903146TCF7L231.6%1.4011.2%
rs5219KCNJ1137.1%1.144.9%
Asthma Moderate population-level contribution
15.0% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs7216389GSDMB40.9%1.3010.9%
rs20541IL1329.7%1.164.5%
Breast Cancer Moderate population-level contribution
13.2% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs2981582FGFR236.0%1.268.6%
rs3803662TOX326.9%1.205.1%
Alzheimer's Disease Moderate population-level contribution
10.4% combined population-attributable risk across 1 lead SNP
SNPGeneRisk allele freqORPAR
rs6733839BIN158.2%1.2010.4%
Rheumatoid Arthritis Low population-level contribution
6.7% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs6910071HLA-DRB18.0%1.806.0%
rs2476601PTPN221.5%1.500.8%
Coronary Artery Disease Low population-level contribution
0.6% combined population-attributable risk across 1 lead SNP
SNPGeneRisk allele freqORPAR
rs10455872LPA1.1%1.600.6%
What PAR does and doesn't mean. A high combined PAR means the included risk alleles are common AND meaningfully increase relative risk in the population — it is not a statement about any individual's personal risk, and ignores gene-environment interaction, polygenic background beyond the listed lead SNPs, and LD/epistasis between loci. Allele frequencies shown are from the supplied J&K database; odds ratios are drawn from the general published GWAS literature for each locus, not J&K-specific studies.

06 — Pathogenic susceptibility

ClinVar cross-reference and quality triage

Official GRCh37 ClinVar P/LP SNVs intersected by position and allele, requiring ≥1 review star.

8,844P/LP frequency cross-references
9practice-guideline reviewed
1,151expert-panel reviewed
10high-frequency QC flags

Prioritized carrier-frequency estimates

HWE marker-level estimate %
VariantGeneCondition(s)ClinVarReviewAllele qHWE carrier
rs75039782CFTRCFTR-related disorder; Cystic fibrosis; not specified; not provided; Bronchiectasis with or without elevaPathogenic★★★★0.3527%0.703%
rs75096551CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; Cystic fibrosis; Congenital bilateral aplasia of Pathogenic★★★★0.1779%0.355%
rs113993959CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; Congenital bilateral aplasia of vas deferens fromPathogenic★★★★0.1770%0.353%
rs75527207CFTRCFTR-related disorder; Congenital bilateral aplasia of vas deferens from CFTR mutation; Cystic fibrosis; Pathogenic★★★★0.1770%0.353%
rs75961395CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; not provided; Congenital bilateral aplasia of vasPathogenic★★★★0.1767%0.353%
rs80224560CFTRCFTR-related disorder; Cystic fibrosis; Respiratory ciliopathies including non-CF bronchiectasis; not spePathogenic★★★★0.1764%0.352%
rs77010898CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; not specified; Congenital bilateral aplasia of vaPathogenic★★★★0.1764%0.352%
rs121908748CFTRCystic fibrosis; Congenital bilateral aplasia of vas deferens from CFTR mutation; not provided; HereditarPathogenic★★★★0.0896%0.179%
rs78655421CFTRCFTR-related disorder; Respiratory ciliopathies including non-CF bronchiectasis; Cystic fibrosis; CongeniPathogenic★★★★0.0893%0.178%
rs1800553ABCA4Syndromic retinitis pigmentosa; Autosomal recessive ABCA4-related disorders; ABCA4-related retinopathy; APathogenic★★★2.1310%4.171%
rs28936700CYP1B1Anterior segment dysgenesis 6; Glaucoma 3A; CYP1B1-related glaucoma with or without anterior segment dysgPathogenic★★★1.7860%3.508%
rs794726877IDUAMucopolysaccharidosis, MPS-I-S; Hurler syndrome; Mucopolysaccharidosis, MPS-I-H/S; Mucopolysaccharidosis Pathogenic★★★1.4710%2.899%
rs193922262GCKMonogenic diabetes; Maturity-onset diabetes of the young type 2Likely pathogenic★★★1.1570%2.287%
rs121909531ACTA1Alpha-actinopathy; Congenital myopathy 2c, severe infantile, autosomal dominant; Progressive scapulohumerPathogenic★★★1.0750%2.127%
rs28940868GAAGAA-related disorder; Cardiovascular phenotype; not provided; Glycogen storage disease, type IIPathogenic★★★1.0680%2.113%
rs786201058CDH1CDH1-related diffuse gastric and lobular breast cancer syndrome; Hereditary cancer-predisposing syndrome;Pathogenic★★★0.9191%1.821%
rs398123256IDUAMucopolysaccharidosis type 1; not provided; Hurler syndromePathogenic★★★0.8961%1.776%
rs727503788ACADVLnot provided; Very long chain acyl-CoA dehydrogenase deficiencyPathogenic★★★0.8079%1.603%
This result is a major QC finding, not a diagnosis. Thousands of rare pathogenic-array sites repeat at frequency steps near 0.176%, consistent with one/few threshold genotype calls across many loci. That burden is biologically implausible if treated as confirmed variants and strongly indicates that rare-site array calls require cluster-plot QC and orthogonal sequencing. This page shows only higher-review, q<5% rows.
ClinVar classifications are updated weekly and can change. Carrier calculations assume HWE and do not encode inheritance, penetrance or compound heterozygosity. Source: NCBI ClinVar.