Population Genomic Interpretationv2.5.1
PLINK-recalculated aggregate J&K data · research/learning only

Evidence based Interpretation of SNPs effect at Population Level

700,604SNP frequency rows
19,524K16 SNPs harmonized

00 — Regional context

Jammu & Kashmir: geography, community and genomic variation

A population-frequency view of the Indian Union Territory—not a definition of any person, community or identity.

Place

A connected Himalayan region

Jammu & Kashmir spans subtropical plains and foothills, the Pir Panjal, and the temperate Kashmir Valley. Elevation, valleys, passes and seasonal accessibility have shaped where people settled and how communities connected, while long-distance trade and migration linked the region with the wider Indian subcontinent and Central and West Asia.

Government of J&K UT landscape profile · Jammu Division · Kashmir Division

People & culture

Diversity is the starting point

The region includes multiple languages, faiths, livelihoods and cultural traditions. Geography can reduce gene flow, while migration increases it; marriage within communities can amplify founder effects and genetic drift. These processes alter some allele frequencies, but they do not create discrete biological “races,” and culture or identity cannot be read from a SNP.

Government of J&K UT tourism and culture · J&K Official Languages Act, 2020 · J&K–India rare-disease review

Population genetics

Frequencies reflect history and sampling

South Asian genomes lie on overlapping ancestry clines shaped by ancient mixture, later mobility and population-specific drift. ANI and ASI are historical statistical constructs—not present-day peoples. A regional average can therefore be informative for research while concealing substantial variation among localities, families and individuals.

Indian population structure, Nature (2009) · Jammu & Kashmir–India mitogenome study

Regional reference scope: introductory sources cover the present Union Territory of Jammu & Kashmir, India, its Jammu and Kashmir administrative divisions, or population research explicitly identified as J&K–India. No Pakistan-administered/PoJK population dataset is used as a J&K population source. The PJL row in the later 1000 Genomes comparison is a standard Lahore reference comparator—not a J&K dataset.
What JKDNA offers here

A privacy-preserving community frequency atlas

This page analyzes aggregate SNP frequencies: rsID/name and GRCh37 position matching, haplogroup-marker evidence, reference-component projections, pharmacogenomic markers, disease-associated markers and a ClinVar quality audit. The published output contains no names, sample identifiers, pedigrees or individual genotypes.

With explicit informed consent, a participant’s own genotype could be run through the same auditable marker pipeline and compared with the community reference to create a personal educational report. Population frequencies are useful priors, not a substitute for that person’s genotype, clinical history, ancestry model or confirmatory testing.

Community-drivenBroader voluntary participation can improve representativeness and expose sampling gaps.
Data-minimizedPublic summaries should remain aggregate; consent, withdrawal and controlled raw-data access belong upstream.
Careful interpretationNo identity assignment, diagnosis, prescribing or reproductive decision should be made from this atlas.

01 — Ancestry & lineage

Expanded haplogroup marker evidence

rsID/name matching first, GRCh37 position fallback second; alleles are strand-checked before use.

Y-DNA lineage marker signals

median derived-marker frequency %

mtDNA candidate lineage signals

median diagnostic-marker frequency %
L3
N
R
U (parent marker inconsistency)
22.85%
The U7 child-marker signal is conservatively reported at U because its intervening parent-marker evidence is inconsistent. M4/M18 child signals are grouped as an M minimum; R5 and R8 remain direct-node signals. Values are nested and must not be summed.
Y-tree nodesignalusable/coveredspreadevidence
R1a31.43%2/30.08%supportive marker concordance
J214.11%3/41.33%strong marker concordance
H1a1a12.27%2/32.22%supportive marker concordance
R2a11.00%3/32.15%strong marker concordance
Q9.56%9/102.41%strong marker concordance
L1a7.78%4/51.96%strong marker concordance
J2b7.11%3/32.48%strong marker concordance
O24.12%4/42.27%strong marker concordance
R1b3.67%7/71.03%strong marker concordance
C2b1a2.55%4/42.43%strong marker concordance
G2a2.43%8/102.27%strong marker concordance
J12.08%2/22.20%supportive marker concordance
T2.03%6/72.46%strong marker concordance
N11.34%10/112.25%strong marker concordance
I11.21%12/212.26%strong marker concordance
O (parent marker inconsistency)6.05%23/27nearest supported ancestral grouping; child marker(s): O1b1a2b
G (parent marker inconsistency)2.90%60/77nearest supported ancestral grouping; child marker(s): G2b2a
I (parent marker inconsistency)2.17%55/79nearest supported ancestral grouping; child marker(s): I2
A0 (parent marker inconsistency)1.97%7/12nearest supported ancestral grouping; child marker(s): A0000
mtDNA nodesignalusable/coveredspreadevidence
U (parent marker inconsistency)22.85%2/3nearest supported ancestral grouping; child marker(s): U, U7
M (parent marker inconsistency)≥5.27%3/4child-marker-supported ancestral minimum; child marker(s): M4, M18
R51.88%3/3strong marker concordance
R81.24%2/4supportive marker concordance
Y-DNA frequency recovery: all 2,643/2,643 supplied Y-manifest rows matched; 2,373 have callable aggregate frequencies. Because FAM sex codes were unavailable, X heterozygosity and Y call rate were used to identify the Y-callable cluster; cluster separation 2.81 SD. Only Y-callable homozygous/haploid calls contributed to Y allele frequencies; strand-ambiguous palindromic markers were excluded from lineage estimates.
Lineage interpretation boundary: each value is a population-level diagnostic-marker signal. When an isolated child marker exceeds the supported parent path, the result is moved to the nearest plausible ancestor and labelled (parent marker inconsistency); extreme isolated signals are excluded as probable array/strand artifacts. Parent and descendant nodes overlap and must not be summed. These are not individual haplogroup calls, a terminal distribution, biological sex, ethnicity or identity. Low-frequency and single-marker calls require sequencing confirmation.
mtDNA expansion: all 1,138 supplied array rows matched (1,138 unique database sites; 539 polymorphic). Concordant child evidence is retained, but an inconsistent parent path is reported at the nearest defensible ancestral node. The ≥ symbol marks a child-marker-supported minimum rather than a complete clade frequency. These are population marker signals—not individual calls or a distribution that sums to 100. Source: HaploGrep PhyloTree 17.

02 — Within-cohort comparison

Comparative population-subgroup parameters

Only labels passing the private minimum-size rule are included; denominators and participant counts are intentionally not displayed.

Aggregate frequency-space PCA

subgroup centroids, not individuals

K16 grouped reference components

model similarity %, stacked
Population codeautosomal call rateobserved heterozygosityF proxynearest centroidRMS allele-frequency Δclosest South Asian reference
JK-199.14%30.84%+0.033JK-30.0453PJL
JK-298.55%31.00%+0.028JK-10.0460PJL
JK-399.02%30.99%+0.024JK-10.0453PJL
JK-498.93%31.80%-0.026JK-20.0511PJL
JK-599.15%31.32%-0.013JK-10.0531PJL
JK-699.83%30.92%-0.020JK-30.0777PJL
JK-799.48%31.40%-0.021JK-10.0649PJL
JK-899.75%31.16%-0.019JK-10.0656PJL
JK-998.00%31.01%-0.012JK-10.0895BEB
JK-1099.50%31.23%-0.026JK-10.0680PJL
JK-1197.92%32.38%-0.061JK-10.0681PJL
JK-1299.71%30.00%-0.019JK-90.1000BEB
JK-1397.81%29.55%+0.047JK-10.1032BEB
Comparative account: the strongest East-Eurasian-related model signals occur in JK-12, JK-9, JK-13; the strongest West-Eurasian-related model signals occur in JK-10, JK-4, JK-2. The closest pairwise frequency centroid includes JK-1 and JK-3.
Within-group diversity: the descriptive heterozygosity F proxy ranges from -0.061 (JK-11) to +0.047 (JK-13). Negative values indicate excess observed heterozygosity relative to within-group allele frequencies; positive values indicate a deficit. It is not a pedigree or clinical inbreeding estimate.
Interpretation boundary: JK-1, JK-2 and subsequent codes anonymize the qualifying covariate strata; the source-label mapping is restricted and is not published. These codes are not identities or rankings. PCA, distances, K16 components and South Asian coefficients describe aggregate frequency-space similarity only. Differences may reflect recruitment, disease ascertainment, geography, relatedness, array batch or uneven representation. No individual classification or subgroup health-risk claim should be made from this section.

03 — Indian population comparison

J&K within modern South Asian frequency space

Common, non-palindromic autosomal SNP frequencies harmonized by rsID, allele and genome build.

3,935complete 1000G SAS markers
3,671GenomeIndia overlaps
2.22%J&K vs pooled GenomeIndia mean |ΔAF|
3.62%five-reference fit RMSE

Best-fitting modern-reference combination

descriptive coefficient %

Aggregate population-frequency PCA

centroids, not individuals
Code1000 Genomes populationfit coefficientchromosome sensitivitymean |ΔAF|RMS ΔAF
PJLPunjabi in Lahore, Pakistan40.9%40.2%–42.6%3.25%4.46%
GIHGujarati Indian in Houston, USA25.7%20.7%–26.7%3.82%5.06%
BEBBengali in Bangladesh23.2%22.6%–24.0%3.73%5.08%
ITUIndian Telugu in the United Kingdom5.6%5.0%–6.5%3.93%5.31%
STUSri Lankan Tamil in the United Kingdom4.6%3.9%–6.2%3.98%5.38%
GenomeIndia overall comparison: 3,671 harmonized markers; mean absolute frequency difference 2.22%, RMS difference 2.98%. The public 9,768-genome archive is pooled across India, so it is used as a national frequency centroid—not as an ancestry source or an 83-group map. Source: GenomeIndia/IBDC.
Affinity, not ancestry fractions. The five coefficients are the constrained combination of PJL, GIH, BEB, ITU and STU population-average frequencies that best reconstructs the aggregate J&K vector. These modern populations are themselves related and admixed; coefficients do not identify genealogy, community, ANI, ASI or AASI. The PCA contains population centroids only and cannot describe variation among J&K individuals. Sources: Ensembl population-frequency API · 1000 Genomes Phase 3.

04 — Global reference projection

Continental-scale and ANI-related components

J&K frequencies projected onto a real 16-component SNPweights panel.

Grouped ancestry-related components

model proportion

K16 components

model proportion
Grouped componentEstimatechromosome sensitivity*
ANI-specific51.5%49.8%–53.1%
West Eurasian-related27.9%26.8%–29.0%
East Eurasian-related15.1%13.9%–16.2%
African-related4.0%3.0%–4.9%
Oceania/Americas-related1.4%0.7%–2.2%
Unresolved ancestor0.1%0.0%–0.6%
K16 componentEstimatesensitivity*
ANI51.5%49.8%–53.1%
Caucasian-HG17.5%16.0%–19.0%
SouthEastAsian10.5%9.6%–11.4%
EHG5.9%4.7%–7.0%
Siberian4.5%3.6%–5.4%
WHG-UHG4.0%3.3%–4.7%
Subsaharian3.6%2.9%–4.3%
Australian1.0%0.3%–1.6%
Neolithic0.6%0.0%–1.4%
EastAfrican0.4%0.0%–1.2%
Oceanic0.2%0.0%–0.8%
Amerindian0.2%0.0%–1.1%
Ancestor0.1%0.0%–0.6%
Arctic0.1%0.0%–0.9%
NearEast0.0%0.0%–0.0%
NorthAfrcian0.0%0.0%–0.6%
Model coverage: 19,524/116,446 panel SNPs harmonized (10,212 direct; 9,312 complement). Projection uses SNPweights normalization and M/M′ correction. *Delete-one-autosome sensitivity intervals are not cohort sampling CIs. Method.
Do not call the remainder “ASI.” K16 has an ANI-labelled component but no AASI/ASI reference. Modern models separate AASI-, Iranian-related and Steppe-related ancestry. These labels are model similarities, not literal genealogy.

05 — Pharmacogenomics

Drug-wise marker distribution and metabolism signal

Genotype frequencies expanded from allele frequency under Hardy-Weinberg; phenotype categories are a single-marker simplification (see caveats) — not full diplotype/CNV-based clinical calling.

Share of population with a non-normal phenotype, by marker

%
abacavir
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
HLA-B (rs2395029)response risk marker6.0%Tag negative 88.4%, Tag positive 11.6%11.6%
azathioprine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
TPMT (rs1800460)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
NUDT15 (rs116855232)no function7.5%Normal Metabolizer 85.6%, Intermediate Metabolizer 13.9%, Poor Metabolizer 0.6%14.4%
capecitabine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
DPYD (rs3918290)no function1.8%Normal Metabolizer 96.5%, Intermediate Metabolizer 3.5%, Poor Metabolizer 0.0%3.5%
clopidogrel
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function2.5%Normal Metabolizer 95.0%, Intermediate Metabolizer 4.9%, Poor Metabolizer 0.1%5.0%
CYP2C19 (rs12248560)increased function16.6%Normal Metabolizer 69.6%, Rapid Metabolizer 27.7%, Ultrarapid Metabolizer 2.8%30.4%
fluorouracil
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
DPYD (rs3918290)no function1.8%Normal Metabolizer 96.5%, Intermediate Metabolizer 3.5%, Poor Metabolizer 0.0%3.5%
mercaptopurine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
TPMT (rs1800460)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
NUDT15 (rs116855232)no function7.5%Normal Metabolizer 85.6%, Intermediate Metabolizer 13.9%, Poor Metabolizer 0.6%14.4%
omeprazole
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function2.5%Normal Metabolizer 95.0%, Intermediate Metabolizer 4.9%, Poor Metabolizer 0.1%5.0%
CYP2C19 (rs12248560)increased function16.6%Normal Metabolizer 69.6%, Rapid Metabolizer 27.7%, Ultrarapid Metabolizer 2.8%30.4%
phenytoin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C9 (rs1057910)decreased function12.6%Normal Metabolizer 76.4%, Normal-Intermediate Metabolizer 22.0%, Intermediate Metabolizer 1.6%23.6%
CYP2C9 (rs1799853)decreased function5.6%Normal Metabolizer 89.0%, Normal-Intermediate Metabolizer 10.7%, Intermediate Metabolizer 0.3%11.0%
simvastatin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
SLCO1B1 (rs4149056)decreased function3.7%Normal Metabolizer 92.8%, Normal-Intermediate Metabolizer 7.0%, Intermediate Metabolizer 0.1%7.2%
tacrolimus
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP3A5 (rs776746)no function75.6%Poor Metabolizer 57.1%, Intermediate Metabolizer 36.9%, Normal Metabolizer 6.0%94.0%
thioguanine
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
TPMT (rs1142345)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
TPMT (rs1800460)no function1.4%Normal Metabolizer 97.3%, Intermediate Metabolizer 2.7%, Poor Metabolizer 0.0%2.7%
NUDT15 (rs116855232)no function7.5%Normal Metabolizer 85.6%, Intermediate Metabolizer 13.9%, Poor Metabolizer 0.6%14.4%
voriconazole
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C19 (rs4986893)no function2.5%Normal Metabolizer 95.0%, Intermediate Metabolizer 4.9%, Poor Metabolizer 0.1%5.0%
warfarin
Gene (variant)FunctionAlt freqPhenotype distributionNon-normal
CYP2C9 (rs1057910)decreased function12.6%Normal Metabolizer 76.4%, Normal-Intermediate Metabolizer 22.0%, Intermediate Metabolizer 1.6%23.6%
CYP2C9 (rs1799853)decreased function5.6%Normal Metabolizer 89.0%, Normal-Intermediate Metabolizer 10.7%, Intermediate Metabolizer 0.3%11.0%
Simplification notice. Phenotype categories here are derived from a single diagnostic SNP per gene under a simplified function model (no-function / decreased-function / increased-function relative to reference). Real clinical PGx phenotyping (esp. for CYP2D6, which needs copy-number detection) combines multiple star alleles per haplotype. Treat this as population-level screening signal, not individual diplotype calls.

06 — Disease Risk

Population-attributable risk summary

Levin's formula combines each lead SNP's population risk-allele frequency with its published odds ratio into a share of disease burden statistically attributable to that variant, at the population level.

Combined PAR by disease

%
Prostate Cancer Moderate population-level contribution
16.4% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs6983267POU5F1B51.7%1.209.4%
rs10993994MSMB55.8%1.157.7%
Type 2 Diabetes Moderate population-level contribution
15.6% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs7903146TCF7L231.1%1.4011.1%
rs5219KCNJ1138.0%1.145.1%
Asthma Moderate population-level contribution
14.9% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs7216389GSDMB40.3%1.3010.8%
rs20541IL1330.6%1.164.7%
Breast Cancer Moderate population-level contribution
13.7% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs2981582FGFR236.3%1.268.6%
rs3803662TOX329.4%1.205.6%
Alzheimer's Disease Moderate population-level contribution
10.6% combined population-attributable risk across 1 lead SNP
SNPGeneRisk allele freqORPAR
rs6733839BIN159.0%1.2010.6%
Rheumatoid Arthritis Low population-level contribution
6.5% combined population-attributable risk across 2 lead SNPs
SNPGeneRisk allele freqORPAR
rs6910071HLA-DRB17.2%1.805.4%
rs2476601PTPN222.2%1.501.1%
Coronary Artery Disease Low population-level contribution
0.5% combined population-attributable risk across 1 lead SNP
SNPGeneRisk allele freqORPAR
rs10455872LPA0.9%1.600.5%
What PAR does and doesn't mean. A high combined PAR means the included risk alleles are common AND meaningfully increase relative risk in the population — it is not a statement about any individual's personal risk, and ignores gene-environment interaction, polygenic background beyond the listed lead SNPs, and LD/epistasis between loci. Allele frequencies shown are from the supplied J&K database; odds ratios are drawn from the general published GWAS literature for each locus, not J&K-specific studies.

07 — Pathogenic susceptibility

ClinVar cross-reference and quality triage

Official GRCh37 ClinVar P/LP SNVs intersected by position and allele, requiring ≥1 review star.

10,689P/LP frequency cross-references
12practice-guideline reviewed
1,538expert-panel reviewed
127effect alleles not observed
7high-frequency QC flags

Prioritized carrier-frequency estimates

HWE marker-level estimate %
VariantGeneCondition(s)ClinVarReviewAllele qHWE carrier
rs75039782CFTRCFTR-related disorder; Cystic fibrosis; not specified; not provided; Bronchiectasis with or without elevaPathogenic★★★★1.5714%3.093%
rs78655421CFTRCFTR-related disorder; Respiratory ciliopathies including non-CF bronchiectasis; Cystic fibrosis; CongeniPathogenic★★★★1.3788%2.720%
rs75096551CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; Cystic fibrosis; Congenital bilateral aplasia of Pathogenic★★★★1.3689%2.700%
rs80224560CFTRCFTR-related disorder; Cystic fibrosis; Respiratory ciliopathies including non-CF bronchiectasis; not spePathogenic★★★★1.3571%2.677%
rs113993959CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; Congenital bilateral aplasia of vas deferens fromPathogenic★★★★1.2950%2.556%
rs75961395CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; not provided; Congenital bilateral aplasia of vasPathogenic★★★★1.2894%2.546%
rs77010898CFTRCFTR-related disorder; Cystic fibrosis diagnostic test; not specified; Congenital bilateral aplasia of vaPathogenic★★★★1.2160%2.402%
rs75527207CFTRCFTR-related disorder; Congenital bilateral aplasia of vas deferens from CFTR mutation; Cystic fibrosis; Pathogenic★★★★1.1511%2.276%
rs77188391CFTRCFTR-related disorder; not specified; Cystic fibrosis; Congenital bilateral aplasia of vas deferens from Pathogenic★★★★0.7184%1.426%
rs78756941CFTRCFTR-related disorder; Respiratory ciliopathies including non-CF bronchiectasis; Cystic fibrosis; CongeniPathogenic★★★★0.4323%0.861%
rs74767530CFTRCFTR-related disorder; Congenital bilateral aplasia of vas deferens from CFTR mutation; Cystic fibrosis; Pathogenic★★★★0.3608%0.719%
rs74551128CFTRnot provided; Cystic fibrosis; Hereditary pancreatitis; Bronchiectasis with or without elevated sweat chlPathogenic★★★★0.2878%0.574%
rs1800553ABCA4Syndromic retinitis pigmentosa; Autosomal recessive ABCA4-related disorders; ABCA4-related retinopathy; APathogenic★★★3.0980%6.004%
rs28936700CYP1B1Anterior segment dysgenesis 6; Glaucoma 3A; CYP1B1-related glaucoma with or without anterior segment dysgPathogenic★★★2.8331%5.506%
rs786201058CDH1CDH1-related diffuse gastric and lobular breast cancer syndrome; Hereditary cancer-predisposing syndrome;Pathogenic★★★2.4481%4.776%
rs121909531ACTA1Alpha-actinopathy; Congenital myopathy 2c, severe infantile, autosomal dominant; Progressive scapulohumerPathogenic★★★2.4052%4.695%
rs398123256IDUAMucopolysaccharidosis type 1; not provided; Hurler syndromePathogenic★★★2.1866%4.278%
rs727503788ACADVLnot provided; Very long chain acyl-CoA dehydrogenase deficiencyPathogenic★★★1.9795%3.881%
This result is a major QC finding, not a diagnosis. Thousands of rare pathogenic-array sites repeat at frequency steps near 0.176%, consistent with one/few threshold genotype calls across many loci. That burden is biologically implausible if treated as confirmed variants and strongly indicates that rare-site array calls require cluster-plot QC and orthogonal sequencing. This page shows only higher-review, q<5% rows.
ClinVar classifications are updated weekly and can change. Carrier calculations assume HWE and do not encode inheritance, penetrance or compound heterozygosity. “Not observed” is an aggregate frequency result, not an individual normal genotype. Browse every screened row in the Marker Atlas. Source: NCBI ClinVar.