Privacy Design®
Knowledge Base → Privacy engineering → Genetic Data Re-identification Risk

Genetic Data Re-identification Risk

ℹ️ Informational — not legal or clinical advice. Informational decision-support, not legal or clinical advice; re-identification risk is context-dependent and thresholds are the paper's illustrative values.

As of 2026-09-10. Framework adapted (CC BY 4.0, attributed) from Thomas M, Mackes N, Preuss-Dodhy A, Wieland T, Bundschus M. JMIR Bioinform Biotechnol 2024. PMC11165293. Thresholds are the paper's illustrative values; re-identification risk is context-dependent — a “lower” result is not a safety guarantee.

A guided self-assessment of re-identification risk in genetic datasets, based on the nine dataset features identified by Thomas et al. (2024) — biological modality, assay, data format, germline vs somatic, SNP content, short tandem repeats, aggregate measures, rare variants and structural variants. Answer the guiding questions for the features relevant to your data and the tool returns a qualitative risk profile (not a numeric score), surfacing the paper’s thresholds (e.g. >~20 SNPs, >~10 STRs, germline ≫ somatic).

Informational — not legal or clinical advice. Re-identification risk is context-dependent and compounds with linkage/auxiliary data; a “lower” result is not a guarantee of non-identifiability. Framework adapted with attribution under CC BY 4.0. See also [Digital Health] · [PETs / anonymisation].

Field guide

The 9 features (overview)

General features (rough estimate of privacy-critical information)

#FeatureWhat raises riskThresholds
1Biological modalityThe type of molecular data (e.g. DNA sequence, RNA, DNA methylation, protein), which determines whether sequence or sensitive attributes can be read or inferred.No full identity-tracing attack starting from data other than DNA sequence has been demonstrated yet.
2Experimental assayThe method used to generate the data, which governs how rich the information is and how much of the genome is covered.Most published privacy attacks used whole-genome sequencing or commercial SNP microarrays; data from common DTC-GT methods target the same variants and are more likely to be in public databases.
3Data format or level of processingThe processing state of the data (raw, semi-processed, or highly processed), where less-processed formats often carry extra exploitable information.Raw or low-processed data often contain information not of primary interest that can be exploited for re-identification attacks.
4Germline versus somatic variation contentWhether the variants are heritable germline variants (present in every cell, passed to offspring) or acquired somatic variants (tissue-specific), the latter being far lower risk.No identity-tracing, inference, or membership attack based on somatic variation has been published; somatic variation can currently be considered low risk.

Higher-risk components (demonstrated privacy attacks)

#FeatureWhat raises riskThresholds
5SNPs (single nucleotide polymorphisms)Common germline variants (present in >1% of the population) that are the most privacy-critical component, since knowing an individual's state at a modest number of independent SNPs can uniquely identify them.Knowing 30-80 statistically independent SNPs can suffice for identification; data-sanitization is recommended for any data set containing >20 SNPs.
6STRs (short tandem repeats)Highly variable repetitive DNA regions used in forensics and genealogy, so identifiable that a small number of loci can single out an individual.Knowing the repeat numbers of as few as 10-30 STRs can suffice for identification; data directly or indirectly containing >10 STR loci could be considered identifiable.
7Aggregated sample measuresVariables aggregated across samples (e.g. SNP/allele frequencies, odds ratios, association-study summary statistics) that enable membership attacks even though no identity-tracing attack has been shown.Can enable membership attacks revealing demographic, genetic, and phenotypic information; no identity-tracing attack based on aggregate data has been demonstrated yet.

Lower-risk components (no attack demonstrated yet)

#FeatureWhat raises riskThresholds
8Rare SNVs (single nucleotide variants)Single nucleotide variants present in <1% of the population, currently low risk because no identity-tracing, completion, or inference attack on them has been published.No identity-tracing, completion, or inference attack on rare SNVs has been published; most DTC-GT providers do not detect or remove them, so they can currently be viewed as low risk.
9Structural variantsLarge-scale genomic variations such as deletions, duplications, and copy number variations (CNVs), currently low risk as no privacy attack based on them has been demonstrated.A privacy attack based on CNVs or any other structural variant remains to be demonstrated; risk can currently be considered low but growing public databases should be monitored.

Source

Thomas M, Mackes N, Preuss-Dodhy A, Wieland T, Bundschus M. JMIR Bioinform Biotechnol 2024. PMC11165293 — https://pmc.ncbi.nlm.nih.gov/articles/PMC11165293/ · licensed CC BY 4.0. Adapted with attribution.