Genetic Data Re-identification Risk
ℹ️ Informational — not legal or clinical advice. Informational decision-support, not legal or clinical advice; re-identification risk is context-dependent and thresholds are the paper's illustrative values.
As of 2026-09-10. Framework adapted (CC BY 4.0, attributed) from Thomas M, Mackes N, Preuss-Dodhy A, Wieland T, Bundschus M. JMIR Bioinform Biotechnol 2024. PMC11165293. Thresholds are the paper's illustrative values; re-identification risk is context-dependent — a “lower” result is not a safety guarantee.
A guided self-assessment of re-identification risk in genetic datasets, based on the nine dataset features identified by Thomas et al. (2024) — biological modality, assay, data format, germline vs somatic, SNP content, short tandem repeats, aggregate measures, rare variants and structural variants. Answer the guiding questions for the features relevant to your data and the tool returns a qualitative risk profile (not a numeric score), surfacing the paper’s thresholds (e.g. >~20 SNPs, >~10 STRs, germline ≫ somatic).
Informational — not legal or clinical advice. Re-identification risk is context-dependent and compounds with linkage/auxiliary data; a “lower” result is not a guarantee of non-identifiability. Framework adapted with attribution under CC BY 4.0. See also [Digital Health] · [PETs / anonymisation].
Field guide
- Risk level — each answer maps to lower / elevated / higher. A feature’s risk = the highest among its answered questions; the overall band = higher if any feature is higher, else elevated if any is elevated, else lower.
- Groups — the paper’s grouping: general features (1–4) · higher-risk components (5–7) · lower-risk components (8–9).
- Thresholds — the paper’s illustrative values (e.g. >~20 SNPs, >~10 STRs, germline ≫ somatic), not hard legal limits.
- Leave features/questions blank where they don’t apply — unanswered features don’t affect the result.
Assessment
Answer the guiding questions for the features relevant to your dataset; leave others blank. The result is a qualitative risk profile, not a score.
The 9 features (overview)
General features (rough estimate of privacy-critical information)
| # | Feature | What raises risk | Thresholds |
|---|---|---|---|
| 1 | Biological modality | The type of molecular data (e.g. DNA sequence, RNA, DNA methylation, protein), which determines whether sequence or sensitive attributes can be read or inferred. | No full identity-tracing attack starting from data other than DNA sequence has been demonstrated yet. |
| 2 | Experimental assay | The method used to generate the data, which governs how rich the information is and how much of the genome is covered. | Most published privacy attacks used whole-genome sequencing or commercial SNP microarrays; data from common DTC-GT methods target the same variants and are more likely to be in public databases. |
| 3 | Data format or level of processing | The processing state of the data (raw, semi-processed, or highly processed), where less-processed formats often carry extra exploitable information. | Raw or low-processed data often contain information not of primary interest that can be exploited for re-identification attacks. |
| 4 | Germline versus somatic variation content | Whether the variants are heritable germline variants (present in every cell, passed to offspring) or acquired somatic variants (tissue-specific), the latter being far lower risk. | No identity-tracing, inference, or membership attack based on somatic variation has been published; somatic variation can currently be considered low risk. |
Higher-risk components (demonstrated privacy attacks)
| # | Feature | What raises risk | Thresholds |
|---|---|---|---|
| 5 | SNPs (single nucleotide polymorphisms) | Common germline variants (present in >1% of the population) that are the most privacy-critical component, since knowing an individual's state at a modest number of independent SNPs can uniquely identify them. | Knowing 30-80 statistically independent SNPs can suffice for identification; data-sanitization is recommended for any data set containing >20 SNPs. |
| 6 | STRs (short tandem repeats) | Highly variable repetitive DNA regions used in forensics and genealogy, so identifiable that a small number of loci can single out an individual. | Knowing the repeat numbers of as few as 10-30 STRs can suffice for identification; data directly or indirectly containing >10 STR loci could be considered identifiable. |
| 7 | Aggregated sample measures | Variables aggregated across samples (e.g. SNP/allele frequencies, odds ratios, association-study summary statistics) that enable membership attacks even though no identity-tracing attack has been shown. | Can enable membership attacks revealing demographic, genetic, and phenotypic information; no identity-tracing attack based on aggregate data has been demonstrated yet. |
Lower-risk components (no attack demonstrated yet)
| # | Feature | What raises risk | Thresholds |
|---|---|---|---|
| 8 | Rare SNVs (single nucleotide variants) | Single nucleotide variants present in <1% of the population, currently low risk because no identity-tracing, completion, or inference attack on them has been published. | No identity-tracing, completion, or inference attack on rare SNVs has been published; most DTC-GT providers do not detect or remove them, so they can currently be viewed as low risk. |
| 9 | Structural variants | Large-scale genomic variations such as deletions, duplications, and copy number variations (CNVs), currently low risk as no privacy attack based on them has been demonstrated. | A privacy attack based on CNVs or any other structural variant remains to be demonstrated; risk can currently be considered low but growing public databases should be monitored. |
Source
Thomas M, Mackes N, Preuss-Dodhy A, Wieland T, Bundschus M. JMIR Bioinform Biotechnol 2024. PMC11165293 — https://pmc.ncbi.nlm.nih.gov/articles/PMC11165293/ · licensed CC BY 4.0. Adapted with attribution.