<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD 2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">EXCLI J</journal-id>
      <journal-title>EXCLI Journal</journal-title>
      <issn pub-type="epub">1611-2156</issn>
      <publisher>
        <publisher-name>Leibniz Research Centre for Working Environment and Human Factors</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">2021-4002</article-id>
      <article-id pub-id-type="doi">10.17179/excli2021-4002</article-id>
      <article-id pub-id-type="pii">Doc1243</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Review article</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Privacy considerations for sharing genomics data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Oestreich</surname>
            <given-names>Marie</given-names>
          </name>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Chen</surname>
            <given-names>Dingfan</given-names>
          </name>
          <xref ref-type="aff" rid="A2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Schultze</surname>
            <given-names>Joachim L.</given-names>
          </name>
          <xref ref-type="aff" rid="A1">1</xref>
          <xref ref-type="aff" rid="A3">3</xref>
          <xref ref-type="aff" rid="A4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Fritz</surname>
            <given-names>Mario</given-names>
          </name>
          <xref ref-type="aff" rid="A2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Becker</surname>
            <given-names>Matthias</given-names>
          </name>
          <uri content-type="orcid">https://orcid.org/0000-0002-7120-4508</uri>
          <xref ref-type="corresp" rid="COR1">&#x0002a;</xref>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="A1">
        <label>1</label>Systems Medicine, Deutsches Zentrum f&#xFC;r Neurodegenerative Erkrankungen (DZNE), Venusberg-Campus 1&#x2F;99, 53127 Bonn, Germany</aff>
      <aff id="A2">
        <label>2</label>CISPA Helmholtz Center for Information Security, Saarbr&#xFC;cken, Germany, Stuhlsatzenhaus 5, 66123 Saarbr&#xFC;cken, Germany</aff>
      <aff id="A3">
        <label>3</label>Genomics and Immunoregulation, Life &#x26; Medical Sciences (LIMES) Institute, University of Bonn, Bonn, Germany, Carl-Troll-Stra&#xDF;e 31, 53115 Bonn, Germany</aff>
      <aff id="A4">
        <label>4</label>PRECISE Platform for Single Cell Genomics and Epigenomics at Deutsches Zentrum f&#xFC;r Neurodegenerative Erkrankungen (DZNE) and the University of Bonn, Germany, Venusberg-Campus 1&#x2F;99, 53127 Bonn, Germany</aff>
      <author-notes>
        <corresp id="COR1">*To whom correspondence should be addressed: Matthias Becker, Systems Medicine, Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Venusberg-Campus 1/99, 53127 Bonn, Germany; Tel. +49 228 43302 642, E-mail: <email>Matthias.Becker@dzne.de</email></corresp>
      </author-notes>
      <pub-date pub-type="epub">
        <day>16</day>
        <month>07</month>
        <year>2021</year>
      </pub-date>
      <pub-date pub-type="collection">
        <year>2021</year>
      </pub-date>
      <volume>20</volume>
      <fpage>1243</fpage>
      <lpage>1260</lpage>
      <history>
        <date date-type="received">
          <day>19</day>
          <month>06</month>
          <year>2021</year>
        </date>
        <date date-type="accepted">
          <day>07</day>
          <month>07</month>
          <year>2021</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Copyright &#xA9; 2021 Oestreich et al.</copyright-statement>
        <copyright-year>2021</copyright-year>
        <license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
          <p>This is an Open Access article distributed under the terms of the Creative Commons Attribution Licence (http://creativecommons.org/licenses/by/4.0/) You are free to copy, distribute and transmit the work, provided the original author and source are credited.</p>
        </license>
      </permissions>
      <self-uri xlink:href="https://www.excli.de/vol20/excli2021-4002.pdf">This article is available from https://www.excli.de/vol20/excli2021-4002.pdf</self-uri>
      <abstract><p>An increasing amount of attention has been geared towards understanding the privacy risks that arise from sharing genomic data of human origin. Most of these efforts have focused on issues in the context of genomic sequence data, but the popularity of techniques for collecting other types of genome-related data has prompted researchers to investigate privacy concerns in a broader genomic context. In this review, we give an overview of different types of genome-associated data, their individual ways of revealing sensitive information, the motivation to share them as well as established and upcoming methods to minimize information leakage. We further discuss the concise threats that are being posed, who is at risk, and how the risk level compares to potential benefits, all while addressing the topic in the context of modern technology, methodology, and information sharing culture. Additionally, we will discuss the current legal situation regarding the sharing of genomic data in a selection of countries, evaluating the scope of their applicability as well as their limitations. We will finalize this review by evaluating the development that is required in the scientific field in the near future in order to improve and develop privacy-preserving data sharing techniques for the genomic context.</p></abstract>
      <kwd-group>
        <kwd>data privacy</kwd>
        <kwd>data sharing</kwd>
        <kwd>genomic data</kwd>
        <kwd>transcriptomic data</kwd>
        <kwd>epigenomic data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec sec-type="intro">
      <title>Introduction</title><p>Since the introduction of high throughput data generation methods, the medical field has become increasingly data-driven. Large biomedical studies often comprise sampling and data generation at different centers, requiring the data to be subsequently shared for joint analysis. Though data sharing has thus become crucial, the inherently private nature of medical data demands caution and elaborate sharing protocols are necessary to ensure patient privacy. This is particularly important when sharing genomics data and the associated medical metadata. The concise risks of privacy breaches and the methods used by adversaries vary for different data types. In this review, we will assess such risks and techniques in the context of genomic data, more precisely, three different data types: genomic sequences, transcriptomic and epigenomic data. We will further elaborate on currently used methods for privately sharing data along with different laws that are meant to protect individuals in case of a data leak. We will additionally introduce current research areas for private sharing of genomic data and give an outlook on the necessary as well as expected changes in the upcoming years.</p></sec>
    <sec>
      <title>Genomic Data Types</title><p>Traditionally, genomic data (Goodwin et al., 2016[<xref ref-type="bibr" rid="R36">36</xref>]) refers to data holding information on the base sequence in an individual&#x27;s genome. Here, we are going to extend this notion of genomics data to also incorporate transcriptomic as well as epigenomic data, which are closely associated with and influenced by the genomic sequence (Figure 1a<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). This serves as a more holistic assessment of potential privacy breaches when working with either category of data by evaluating how these different data types impact and relate to one another and how weak points may translate across categories.</p><sec><title>Sequence data</title><p>Sequence data contains qualitative information, identifying specific bases and their positions along the genome. This data type describes an individual&#x27;s genomic sequence to a different extent, depending on the protocol underlying the data generation process. Whole genome sequencing (WGS) produces genomic sequences in their entirety, while partial genomic sequencing focuses only on particular areas, e.g., exon regions in case of whole exome sequencing (WES) (Mar&#xF3;ti et al., 2018[<xref ref-type="bibr" rid="R53">53</xref>]). Even more condensed is sequence data storing only single nucleotide polymorphisms (SNPs) (Robert and Pelletier, 2018[<xref ref-type="bibr" rid="R63">63</xref>]), which are specific sites in the genome for which different allele frequencies have been observed between populations. Despite its strongly reduced dimensionality, SNP data often allows for accurate genomic fingerprinting of individuals (Lin et al., 2004[<xref ref-type="bibr" rid="R50">50</xref>]).</p></sec><sec><title>Transcriptomic data</title><p>Unlike sequence data, transcriptomic data is of a quantitative nature. It stores information on the abundance of RNA transcripts in a sample, mostly gene transcripts in the form of mRNA, but also other types such as long non-coding RNAs or microRNAs. Transcript abundances are directly influenced by external factors such as environmental stressors, pathogens, or drugs but are also closely impacted by the underlying genomic sequences. Genomic mutations can change a gene&#x27;s susceptibility for transcription and are therefore indirectly reflected in the transcriptomic landscape. The genomic locations that are associated with variation in transcript abundances are referred to as expression quantitative trait loci (eQTLs) (Nica and Dermitzakis, 2013[<xref ref-type="bibr" rid="R56">56</xref>]). Transcriptomic data can be collected at different levels of resolution: more coarse grained data is measured using bulk-RNA-sequencing, where transcript abundances are measured across many cells, which is in contrast to the higher resolution of single-cell RNA-sequencing, where transcripts are quantified per cell (Tang et al., 2009[<xref ref-type="bibr" rid="R70">70</xref>]). </p></sec><sec><title>Epigenomic data</title><p>Epigenomic data holds information on heritable alterations that affect gene expression but are not based on changes in the genomic sequence itself (Weinhold, 2006[<xref ref-type="bibr" rid="R77">77</xref>]). These alterations typically take place on the DNA or on histone proteins that are responsible for binding and condensing the DNA, therefore impacting its accessibility for transcription. One important modification on the DNA itself is DNA methylation, which commonly occurs on cytosine bases that are followed by a guanine, so called CpG dinucleo-tides. These often occur in dense clusters referred to as CpG islands. In general, DNA methylation is associated with a reduction of gene expression whereas the absence of methylation promotes RNA transcription. Histone modifications are post-translational alterations on histone proteins that affect the histone&#x27;s ability to remodel the chromatin. They include - among others - methylation and acetylation on amino acid rests and impact how densely the histone can adhere to the DNA (Cazaly et al., 2019[<xref ref-type="bibr" rid="R11">11</xref>]). There are other levels of epigenomics data that are determined by sequencing, for example, so-called open chromatin regions, regions that can be accessed by transcription factors to regulate gene transcription. Out of these types of epigenetic data, in this review, we will focus on DNA methylation data, given that it may be directly impacted by changes in the genomic sequence.</p></sec></sec>
    <sec>
      <title>Why Share Genomic Data?</title><p>In order to assess privacy issues that arise when sharing genomic data, there is a need to discuss the motivation behind sharing it in the first place (Figure 1b<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). </p><p>Data sharing is required during multi-party studies, where separate datasets have been collected in a decentralized fashion and subsequently have to be joined together. Further, sharing is required to improve the reproducibility of results by publicly (or on request) providing the dataset that was used in a particular analysis. A set of principles targeting the findability, accessibility, interoperability and reusability of data, particularly by machines, was introduced by Wilkinson et al. as the FAIR-principles (Wilkinson et al., 2016[<xref ref-type="bibr" rid="R78">78</xref>]). Other scenarios that require data sharing are the reusing of data, in which case datasets are shared for further use in order to answer questions other than those originally posed when collecting the data, or sharing for reduction of model bias. The latter is especially important when the data was collected from minority groups, in which case public accumulation can battle common bias problems in population-based modeling approaches (Mehrabi et al., 2019[<xref ref-type="bibr" rid="R54">54</xref>]).</p><p>A rather recent development that urges the distribution of genomic data is the increasing application of machine learning techniques (Libbrecht and Noble, 2015[<xref ref-type="bibr" rid="R49">49</xref>]). These approaches typically require a large amount of training data, especially in the context of deep learning (Eraslan et al., 2019[<xref ref-type="bibr" rid="R24">24</xref>]), as otherwise the model tends to bias towards local minima that generalize poorly on unseen samples. The demand for big data is particularly strong in genomics because of its immense feature space, and oftentimes exceeds the scope of a single study by far, prompting the collection of data from multiple studies. </p><p>Moreover, researchers (Brittain et al., 2017[<xref ref-type="bibr" rid="R10">10</xref>]; Dand et al., 2019[<xref ref-type="bibr" rid="R18">18</xref>]) and politicians (The White House, 2015[<xref ref-type="bibr" rid="R71">71</xref>]; BMBF, 2018[<xref ref-type="bibr" rid="R8">8</xref>]) alike are promoting the advent of personalized medicine to enable the development of treatment plans tailored to the individual. To provide the knowledge base for personalized medicine, a great wealth of biomedical data, including genomics data, will have to be generated and shared, given that the sampling will have to be conducted in a decentralized fashion to properly reflect human diversity and the data has to be accumulated for subsequent analysis. Given the highly private nature of the data, this naturally demands a thorough discussion on how to guarantee sufficient privacy compliance.</p></sec>
    <sec>
      <title>Privacy Concerns when Sharing Data</title><p>In light of the aforementioned reasons to share genomic data and before introducing concrete data sharing methods, the problems that arise from privacy breaches in the context of genomics must be elucidated. The root of these problems and the reason why genomic data privacy requires thorough discussion is the issue of subject re-identification. Identifying the individual which the data was taken from can potentially have severe consequences starting with social issues such as stigmatization, which is often observed when there is knowledge about the presence of certain risk alleles in an individual&#x27;s genome, especially in the context of mental health issues (Ward et al., 2019[<xref ref-type="bibr" rid="R75">75</xref>]). Additionally, known preconditions and increased disease risks can negatively affect chances for employment or health insurance (Godard et al., 2003[<xref ref-type="bibr" rid="R34">34</xref>]). There have also been cases where adopted children or children received by use of sperm donations identified their biological parent(s) due to publicly available genomic information (Erlich and Narayanan, 2014[<xref ref-type="bibr" rid="R26">26</xref>]). Another issue is posed by the longevity of genomic data: given the heritability of genomic and epigenomic alterations, re-identification does not only pose a risk to the individual itself but also to close relatives (Ayday et al., 2015[<xref ref-type="bibr" rid="R3">3</xref>]; Oprisanu et al., 2019[<xref ref-type="bibr" rid="R58">58</xref>]). While all living relatives can - in theory - be asked to give their consent to the collection of the data in addition to the consent given by the individual, privacy breaches may also impact unborn relatives or underage relatives whose consent cannot be given. </p><p>The risk of re-identification and the corresponding instantiated attack methods are different for the previously introduced data types. We will first discuss these with respect to sequence data and we will see that most re-identification risks associated with transcriptomic and epigenomic data fall back onto techniques that infer the underlying sequence information. </p><sec><title>Sequence data</title><p>One particular privacy complication often discussed in the context of sequence data is SNP &#x201C;barcoding&#x201D; or &#x201C;fingerprinting&#x201D;. This refers to the observation that a limited number of SNPs - 30 to 80 independent SNPs according to (Lin et al., 2004[<xref ref-type="bibr" rid="R50">50</xref>]) - is sufficient to unequivocally identify an individual, therefore providing a genome-based fingerprint of that person. Following this concept, an attacker can construct a SNP fingerprint-library from all publicly available sequence datasets and identify individuals - given their DNA - by checking for matches with fingerprints of publicly deposited sequences, which are often associated with study metadata that reveals some of the individual&#x27;s medical information (Figure 1c<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). This scenario raises the question of how an attacker would obtain the victim&#x27;s DNA sequence in the first place. Some examples here are DNA theft or careless publication of sequence data. The former is often illustrated by means of the coffee-cup example, where a malicious attacker acquires a person&#x27;s DNA by collecting it from a thrown-away coffee cup (G&#xFC;rsoy et al., 2020[<xref ref-type="bibr" rid="R37">37</xref>]). There are of course many other ways to steal DNA, be it from a tossed-out tissue, blood, or other sources as long as they contain intact cells. To then go from the sample to the DNA sequence is easier than oftentimes assumed, nowadays, where companies have specialized in direct-to-consumer individual genome sequencing for a moderate price, often only requiring as much as a saliva sample (Eissenberg, 2017[<xref ref-type="bibr" rid="R22">22</xref>]). The privacy concerns that accompany SNP fingerprinting translate further onto other types of genomic data, if the data allows the inference of the underlying SNPs.</p><p>A prominent example of the careless publication of private sequencing data occurred during the Personal Genome Project. Here, a group from Harvard College demonstrated that many of the uploaded, publicly available files contain personal information, such as the individual&#x27;s first and last name that was part of the file identifiers after decompression (Sweeney et al., 2013[<xref ref-type="bibr" rid="R69">69</xref>]).  </p><p>But re-identification using sequence data does not always have to be part of a malicious attack. A rather recent example is the apprehension of the &#x201C;Golden State Killer&#x201D;, who was caught several decades after his first known crimes due to the use of consumer genetic databases. Using the criminal&#x27;s sequence data and the data available in those databases, the investigators reconstructed several family trees by pairing the database information with public records, ultimately leading to the apprehension through distant relatives who had uploaded their genomic sequences (Zabel, 2019[<xref ref-type="bibr" rid="R82">82</xref>]). Though this case of re-identification helped to catch a serial killer, it also explicitly emphasizes privacy concerns in the context of shared genomic sequences.</p><p>There have been cases of deliberate sequence data publication in which particularly sensitive parts of the genome were removed to prevent the individual&#x27;s risk assessment with respect to certain diseases. A well-known example is the publication of James Watson&#x27;s whole genome with the exception of the APOE-gene. The gene has been associated with late onset Alzheimer&#x27;s disease - information that was not meant to be revealed. Shortly thereafter, Nyholt et al. (2009[<xref ref-type="bibr" rid="R57">57</xref>]) demonstrated that the removal of the gene alone does not necessarily prevent the accurate prediction of the disease risk posed by different alleles of the gene, since these were shown to be in linkage disequilibrium with other SNPs which had not been removed.</p><p>In cases where publicly available sequence data is available in combination with corresponding meta information, an individual&#x27;s sequence does not necessarily have to be acquired to re-identify the individuals in the dataset. This is due to the personal information contained in the meta information. The risk posed by metadata has been significantly mitigated by means of legal regulations, however, free-form text, longitudinal data, low sample size (El Emam, 2011[<xref ref-type="bibr" rid="R23">23</xref>]), and non-random generation of accession numbers (Erlich and Narayanan, 2014[<xref ref-type="bibr" rid="R26">26</xref>]) have been shown to remain problematic. </p><p>A thorough review of the requirements for secure genomic sequence sharing, storing, and testing methods is provided by Ayday et al. (2015[<xref ref-type="bibr" rid="R3">3</xref>]).</p></sec><sec><title>Transcriptomic data</title><p>Already in 2012, Schadt et al. demonstrated how to infer an individual&#x27;s genotype at eQTL positions and therefore to SNP-fingerprint the individual using a Bayesian approach that solely relied on publicly available data on putative <italic>cis</italic>-eQTLs and RNA expression data (Schadt et al., 2012[<xref ref-type="bibr" rid="R65">65</xref>]). They pointed out that this is particularly problematic given that gene expression data is commonly assumed to allow too little insight into a study participant&#x27;s genomic sequence to reveal their identity and has therefore been shared publicly on platforms such as ArrayExpress and Gene Expression Omnibus, while sequence data is held under controlled access. However, if contrary to common belief, gene expression data does allow SNP-fingerprinting, then the publicly available data for building a SNP-fingerprint library for re-identification attacks mentioned in the sequence section above is expanded from publicly accessible sequence data to both sequence data and gene expression data, resulting in a much larger pool of individuals at risk of re-identification. The paper was reviewed by Erlich and Narayanan (2014[<xref ref-type="bibr" rid="R26">26</xref>]) who concluded that the threat posed to individuals whose gene expression data has been published is low. This conclusion was based on the fact that the Bayesian method introduced by Schadt et al. only performed well when the eQTL data and gene expression data were measured on the same platform, while at the time both sequencing on microarray and Next Generation Sequencing (NGS) were common, making many of the datasets incompatible for the inference approach. In 2017, however, Lowe et al. (2017[<xref ref-type="bibr" rid="R52">52</xref>]) showed that the use of RNA-seq surpassed that of microarrays in 2016, showing increasing trends for RNA-seq while microarray publications decreased further in count. Corchete et al. (2020[<xref ref-type="bibr" rid="R15">15</xref>]) referred to RNA-seq as the &#x201C;first choice in transcriptomic analysis&#x201D; in 2020. Thus, while being reasonable at the time, the argument that platform heterogeneity hinders phenotype inference in many cases progressively loses its validity when the sequence data landscape becomes increasingly homogeneous. Erlich and Narayanan further argued that the amount of public gene expression data to be downloaded and processed by an adversary to generate the SNP barcodes requires immense computational power rendering it an unlikely endeavor. Putting this into today&#x27;s perspective, with the currently available hardware, new compute architectures (Becker et al., 2020[<xref ref-type="bibr" rid="R5">5</xref>]), and cloud space for rent, this argument as well requires reassessment. This emphasizes the need for re-evaluation of the re-identification risk posed by the method proposed by Schadt et al. considering the technological changes that have occurred since its original publication. The computational feasibility of SNP-fingerprinting entire databases such as GEO would enable an attacker to test whether a given individual - e.g., crime suspect, victim of DNA theft, person with public genome sequence - is part of a study in which the gene expression data has been published. Gene expression data could then also be utilized for associating individuals with certain traits such as BMI, sex, age, insulin levels and glucose levels (Schadt et al., 2012[<xref ref-type="bibr" rid="R65">65</xref>]) or to match individuals of one study with those of another, creating cross-links that potentially reveal further meta information on an individual. These genotype-phenotype linking attacks have been assessed in more detail by (Harmanci and Gerstein, 2016[<xref ref-type="bibr" rid="R40">40</xref>]), urging to use methods to quantify private information leakage before publishing a dataset.</p></sec><sec><title>Epigenomic data</title><p>The methylation of DNA is measured using bisulfite conversion. Here, DNA is treated with sodium bisulfite which does not impact methylated Cytosines (C) but converts unmethylated ones into Uracil (U). Adenine (A), Guanine (G) and Thymine (T) are not affected. C&#x2F;T-SNPs, i.e. SNPs where a C was changed to a T, are the most common transitions in human genomes (LaBarre et al., 2019[<xref ref-type="bibr" rid="R47">47</xref>]) and they can change a CpG dinucleotide into a TpG dinucleotide. These SNPs often induce a three-tier pattern in the measured methylation according to homozygous TpG individuals with low, heterozygous CpG&#x2F;TpG individuals with medium and homozygous CpG individuals with high methylation signals. Similar effects could be observed with C&#x2F;A or C&#x2F;G mutations (Philibert et al., 2014[<xref ref-type="bibr" rid="R60">60</xref>]; Daca-Roszak et al., 2015[<xref ref-type="bibr" rid="R17">17</xref>]). Based on these tri-modal patterns, LaBarre et al. developed a method that recognizes C&#x2F;T-SNPs in methylation data and subsequently removes them. However, the three-tier patterns could also result from differential methylation rather than underlying genotypes, and SNPs in CpGs that are always unmethylated are unlikely to be detected, since it is not possible to distinguish between a C&#x2F;T-SNP and a C that is converted to a T during bisulfite conversion. </p><p>In 2019, Hagestedt et al. showcased an effective membership inference attack on DNA methylation data, i.e., inferring the presence of an individual in a genomic dataset, disclosing the privacy risks posed by their public release (Hagestedt et al., 2019[<xref ref-type="bibr" rid="R38">38</xref>]). They further introduced MBeacon, a platform where researchers can query DNA methylation data deposited at Beacon sites with regard to whether or not they contain samples with specific methylation. MBeacon was designed using a privacy-by-design approach, substantially decreasing the success of membership inference attacks while maintaining a decent level of data utility.</p><p>Besides the potential inference of the underlying genotype, another privacy concern was raised by Philibert et al. (2014[<xref ref-type="bibr" rid="R60">60</xref>]). They claim that the methylation status of an individual at specific sites could be used to infer the smoking status and alcohol consumption of the individual. A response was issued in 2015 by Joly et al. (2015[<xref ref-type="bibr" rid="R44">44</xref>]), members of the International Human Epigenetic Consortium (IHEC), criticizing that the prediction of smoking status was not accurate enough to give reliable results and that the risk of re-identification is low in the absence of access to the corresponding genomic sequence information. Dyke, together with other authors that also participated in the response letter of Joly et al., issued a study that year which illustrates the inference of SNPs from DNA-methylation data using imbalances in methylation signal from forward and reverse strands when the data was collected using strand-specific Whole Genome Bisulfite Sequencing (WGBS) (Dyke et al., 2015[<xref ref-type="bibr" rid="R21">21</xref>]). They emphasize that the risk of re-identification is low but may be increased in combination with meta information on the individual&#x27;s demographic or health status, especially in the case of rare disease phenotypes. To maximize the privacy of participants in DNA-methylation studies, the authors ask for a more synchronized metadata vocabulary, in order to avoid different levels of information being revealed by different authors based on the descriptors they used. They also recommend that CpGs overlapping known SNPs are removed from DNA-methylation data prior to publication and offer a list of questions to consider to protect people with rare diseases in particular.</p><p>Berrang et al. provided a privacy risk assessment spanning different types of genomic data along a temporal axis and between related individuals (Berrang et al., 2018[<xref ref-type="bibr" rid="R6">6</xref>]). They demonstrated a Bayesian framework that utilizes correlations present between different data types to infer the values of one data type, using the information from another, e.g., inferring the methylation pattern from sequence data and vice versa. </p></sec></sec>
    <sec>
      <title>Laws and Limitations</title><p>Many countries have developed sophisticated laws and guidelines to protect individuals from misuse of their genomic information (Figure 1d<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). However, the wording is often ambiguous towards the different types of genomic data and has therefore prompted researchers to question their validity in particular situations. </p><p>For example, the German &#x201C;Gendiagnostikgesetz&#x201D; prohibits discrimination of citizens based on their genetic characteristics. Genetic characteristics are here defined as inherited or between conception and birth acquired hereditary information of human origin (Deutscher Bundestag, 2009[<xref ref-type="bibr" rid="R19">19</xref>]). This unambiguously covers genomic sequence information as it is present at birth. It is unclear, however, if this also covers quantitative gene expression information which can be affected by the environment and not only by underlying sequence information. Further, if it covers epigenomic data, which is often altered by environmental factors and therefore can be lost or acquired after birth and is not necessarily inherited. Lastly, if it also refers to genomic sequence information that has changed after birth, e.g., mutations due to external stimuli such as UV-radiation. The law also states that it does not apply to genetic examination and analyses and the handling of genetic samples and data for the purpose of research. The Council of Europe had previously issued a convention in 1997 that prohibits &#x201C;any form of discrimination against a person on grounds of his or her genetic heritage&#x201D;, where the term genetic heritage is not defined further. This convention was signed and ratified by 29 countries (Council of Europe, 1997[<xref ref-type="bibr" rid="R16">16</xref>]). In the United States, a similar effort has been made to protect citizens from discrimination in health insurance and employment based on genetic information by introducing the Genetic Information Nondiscrimination Act (GINA) (U.S. Equal Employment Opportunity Commission, 2008[<xref ref-type="bibr" rid="R73">73</xref>]). Genetic information is here defined as information on an individual&#x27;s genetic test, genetic tests of their family members and diseases or disorders that have manifested in family members of the individual. A genetic test is defined as &#x201C;an analysis of human DNA, RNA, chromosomes, proteins or metabolites, that detects genotypes, mutations, or chromosomal changes&#x201D;. This definition leaves the same questions unanswered as above. Additionally, it explicitly excludes the U.S. military from the list of employers that are prohibited to use genetic information as well as any employer with less than 15 employees. </p><p>In addition to GINA, disclosure of health information is regularized by the Health Insurance Portability and Accountability Act (HIPAA). It provides three standards for the disclosure of patient health data that do not require authorization by the patient. Those three standards are the Safe Harbor standard, the Limited Dataset standard, and the statistical standard (El Emam, 2011[<xref ref-type="bibr" rid="R23">23</xref>]). Safe Harbor regulates the de-identification of health data by removing 18 different identifying elements, among those the name, certain geographic information, all elements of dates except for the year, phone and fax numbers, e-mail addresses, social security numbers, and many more (El Emam, 2011[<xref ref-type="bibr" rid="R23">23</xref>]). The Limited Dataset standard only removes 16 potential identifiers but additionally requests a data sharing agreement between data custodian and data recipient and the statistical standard requires expert evaluation and classification of the de-identification risk as very small (El Emam, 2011[<xref ref-type="bibr" rid="R23">23</xref>]). Given its clarity and simplicity, Safe Harbor is often referred to for data de-identification. However, as pointed out by El Emam, not only does it often result in the removal of information that could have been useful for the data evaluation, it also does not sufficiently ensure the protection of the individual with regard to re-identification. In this context, there is special emphasis on the lack of protection through genetic data, such as SNPs, longitudinal data, widely used diagnosis codes, sampling size, and free-form text. A more detailed elaboration can be found in El Emam (2011[<xref ref-type="bibr" rid="R23">23</xref>]).</p><p>In 2018, major advancements have been made in the EU with respect to data protection due to the release of the General Data Protection Regulation, GDPR (European Union, 2018[<xref ref-type="bibr" rid="R30">30</xref>]; Shabani and Borry, 2018[<xref ref-type="bibr" rid="R66">66</xref>]). It applies in all EU member states and aims to unify data protection across countries. GDPR overcomes the ambiguous definition of genetic data as outlined above by defining it as &#x201C;personal data relating to the inherited or acquired genetic characteristics of a natural person which give unique information about the physiology or the health of that natural person and which result, in particular, from an analysis of a biological sample from the natural person in question&#x201D; (&#x201C;Recital 34 - Genetic data - GDPR.eu,&#x201D; 2018[<xref ref-type="bibr" rid="R33">33</xref>]). This definition includes the wide range of modern genomics data types and therefore offers more thorough protection from open sharing and un-consented processing than the regulations discussed above. While passed in the EU, the law applies worldwide if the data that is processed was sampled from a citizen of an EU member state.</p><p>Additional EU guidelines have been introduced recently that specifically aim to protect EU citizens from threats posed by Artificial Intelligence (AI) systems. The &#x201C;Assessment List for Trustworthy Artificial Intelligence&#x201D; (ALTAI) was published by the European Commission in 2020 as a guideline for the development of trustworthy AI (European Commission, 2020[<xref ref-type="bibr" rid="R28">28</xref>]). Early in 2021, the European Commission additionally published a draft of the Artificial Intelligence Act (AIA) (European Commission, 2021[<xref ref-type="bibr" rid="R29">29</xref>]). Subject to this act is any provider of AI applications worldwide, if those applications are used by EU citizens. It categorizes AI systems into different categories of threat, connected with strict obligations that have to be met by the provider. These regulations also address AI used in the health context, therefore including genomic data. While the motivation behind these guidelines and regulations is most reasonable, there are some concerns regarding the implications they might have on data privacy. This is particularly important because they demand the models to be fair and unbiased by using representative, non-discriminatory and complete training, validation and test datasets. However, in order to test a model&#x27;s fulfillment of these requirements, the inspecting authority is likely to need access to the datasets, which - e.g., in the context of genomic data - are often highly private.</p></sec>
    <sec>
      <title>Risk-Benefit Considerations</title><p>When assessing privacy risks, it is always essential to weigh the risk of an individual against the possible gain that is associated with the vulnerable position the individual finds themself in. The group of individuals that was - and in many scenarios still is - the main target groups for genomic data collection are those that have a personal reason to participate in the respective studies, e.g., a difficult-to-treat or poorly researched disease (Esplin et al., 2014[<xref ref-type="bibr" rid="R27">27</xref>]). For these individuals, the potential gain that comes from participating in these studies often substantially outweighs the risk of re-identification. But the focus of the target group appears to gradually shift, people have started to have their genome sequenced out of curiosity rather than acute medical reasons and the field of personalized medicine advocates genomic data collection to become part of a medical care routine (Brittain et al., 2017[<xref ref-type="bibr" rid="R10">10</xref>]; Suwinski et al., 2019[<xref ref-type="bibr" rid="R68">68</xref>]). Therefore, the group at risk of re-identification is bound to change from those that have a high benefit-risk ratio to a more heterogeneous population. Additionally, the risk of cross-referencing genomic data and metadata to narrow down the set of individuals a data instance potentially belongs to, is likely to increase due to oversharing of personal information on social media, be it voluntary or involuntary. This can start with information as subtle as height, sex, and weight which can be inferred from pictures, geotags, and dates that put an individual into close spatial and temporal proximity of the conduction of a given study, and it can go as far as people openly sharing their health status or study participation. This can be expected to substantially increase the risk of re-identification and it is an issue that has lacked thorough attention in prior risk-evaluation strategies.</p></sec>
    <sec>
      <title>Established Data Sharing Techniques</title><p>Current data sharing techniques come at different levels of security, as is illustrated in Figure 1e<xref ref-type="fig" rid="F1">(Fig. 1)</xref>. As discussed above, some data types are often shared publicly in plain text without restrictions other than de-identified sample descriptors. This is the case when the risk of re-identification for the participants is considered minimal. Genomic data with increased re-identification risk such as sequence data or data that gives direct information on partial sequences (reads, SNPs) are published under controlled access. In this scenario, the legitimacy of an access request is evaluated based on the applicant&#x27;s personal information and the research project the data is intended to be used for. The use of the data also often comes with a series of constraints that regard a safe storing location, no sharing and no re-identification attempts. While there is still a lack of true oversight with respect to whether or not the data is shared after downloading, another controlled access strategy is to not allow the data to be downloaded but instead run protocolled queries on the data and only retrieve the results. This may severely restrict the flexibility with which the data can be analyzed (Erlich and Narayanan, 2014[<xref ref-type="bibr" rid="R26">26</xref>]). </p><p>To decrease the risk of subject identification, be it in publicly shared or controlled-access data, efforts are made to de-identify the data, i.e., to remove information or reduce its granularity such that identification of the individual becomes very unlikely. A common set of guidelines for de-identification is the mentioned Safe Harbor standard included in the HIPAA Privacy Rule. </p><p>G&#xFC;rsoy et al. developed a method that sanitizes the raw reads underlying gene expression data such that sharing with reduced re-identification risk is possible, while keeping the necessary data manipulation minimal (G&#xFC;rsoy et al., 2020[<xref ref-type="bibr" rid="R37">37</xref>]). They achieve this by transforming the original BAM file into a sanitized file, where information that reveals the presence of a variant (SNPs, insertions, deletions) is masked. For instance, information on variants as it is contained in reads is removed by replacing the called base at the site of the variant with that present in the reference genome. The true values of the sanitized elements are stored in a separate file which is meant to be under controlled access. This allows for sharing of the sanitized data while being able to reconstruct the original if access to the additional file is granted, though no formal or statistical guarantees on privacy are provided.</p><p>Classical encryption approaches have been leveraged as well to enable secure sharing of genomics data. One example is the crypt4gh file format introduced in 2019 by the Global Alliance for Genomics and Health (GA4GH) (&#x201C;GA4GH File Encryption Standard,&#x201D; 2019[<xref ref-type="bibr" rid="R32">32</xref>]). The format allows the genomic data to be encrypted while in storage, in transit, during reading and writing. It uses symmetric encryption as well as public-key encryption, it is confidential in the sense that it is only readable by holders of a secret decryption key, but it does not obscure the length of the file. The data is stored in blocks, the integrity of which is ensured using message authentication codes, however, blocks can be rearranged, removed, or added. Authentication of files encrypted with this method is not provided. Though as with any system that relies on secret decryption keys, privacy is lost in the case of a system breach that results in an untrusted party acquiring the key. </p><p>Others have explored fully homomorphic encryption (FHE) to securely operate on genomic data. FHE allows to conduct computations on the genomic data while it remains in its encrypted state, receiving the encrypted results and decrypting them with a personal key (Erlich and Narayanan, 2014[<xref ref-type="bibr" rid="R26">26</xref>]). This allows for secure computations on cloud systems even if the system itself is not. While this approach was long assumed too computationally expensive to be reasonable in the context of genomic data, Blatt et al. recently presented an improvement in run-time of FHE for Genome Wide Association Studies (GWAS) by introducing parallelization and crypto-engineering optimizations, which allegedly outperforms secure multiparty computations (SMPC) (Blatt et al., 2020[<xref ref-type="bibr" rid="R7">7</xref>]). SMPC is another approach to secure computation, in which two or more parties that hold private data can compare and perform computations on the data without ever revealing the data itself to the other party or another third party (Erlich and Narayanan, 2014[<xref ref-type="bibr" rid="R26">26</xref>]). While a prominent point of criticism with SMPC models is the oftentimes extensive communication required between computing parties, Cho et al. demonstrated an approach for SMPC in GWAS in which run-time scaled linearly to the number of samples and was reduced to 80 days for 1 million individuals and 500,000 SNPs (Cho et al., 2018[<xref ref-type="bibr" rid="R14">14</xref>]). Though this is still a substantial amount of time, efforts such as this have worked on optimizing the procedure in the past years. However, in the case of Cho et al., the privacy guarantee only holds for the semi-honest security model, in which participants are assumed to not deviate from the conduction protocol. </p><p>Besides homomorphic encryption and multi-party computing, there are also hardware-based approaches to handling sensitive information such as genomic data. An example is the Intel Software Guard Extension (SGX), which has also been used in combination with homomorphic encryption on GWAS data (Sadat et al., 2019[<xref ref-type="bibr" rid="R64">64</xref>]). While such techniques based on SGX can also be used to protect model and&#x2F;or data (Hanzlik et al., 2021[<xref ref-type="bibr" rid="R39">39</xref>]), implementations of SGX however have been troubled with serious security issues themselves (Van Bulck et al., 2018[<xref ref-type="bibr" rid="R74">74</xref>]; Lipp et al., 2021[<xref ref-type="bibr" rid="R51">51</xref>]).</p><p>Another widely applied concept to allow privacy preserving data analysis is that of Differential Privacy (DP) (Dwork and Roth, 2014[<xref ref-type="bibr" rid="R20">20</xref>]). The idea behind DP is to allow private data analysis by assuring that the addition or removal of a subject to or from the dataset does not significantly alter potential query results and therefore does not disclose whether or not the individual is part of the dataset. To achieve this, different levels of noise have to be added to the data before its release, where the amount of noise necessary increases when the sample size of a dataset decreases. The privacy loss of the noised dataset is quantified using the epsilon parameter, where a value of 0 indicates total privacy, though decreasing values come with increasing added noise and therefore less utility. In contrast to security measures, cryptography or privacy heuristics, DP comes with strong guarantees that are not susceptible to misuse of the system or typical breaches. Although privacy preserving and statistical analyses do not follow an adverse goal, today&#x27;s techniques are often affected by a reduced utility of the analysis. For an exhaustive review of privacy-enhancing technologies in the genomic context, please refer to the work done by Mittos et al. (2019[<xref ref-type="bibr" rid="R55">55</xref>]).</p></sec>
    <sec>
      <title>Current Research and Future Vision</title><sec><title>Sharing-free solutions</title><p>In the machine learning domain, increasing interest has been focused on developing solutions that do not require data to be shared but that store the data locally and have the model migrate instead, performing local training and only sharing updated model parameters (Figure 1f<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). The most prominent sharing-free learning framework is federated learning (FL), whose advantages and major open problems have been discussed thoroughly in recent works (Kairouz et al., 2019[<xref ref-type="bibr" rid="R45">45</xref>]; Rieke et al., 2020[<xref ref-type="bibr" rid="R62">62</xref>]). FL is a general learning paradigm that can be built on top of various training algorithms, with no restriction on the adopted model architecture, therefore offering a broad application spectrum spanning across diverse data modalities. There have been isolated use cases in the health sector, such as FL for predictions based on electronic health records (Brisimi et al., 2018[<xref ref-type="bibr" rid="R9">9</xref>]; Xu et al., 2021[<xref ref-type="bibr" rid="R80">80</xref>]), medical images (Sheller et al., 2019[<xref ref-type="bibr" rid="R67">67</xref>]; Li et al., 2020[<xref ref-type="bibr" rid="R48">48</xref>]) and compound-protein-binding data (Rieke et al., 2020[<xref ref-type="bibr" rid="R62">62</xref>]). However, a broad application has so far been hindered by technical obstacles such as the requirement for a somewhat standardized data format that the different sites adhere to, different hard- and software environments, and most importantly, related privacy issues. While FL advertises inherent data security by leaving the data in place, potential information leakages could originate in the transmission of model parameter updates and threats are posed by the untrusted central server that receives and combines the local model updates reported by each site. For example, untrusted servers are shown vulnerable to attacks that reconstruct raw user data from the parameter updates sent by local clients (Zhu and Han, 2020[<xref ref-type="bibr" rid="R83">83</xref>]). To minimize the potential privacy risk posed by a vulnerable centralized server, efforts have been made to explore fully decentralized (peer-to-peer) topologies where no central server is required. One completely new data-sharing free approach combining peer-to-peer functionality using blockchain technology with AI is swarm learning (SL) (Warnat-Herresthal et al., 2021[<xref ref-type="bibr" rid="R76">76</xref>]). In particular, SL addresses the issue of untrusted participants by registering and authorizing all participating sites via blockchain technology in a swarm network.</p></sec><sec><title>Synthetic data</title><p>Another ongoing attempt to boost patient privacy when sharing genomics data is the use of synthetic data instead of the original data. The idea behind this is to create a synthetic cohort that follows a similar distribution as the original cohort, producing accurate analytical results while protecting patient privacy by generating novel samples distinct from the original data (Figure 1f<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). In this regard, researchers have utilized the concept of generative modeling, applying for example Restricted Boltzmann Machines (RBMs) (Hinton et al., 2006[<xref ref-type="bibr" rid="R43">43</xref>]) or Generative Adversarial Networks (GANs) (Goodfellow et al., 2014[<xref ref-type="bibr" rid="R35">35</xref>]; Yelmen et al., 2021[<xref ref-type="bibr" rid="R81">81</xref>]). Though not with the goal to create entire synthetic datasets, variational autoencoders (VAEs) (Kingma and Welling, 2013[<xref ref-type="bibr" rid="R46">46</xref>]) have been used to impute missing counts in single-cell expression data (Eraslan et al., 2019[<xref ref-type="bibr" rid="R25">25</xref>]; Qiu et al., 2020[<xref ref-type="bibr" rid="R61">61</xref>]) as well as in bulk-RNA sequencing and methylome data (Qiu et al., 2020[<xref ref-type="bibr" rid="R61">61</xref>]). However, this research is still in its infancy and requires thorough assessment with respect to both the utility and the privacy properties of the generated data. Conventionally, generative models are evaluated along three fronts: (1) fidelity - whether the generated samples can faithfully represent real data; (2) diversity - whether the generated data is diverse enough to cover the variability of real data; and (3) generalization - whether the generated samples are merely copies of real data, i.e., the model overfits and memorizes training data (Alaa et al., 2021[<xref ref-type="bibr" rid="R2">2</xref>]). The privacy property started to be considered and investigated in recent works, mostly in the form of membership inference attacks. Specifically, Hayes et al. introduce membership inference attacks against GANs trained on image data (Hayes et al., 2019[<xref ref-type="bibr" rid="R41">41</xref>]; Hilprecht et al., 2019[<xref ref-type="bibr" rid="R42">42</xref>]) and a systematic analysis has been conducted by Chen et al. (2020[<xref ref-type="bibr" rid="R13">13</xref>]). They assess an attacker&#x27;s ability to infer the presence of a given sample in the GAN&#x27;s training set with respect to different threat models, dataset sizes, and GAN model architectures. They show that membership inference is facilitated if the training dataset is small, which they explain by the GAN&#x27;s inability to generalize and instead memorize. This could be an especially limiting factor in the generation of synthetic genomics data since the real data available for training is often limited, especially in the human context, emphasizing the need for privacy risk assessment prior to the public release of the synthetic data or trained models. </p><p>Recently, Oprisanu et al. extended the investigation of GANs&#x27; privacy properties to genomic sequence data (Oprisanu et al., 2021[<xref ref-type="bibr" rid="R59">59</xref>]). Generally, more research is necessary to improve the utility of generative models while avoiding information leakage and to deploy them to genomics data other than sequences. </p><p>Other work has focussed on private deep learning in general (Abadi et al., 2016[<xref ref-type="bibr" rid="R1">1</xref>]) and GANs in particular (Xie et al., 2018[<xref ref-type="bibr" rid="R79">79</xref>]; Beaulieu-Jones et al., 2019[<xref ref-type="bibr" rid="R4">4</xref>]; Frigerio et al., 2019[<xref ref-type="bibr" rid="R31">31</xref>]; Torkzadehmahani et al., 2019[<xref ref-type="bibr" rid="R72">72</xref>]) over the past years, among those a recent publication that addresses privacy issues when sharing sensitive data or when sharing generators trained on such data and proposes the training and release of differentially private generators instead (Chen et al., 2020[<xref ref-type="bibr" rid="R12">12</xref>]). They further illustrate how the approach can naturally adapt to a federated learning setting. The differential privacy is achieved by restraining the impact a single sample can have on the fully trained model by using differentially private stochastic gradient descent during the training procedure. They further tackle the utility-privacy tradeoff by only training the generator - which may subsequently be publicly released - in a differentially private manner while training the discriminator optimally, therefore not differentially private and discarding it afterwards. Additionally, they demonstrate the value of the approach in a decentralized learning context such as in federated learning approaches. One of the main benefits of a differentially private generator and the obtained samples is that such synthetic data can be accessed without further privacy costs, while, in contrast, differentially private analysis of data needs to keep track of the incurred privacy cost. Also, established tools can be used on such synthetic data, while conventional differential private analysis typically needs to adapt the whole toolchain.</p></sec><sec><title>Protecting privacy and preventing discrimination</title><p>Beyond these research-based approaches, more appropriate laws are required worldwide. There needs to be active exchange between the scientific community - particularly medical and life science researchers as well as scientists from the fields of data privacy and data security - and the lawmakers to ensure that the regulations are up to date with the scientific progress and that their phrasing is less ambiguous. The terminology as it is now is often outdated or kept too broad to be a good guidance for researchers and to successfully protect study participants. In a rapidly evolving field such as this, regular risk assessments and revision of laws based on correspondence with active researchers of the topic is inevitable. Clear and thorough laws against genome-based discrimination are particularly important since protecting an individual&#x27;s privacy can, in practice, oftentimes not be fully assured, simply due to the fact that technologies and methods available to attackers in the future can only be speculated about. Therefore, it is even more important to instantiate laws that explicitly prevent discrimination in the case of data leakage, to provide a fail-safe system in cases where privacy protecting measures fall short (Figure 1f<xref ref-type="fig" rid="F1">(Fig. 1)</xref>). First steps into this direction have already been made by means of the GDPR, ALTAI and AI Act as outlined above.</p></sec></sec>
    <sec sec-type="conclusions">
      <title>Conclusion</title><p>Putting the risks of re-identification using genomic sequence data into perspective, data privacy is a concern that needs to be taken seriously. At this point in time, while all the above-mentioned methods are eligible threats, the costs involved and the expertise needed are still rather high and one could wonder, how realistic the threat really is. But as always in the field of data security and privacy, it is not only important to assess what <italic>is</italic> but also what potentially <italic>will be</italic>. While today, the motivation to acquire the genomic sequence of most people is considerably low given the entailed costs and effort, further decrease in sequencing price, increased insight into the genome and better computing performance are likely to make it more attractive in the future. The future in mind, present day scientists are required to address the issues of genomic data privacy to assure responsible research. This entails active communication with lawmakers to provide non-discrimination laws that protect study participants in the case of data leakage and which are up to date with the science. Further, increased collaboration of data security and privacy researchers with life scientists is essential to develop privacy-preserving data sharing techniques that are specifically tailored to genomics data in the near future. In this context, we can expect an increased necessity for bioinformaticians, computational biologists, biomathematicians, and others to optimally communicate the needs of the genomics community to computer scientists, in order to enable easier, yet secure sharing of genomic data. </p></sec>
    <sec>
      <title>Acknowledgements</title><p>This work was supported by the HGF  Helmholtz AI grant Pro-Gene-Gen (ZT-I-PF-5-23), the HGF Incubator grant sparse2big (ZT-I-0007), HGF Incubator grant &#x201C;Trustworthy Federated Data Analytics (TFDA)&#x22; (ZT-I-OO1 4), by NaFoUniMedCovid19 (FKZ: 01KX2021, project acronym COVIM), by the German Research Foundation (DFG) (INST 37&#x2F;1049-1, INST 216&#x2F;981-1, INST 257&#x2F;605-1, INST 269&#x2F;768-1, INST 217&#x2F;988-1, INST 217&#x2F;577-1, INST 217&#x2F;1011-1, INST 217&#x2F;1017-1 and INST 217&#x2F;1029-1), ImmunoSep (grant 84722) and the BMBF-funded excellence project Diet-Body-Brain (DietBB) (grant 01EA1809A).</p><p>The figure was created with BioRender.com.</p></sec>
    <sec>
      <title>Conflict of interest</title><p>The authors declare that they have no conflict of interest. </p></sec>
  </body>
  <back>
    <ref-list>
      <ref id="R1">
        <label>1</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Abadi</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Chu</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Goodfellow</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>McMahan</surname>
              <given-names>HB</given-names>
            </name>
            <name>
              <surname>Mironov</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Talwar</surname>
              <given-names>K</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Deep learning with differential privacy</article-title>
          <year>2016</year>
          <conf-name>Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security - CCS&#x2019;16</conf-name>
          <publisher-loc>New York, NY</publisher-loc>
          <publisher-name>ACM Press</publisher-name>
          <fpage>308</fpage>
          <lpage>318</lpage>
        </citation>
      </ref>
      <ref id="R2">
        <label>2</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Alaa</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>van Breugel</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Saveliev</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>van der Schaar</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>How faithful is your synthetic data&#x3F; Sample-level metrics for evaluating and auditing generative models</article-title>
          <year>2021</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2102.08921">https://arxiv.org/abs/2102.08921</ext-link></comment>
        </citation>
      </ref>
      <ref id="R3">
        <label>3</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Ayday</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>DeCristofaro</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>Hubaux</surname>
              <given-names>J-P</given-names>
            </name>
            <name>
              <surname>Tsudik</surname>
              <given-names>G</given-names>
            </name>
          </person-group>
          <article-title>The chills and thrills of whole genome sequencing</article-title>
          <year>2015</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1306.1264">https://arxiv.org/abs/1306.1264</ext-link></comment>
        </citation>
      </ref>
      <ref id="R4">
        <label>4</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Beaulieu-Jones</surname>
              <given-names>BK</given-names>
            </name>
            <name>
              <surname>Wu</surname>
              <given-names>ZS</given-names>
            </name>
            <name>
              <surname>Williams</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Lee</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Bhavnani</surname>
              <given-names>SP</given-names>
            </name>
            <name>
              <surname>Byrd</surname>
              <given-names>JB</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Privacy-preserving generative deep neural networks support clinical data sharing</article-title>
          <source>Circ Cardiovasc Qual Outcomes</source>
          <year>2019</year>
          <volume>12</volume>
          <fpage>e005122</fpage>
        </citation>
      </ref>
      <ref id="R5">
        <label>5</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Becker</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Schultze</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Bresniker</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Singhal</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Ulas</surname>
              <given-names>T</given-names>
            </name>
            <name>
              <surname>Schultze</surname>
              <given-names>JL</given-names>
            </name>
          </person-group>
          <article-title>A novel computational architecture for large-scale genomics</article-title>
          <source>Nat Biotechnol</source>
          <year>2020</year>
          <volume>38</volume>
          <fpage>1239–41</fpage>
        </citation>
      </ref>
      <ref id="R6">
        <label>6</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Berrang</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Humbert</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Lehmann</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Eils</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Backes</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Dissecting privacy risks in biomedical data</article-title>
          <year>2018</year>
          <conf-name>2018 IEEE European Symposium on Security and Privacy (EuroS&#x26;P)</conf-name>
          <publisher-loc>New York, NY</publisher-loc>
          <publisher-name>IEEE</publisher-name>
          <fpage>62</fpage>
          <lpage>76</lpage>
        </citation>
      </ref>
      <ref id="R7">
        <label>7</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Blatt</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Gusev</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Polyakov</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Goldwasser</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>Secure large-scale genome-wide association studies using homomorphic encryption</article-title>
          <source>Proc Natl Acad Sci U S A</source>
          <year>2020</year>
          <volume>117</volume>
          <fpage>11608–13</fpage>
        </citation>
      </ref>
      <ref id="R8">
        <label>8</label>
        <citation citation-type="web">
          <collab>BMBF</collab>
          <article-title>Rahmenprogramm Gesundheitsforschung der Bundesregierung</article-title>
          <year>2018</year>
          <access-date>2021 Jun 19</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://www.bundesregierung.de/breg-de/service/publikationen/rahmenprogramm-gesundheitsforschung-der-bundesregierung-731062">https://www.bundesregierung.de/breg-de/service/publikationen/rahmenprogramm-gesundheitsforschung-der-bundesregierung-731062</ext-link></comment>
        </citation>
      </ref>
      <ref id="R9">
        <label>9</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Brisimi</surname>
              <given-names>TS</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Mela</surname>
              <given-names>T</given-names>
            </name>
            <name>
              <surname>Olshevsky</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Paschalidis</surname>
              <given-names>IC</given-names>
            </name>
            <name>
              <surname>Shi</surname>
              <given-names>W</given-names>
            </name>
          </person-group>
          <article-title>Federated learning of predictive models from federated Electronic Health Records</article-title>
          <source>Int J Med Inform</source>
          <year>2018</year>
          <volume>112</volume>
          <fpage>59–67</fpage>
        </citation>
      </ref>
      <ref id="R10">
        <label>10</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Brittain</surname>
              <given-names>HK</given-names>
            </name>
            <name>
              <surname>Scott</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Thomas</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>The rise of the genome and personalised medicine</article-title>
          <source>Clin Med</source>
          <year>2017</year>
          <volume>17</volume>
          <fpage>545–51</fpage>
        </citation>
      </ref>
      <ref id="R11">
        <label>11</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Cazaly</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>Saad</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Heckman</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Ollikainen</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Tang</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Making sense of the epigenome using data integration approaches</article-title>
          <source>Front Pharmacol</source>
          <year>2019</year>
          <volume>10</volume>
          <fpage>126</fpage>
        </citation>
      </ref>
      <ref id="R12">
        <label>12</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Chen</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Orekondy</surname>
              <given-names>T</given-names>
            </name>
            <name>
              <surname>Fritz</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>GS-WGAN: A gradient-sanitized approach for learning differentially private generators</article-title>
          <year>2020</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2006.08265">https://arxiv.org/abs/2006.08265</ext-link></comment>
        </citation>
      </ref>
      <ref id="R13">
        <label>13</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Chen</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Fritz</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <person-group person-group-type="editor">
            <name>
              <surname>Ligatti</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Ou</surname>
              <given-names>X</given-names>
            </name>
          </person-group>
          <article-title>GAN-Leaks: A taxonomy of membership inference attacks against generative models</article-title>
          <source>CCS &#x2019;20. Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security 2020</source>
          <year>2020</year>
          <conf-name>Virtual Event USA, November 9 - 13</conf-name>
          <fpage>343</fpage>
          <lpage>362</lpage>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://dl.acm.org/doi/10.1145/3372297.3417238">https://dl.acm.org/doi/10.1145/3372297.3417238</ext-link></comment>
        </citation>
      </ref>
      <ref id="R14">
        <label>14</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Cho</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Wu</surname>
              <given-names>DJ</given-names>
            </name>
            <name>
              <surname>Berger</surname>
              <given-names>B</given-names>
            </name>
          </person-group>
          <article-title>Secure genome-wide association analysis using multiparty computation</article-title>
          <source>Nat Biotechnol</source>
          <year>2018</year>
          <volume>36</volume>
          <fpage>547–51</fpage>
        </citation>
      </ref>
      <ref id="R15">
        <label>15</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Corchete</surname>
              <given-names>LA</given-names>
            </name>
            <name>
              <surname>Rojas</surname>
              <given-names>EA</given-names>
            </name>
            <name>
              <surname>Alonso-L&#xF3;pez</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>De Las Rivas</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Guti&#xE9;rrez</surname>
              <given-names>NC</given-names>
            </name>
            <name>
              <surname>Burguillo</surname>
              <given-names>FJ</given-names>
            </name>
          </person-group>
          <article-title>Systematic comparison and assessment of RNA-seq procedures for gene expression quantitative analysis</article-title>
          <source>Sci Rep</source>
          <year>2020</year>
          <volume>10</volume>
          <fpage>19737</fpage>
        </citation>
      </ref>
      <ref id="R16">
        <label>16</label>
        <citation citation-type="web">
          <collab>Council of Europe</collab>
          <article-title>Convention on Human Rights and Biomedicine. Convention for the protection of human rights and dignity of the human being with regard to the application of biology and medicine: convention on human rights and biomedicine</article-title>
          <year>1997</year>
          <access-date>2021 Jun 11</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://rm.coe.int/168007cf98">https://rm.coe.int/168007cf98</ext-link></comment>
        </citation>
      </ref>
      <ref id="R17">
        <label>17</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Daca-Roszak</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Pfeifer</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>&#x17B;ebracka-Gala</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Rusinek</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Szybi&#x144;ska</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Jarz&#x105;b</surname>
              <given-names>B</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Impact of SNPs on methylation readouts by Illumina Infinium HumanMethylation450 BeadChip Array: implications for comparative population studies</article-title>
          <source>BMC Genomics</source>
          <year>2015</year>
          <volume>16</volume>
          <fpage>1003</fpage>
        </citation>
      </ref>
      <ref id="R18">
        <label>18</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Dand</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Duckworth</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Baudry</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Russell</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Curtis</surname>
              <given-names>CJ</given-names>
            </name>
            <name>
              <surname>Lee</surname>
              <given-names>SH</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>HLA-C&#x2A;06:02 genotype is a predictive biomarker of biologic treatment response in psoriasis</article-title>
          <source>J Allergy Clin Immunol</source>
          <year>2019</year>
          <volume>143</volume>
          <fpage>2120–30</fpage>
        </citation>
      </ref>
      <ref id="R19">
        <label>19</label>
        <citation citation-type="journal">
          <collab>Deutscher Bundestag</collab>
          <article-title>Gesetz &#xFC;ber genetische Untersuchungen bei Menschen (Gendiagnostikgesetz &#x2013; GenDG)</article-title>
          <source>Jahrbuch f&#xFC;r Wissenschaft und Ethik</source>
          <year>2009</year>
          <volume>14</volume>
          <fpage>347</fpage>
          <lpage>361</lpage>
        </citation>
      </ref>
      <ref id="R20">
        <label>20</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Dwork</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Roth</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>The algorithmic foundations of differential privacy</article-title>
          <source>FNT Theorl Computer Sci</source>
          <year>2014</year>
          <volume>9</volume>
          <fpage>211–407</fpage>
        </citation>
      </ref>
      <ref id="R21">
        <label>21</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Dyke</surname>
              <given-names>SOM</given-names>
            </name>
            <name>
              <surname>Cheung</surname>
              <given-names>WA</given-names>
            </name>
            <name>
              <surname>Joly</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Ammerpohl</surname>
              <given-names>O</given-names>
            </name>
            <name>
              <surname>Lutsik</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Rothstein</surname>
              <given-names>MA</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Epigenome data release: a participant-centered approach to privacy protection</article-title>
          <source>Genome Biol</source>
          <year>2015</year>
          <volume>16</volume>
          <fpage>142</fpage>
        </citation>
      </ref>
      <ref id="R22">
        <label>22</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Eissenberg</surname>
              <given-names>JC</given-names>
            </name>
          </person-group>
          <article-title>Direct-to-consumer genomics: Harmful or empowering&#x3F; It is important to stress that genetic risk is not the same as genetic destiny</article-title>
          <source>Mo Med</source>
          <year>2017</year>
          <volume>114</volume>
          <issue>1</issue>
          <fpage>26–32</fpage>
        </citation>
      </ref>
      <ref id="R23">
        <label>23</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>El Emam</surname>
              <given-names>K</given-names>
            </name>
          </person-group>
          <article-title>Methods for the de-identification of electronic health records for genomic research</article-title>
          <source>Genome Med</source>
          <year>2011</year>
          <volume>3</volume>
          <fpage>25</fpage>
        </citation>
      </ref>
      <ref id="R24">
        <label>24</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Eraslan</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Avsec</surname>
              <given-names>&#x17D;</given-names>
            </name>
            <name>
              <surname>Gagneur</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Theis</surname>
              <given-names>FJ</given-names>
            </name>
          </person-group>
          <article-title>Deep learning: New computational modelling techniques for genomics</article-title>
          <source>Nat Rev Genet</source>
          <year>2019</year>
          <volume>20</volume>
          <fpage>389–403</fpage>
        </citation>
      </ref>
      <ref id="R25">
        <label>25</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Eraslan</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Simon</surname>
              <given-names>LM</given-names>
            </name>
            <name>
              <surname>Mircea</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Mueller</surname>
              <given-names>NS</given-names>
            </name>
            <name>
              <surname>Theis</surname>
              <given-names>FJ</given-names>
            </name>
          </person-group>
          <article-title>Single-cell RNA-seq denoising using a deep count autoencoder</article-title>
          <source>Nat Commun</source>
          <year>2019</year>
          <volume>10</volume>
          <fpage>390</fpage>
        </citation>
      </ref>
      <ref id="R26">
        <label>26</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Erlich</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Narayanan</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>Routes for breaching and protecting genetic privacy</article-title>
          <source>Nat Rev Genet</source>
          <year>2014</year>
          <volume>15</volume>
          <fpage>409–21</fpage>
        </citation>
      </ref>
      <ref id="R27">
        <label>27</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Esplin</surname>
              <given-names>ED</given-names>
            </name>
            <name>
              <surname>Oei</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Snyder</surname>
              <given-names>MP</given-names>
            </name>
          </person-group>
          <article-title>Personalized sequencing and the future of medicine: discovery, diagnosis and defeat of disease</article-title>
          <source>Pharmacogenomics</source>
          <year>2014</year>
          <volume>15</volume>
          <fpage>1771–90</fpage>
        </citation>
      </ref>
      <ref id="R28">
        <label>28</label>
        <citation citation-type="web">
          <collab>European Commission</collab>
          <article-title>Assessment List for Trustworthy Artificial Intelligence (ALTAI) for self-assessment &#x7C; Shaping Europe&#x2019;s digital future</article-title>
          <year>2020</year>
          <access-date>2021 Jun 18</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://digital-strategy.ec.europa.eu/en/library/assessment-list-trustworthy-artificial-intelligence-altai-self-assessment">https://digital-strategy.ec.europa.eu/en/library/assessment-list-trustworthy-artificial-intelligence-altai-self-assessment</ext-link></comment>
        </citation>
      </ref>
      <ref id="R29">
        <label>29</label>
        <citation citation-type="web">
          <collab>European Commission</collab>
          <article-title>Proposal for a Regulation laying down harmonised rules on artificial intelligence &#x7C; Shaping Europe&#x2019;s digital future</article-title>
          <year>2021</year>
          <access-date>2021 Jun 18</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://digital-strategy.ec.europa.eu/en/library/proposal-regulation-laying-down-harmonised-rules-artificial-intelligence">https://digital-strategy.ec.europa.eu/en/library/proposal-regulation-laying-down-harmonised-rules-artificial-intelligence</ext-link></comment>
        </citation>
      </ref>
      <ref id="R30">
        <label>30</label>
        <citation citation-type="web">
          <collab>European Union</collab>
          <article-title>GDPR Archives - GDPR.eu</article-title>
          <year>2018</year>
          <access-date>2021 Jun 7</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://gdpr.eu/tag/gdpr/">https://gdpr.eu/tag/gdpr/</ext-link></comment>
        </citation>
      </ref>
      <ref id="R31">
        <label>31</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Frigerio</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>de Oliveira</surname>
              <given-names>AS</given-names>
            </name>
            <name>
              <surname>Gomez</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Duverger</surname>
              <given-names>P</given-names>
            </name>
          </person-group>
          <person-group person-group-type="editor">
            <name>
              <surname>Dhillon</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Karlsson</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Hedstr&#xF6;m</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Z&#xFA;quete</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>Differentially private generative adversarial networks for time series, continuous, and discrete open data</article-title>
          <source>ICT Systems Security and Privacy Protection: 34th IFIP TC 11 International Conference, SEC 2019, Lisbon, Portugal, June 25-27, 2019, Proceedings</source>
          <year>2019</year>
          <publisher-loc>Cham</publisher-loc>
          <publisher-name>Springer International Publishing</publisher-name>
          <fpage>151</fpage>
          <lpage>164</lpage>
        </citation>
      </ref>
      <ref id="R32">
        <label>32</label>
        <citation citation-type="web">
          <collab>GA4GH</collab>
          <article-title>GA4GH File Encryption Standard</article-title>
          <year>2019</year>
          <access-date>2021 Jun 9</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/samtools/hts-specs/blob/master/crypt4gh.pdf">https://github.com/samtools/hts-specs/blob/master/crypt4gh.pdf</ext-link></comment>
        </citation>
      </ref>
      <ref id="R33">
        <label>33</label>
        <citation citation-type="web">
          <collab>GDPR.EU</collab>
          <article-title>Recital 34 - Genetic data - GDPR.eu</article-title>
          <year>2018</year>
          <access-date>2021 Jun 7</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://gdpr.eu/Recital-34-Genetic-data">https://gdpr.eu/Recital-34-Genetic-data</ext-link></comment>
        </citation>
      </ref>
      <ref id="R34">
        <label>34</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Godard</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Raeburn</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Pembrey</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Bobrow</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Farndon</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Aym&#xE9;</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>Genetic information and testing in insurance and employment: technical, social and ethical issues</article-title>
          <source>Eur J Hum Genet</source>
          <year>2003</year>
          <volume>11</volume>
          <issue>Suppl 2</issue>
          <fpage>S123</fpage>
          <lpage>S142</lpage>
        </citation>
      </ref>
      <ref id="R35">
        <label>35</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Goodfellow</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Pouget-Abadie</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Mirza</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Xu</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Warde-Farley</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Ozair</surname>
              <given-names>S</given-names>
            </name>
            <etal />
          </person-group>
          <person-group person-group-type="editor">
            <name>
              <surname>Ghahramani</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Welling</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Cortes</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Lawrence</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Weinberger</surname>
              <given-names>KQ</given-names>
            </name>
          </person-group>
          <article-title>Generative adversarial networks</article-title>
          <source>Advances in Neural Information Processing Systems 27 (NIPS 2014)</source>
          <year>2014</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1406.2661">https://arxiv.org/abs/1406.2661</ext-link></comment>
        </citation>
      </ref>
      <ref id="R36">
        <label>36</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Goodwin</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>McPherson</surname>
              <given-names>JD</given-names>
            </name>
            <name>
              <surname>McCombie</surname>
              <given-names>WR</given-names>
            </name>
          </person-group>
          <article-title>Coming of age: Ten years of next-generation sequencing technologies</article-title>
          <source>Nat Rev Genet</source>
          <year>2016</year>
          <volume>17</volume>
          <fpage>333–51</fpage>
        </citation>
      </ref>
      <ref id="R37">
        <label>37</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>G&#xFC;rsoy</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Emani</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Brannon</surname>
              <given-names>CM</given-names>
            </name>
            <name>
              <surname>Jolanki</surname>
              <given-names>OA</given-names>
            </name>
            <name>
              <surname>Harmanci</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Strattan</surname>
              <given-names>JS</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Data sanitization to reduce private information leakage from functional genomics</article-title>
          <source>Cell</source>
          <year>2020</year>
          <volume>183</volume>
          <fpage>905</fpage>
          <lpage>17.e16</lpage>
        </citation>
      </ref>
      <ref id="R38">
        <label>38</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Hagestedt</surname>
              <given-names>I</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Humbert</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Berrang</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Tang</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>X</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>MBeacon: Privacy-preserving beacons for DNA methylation data</article-title>
          <year>2019</year>
          <conf-name>Network and Distributed Systems Security (NDSS) Symposium 2019, 24-27 February 2019, San Diego, CA, USA</conf-name>
          <publisher-loc>Reston, VA</publisher-loc>
          <publisher-name>Internet Society</publisher-name>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://dx.doi.org/10.14722/ndss.2019.23064">https://dx.doi.org/10.14722/ndss.2019.23064</ext-link></comment>
        </citation>
      </ref>
      <ref id="R39">
        <label>39</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Hanzlik</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Grosse</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Salem</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Augustin</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Backes</surname>
              <given-names>M</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>MLCapsule: Guarded offline deployment of machine learning as a service</article-title>
          <year>2021</year>
          <publisher-name>Computer Vision Foundation</publisher-name>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content/CVPR2021W/TCV/papers/Hanzlik_MLCapsule_Guarded_Offline_Deployment_of_Machine_Learning_as_a_Service_CVPRW_2021_paper.pdf">https://openaccess.thecvf.com/content/CVPR2021W/TCV/papers/Hanzlik_MLCapsule_Guarded_Offline_Deployment_of_Machine_Learning_as_a_Service_CVPRW_2021_paper.pdf</ext-link></comment>
        </citation>
      </ref>
      <ref id="R40">
        <label>40</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Harmanci</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Gerstein</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Quantification of private information leakage from phenotype-genotype data: Linking attacks</article-title>
          <source>Nat Methods</source>
          <year>2016</year>
          <volume>13</volume>
          <fpage>251–6</fpage>
        </citation>
      </ref>
      <ref id="R41">
        <label>41</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Hayes</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Melis</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Danezis</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>De Cristofaro</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>LOGAN: Membership inference attacks against generative models</article-title>
          <year>2019</year>
          <conf-name>Proceedings on Privacy Enhancing Technologies (PoPETs)</conf-name>
          <fpage>133</fpage>
          <lpage>152</lpage>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://www.sciendo.com/article/10.2478/popets-2019-0008">https://www.sciendo.com/article/10.2478/popets-2019-0008</ext-link></comment>
        </citation>
      </ref>
      <ref id="R42">
        <label>42</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Hilprecht</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>H&#xE4;rterich</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Bernau</surname>
              <given-names>D</given-names>
            </name>
          </person-group>
          <article-title>Monte Carlo and reconstruction membership inference attacks against generative models</article-title>
          <year>2019</year>
          <conf-name>Proceedings on Privacy Enhancing Technologies (PoPETs)</conf-name>
          <fpage>232–49</fpage>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2478/popets-2019-0067">https://doi.org/10.2478/popets-2019-0067</ext-link></comment>
        </citation>
      </ref>
      <ref id="R43">
        <label>43</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Hinton</surname>
              <given-names>GE</given-names>
            </name>
            <name>
              <surname>Osindero</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Teh</surname>
              <given-names>Y-W</given-names>
            </name>
          </person-group>
          <article-title>A fast learning algorithm for deep belief nets</article-title>
          <source>Neural Comput</source>
          <year>2006</year>
          <volume>18</volume>
          <fpage>1527–54</fpage>
        </citation>
      </ref>
      <ref id="R44">
        <label>44</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Joly</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Dyke</surname>
              <given-names>SO</given-names>
            </name>
            <name>
              <surname>Cheung</surname>
              <given-names>WA</given-names>
            </name>
            <name>
              <surname>Rothstein</surname>
              <given-names>MA</given-names>
            </name>
            <name>
              <surname>Pastinen</surname>
              <given-names>T</given-names>
            </name>
          </person-group>
          <article-title>Risk of re-identification of epigenetic methylation data: a more nuanced response is needed</article-title>
          <source>Clin Epigenetics</source>
          <year>2015</year>
          <volume>7</volume>
          <fpage>45</fpage>
        </citation>
      </ref>
      <ref id="R45">
        <label>45</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Kairouz</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>McMahan</surname>
              <given-names>HB</given-names>
            </name>
            <name>
              <surname>Avent</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Bellet</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Bennis</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Bhagoji</surname>
              <given-names>AN</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Advances and open problems in federated learning</article-title>
          <source>Foundations and Trends&#xAE; in Machine Learning</source>
          <year>2019</year>
          <volume>14</volume>
          <issue>1-2</issue>
          <fpage>1</fpage>
          <lpage>210</lpage>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1912.04977">https://arxiv.org/abs/1912.04977</ext-link></comment>
        </citation>
      </ref>
      <ref id="R46">
        <label>46</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Kingma</surname>
              <given-names>DP</given-names>
            </name>
            <name>
              <surname>Welling</surname>
              <given-names>M</given-names>
            </name>
          </person-group>
          <article-title>Auto-Encoding variational bayes</article-title>
          <year>2013</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1312.6114">https://arxiv.org/abs/1312.6114</ext-link></comment>
        </citation>
      </ref>
      <ref id="R47">
        <label>47</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>LaBarre</surname>
              <given-names>BA</given-names>
            </name>
            <name>
              <surname>Goncearenco</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Petrykowska</surname>
              <given-names>HM</given-names>
            </name>
            <name>
              <surname>Jaratlerdsiri</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Bornman</surname>
              <given-names>MSR</given-names>
            </name>
            <name>
              <surname>Hayes</surname>
              <given-names>VM</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>MethylToSNP: Identifying SNPs in Illumina DNA methylation array data</article-title>
          <source>Epigenetics Chromatin</source>
          <year>2019</year>
          <volume>12</volume>
          <fpage>79</fpage>
        </citation>
      </ref>
      <ref id="R48">
        <label>48</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Li</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Gu</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Dvornek</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Staib</surname>
              <given-names>LH</given-names>
            </name>
            <name>
              <surname>Ventola</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Duncan</surname>
              <given-names>JS</given-names>
            </name>
          </person-group>
          <article-title>Multi-site fMRI analysis using privacy-preserving federated learning and domain adaptation: ABIDE results</article-title>
          <source>Med Image Anal</source>
          <year>2020</year>
          <volume>65</volume>
          <fpage>101765</fpage>
        </citation>
      </ref>
      <ref id="R49">
        <label>49</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Libbrecht</surname>
              <given-names>MW</given-names>
            </name>
            <name>
              <surname>Noble</surname>
              <given-names>WS</given-names>
            </name>
          </person-group>
          <article-title>Machine learning applications in genetics and genomics</article-title>
          <source>Nat Rev Genet</source>
          <year>2015</year>
          <volume>16</volume>
          <fpage>321–32</fpage>
        </citation>
      </ref>
      <ref id="R50">
        <label>50</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Lin</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Owen</surname>
              <given-names>AB</given-names>
            </name>
            <name>
              <surname>Altman</surname>
              <given-names>RB</given-names>
            </name>
          </person-group>
          <article-title>Genomic research and human subject privacy</article-title>
          <source>Science</source>
          <year>2004</year>
          <volume>305</volume>
          <fpage>183</fpage>
        </citation>
      </ref>
      <ref id="R51">
        <label>51</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Lipp</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Kogler</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Oswald</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Schwarz</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Easdon</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Canella</surname>
              <given-names>C</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>PLATYPUS: Software-based power side-channel attacks on x86</article-title>
          <year>2021</year>
          <conf-name>42th IEEE Symposium on Security and Privacy - San Francisco, Virtuell, USA, May 20-21, 2021</conf-name>
          <publisher-loc>New York, NY</publisher-loc>
          <publisher-name>IEEE Computer Society</publisher-name>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://platypusattack.com/platypus.pdf">https://platypusattack.com/platypus.pdf</ext-link></comment>
        </citation>
      </ref>
      <ref id="R52">
        <label>52</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Lowe</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Shirley</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Bleackley</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Dolan</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Shafee</surname>
              <given-names>T</given-names>
            </name>
          </person-group>
          <article-title>Transcriptomics technologies</article-title>
          <source>PLoS Comput Biol</source>
          <year>2017</year>
          <volume>13</volume>
          <fpage>e1005457</fpage>
        </citation>
      </ref>
      <ref id="R53">
        <label>53</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Mar&#xF3;ti</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Boldogk&#x151;i</surname>
              <given-names>Z</given-names>
            </name>
            <name>
              <surname>Tomb&#xE1;cz</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Snyder</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Kalm&#xE1;r</surname>
              <given-names>T</given-names>
            </name>
          </person-group>
          <article-title>Evaluation of whole exome sequencing as an alternative to BeadChip and whole genome sequencing in human population genetic analysis</article-title>
          <source>BMC Genomics</source>
          <year>2018</year>
          <volume>19</volume>
          <fpage>778</fpage>
        </citation>
      </ref>
      <ref id="R54">
        <label>54</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Mehrabi</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Morstatter</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Saxena</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Lerman</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Galstyan</surname>
              <given-names>A</given-names>
            </name>
          </person-group>
          <article-title>A survey on bias and fairness in machine learning</article-title>
          <year>2019</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1908.09635">https://arxiv.org/abs/1908.09635</ext-link></comment>
        </citation>
      </ref>
      <ref id="R55">
        <label>55</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Mittos</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Malin</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>De Cristofaro</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>Systematizing genome privacy research: A privacy-enhancing technologies perspective</article-title>
          <source>Proceedings on Privacy Enhancing Technologies (PoPETs)</source>
          <year>2019</year>
          <volume>2019</volume>
          <fpage>87–107</fpage>
        </citation>
      </ref>
      <ref id="R56">
        <label>56</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Nica</surname>
              <given-names>AC</given-names>
            </name>
            <name>
              <surname>Dermitzakis</surname>
              <given-names>ET</given-names>
            </name>
          </person-group>
          <article-title>Expression quantitative trait loci: present and future</article-title>
          <source>Philos Trans R Soc Lond B Biol Sci</source>
          <year>2013</year>
          <volume>368</volume>
          <fpage>20120362</fpage>
        </citation>
      </ref>
      <ref id="R57">
        <label>57</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Nyholt</surname>
              <given-names>DR</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>C-E</given-names>
            </name>
            <name>
              <surname>Visscher</surname>
              <given-names>PM</given-names>
            </name>
          </person-group>
          <article-title>On Jim Watson&#x2019;s APOE status: Genetic information is hard to hide</article-title>
          <source>Eur J Hum Genet</source>
          <year>2009</year>
          <volume>17</volume>
          <fpage>147–9</fpage>
        </citation>
      </ref>
      <ref id="R58">
        <label>58</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Oprisanu</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Dessimoz</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>De Cristofaro</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>How much does Genoguard really &#x201C;guard&#x201D;&#x3F; An empirical analysis of long-term security for genomic data</article-title>
          <year>2019</year>
          <conf-name>Proceedings of the 18th ACM Workshop on Privacy in the Electronic Society - WPES&#x2019;19 (pp 93-105)</conf-name>
          <publisher-loc>New York, NY</publisher-loc>
          <publisher-name>ACM Press</publisher-name>
        </citation>
      </ref>
      <ref id="R59">
        <label>59</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Oprisanu</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Ganev</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>De Cristofaro</surname>
              <given-names>E</given-names>
            </name>
          </person-group>
          <article-title>Measuring utility and privacy of synthetic genomic data</article-title>
          <year>2021</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2102.03314">https://arxiv.org/abs/2102.03314</ext-link></comment>
        </citation>
      </ref>
      <ref id="R60">
        <label>60</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Philibert</surname>
              <given-names>RA</given-names>
            </name>
            <name>
              <surname>Terry</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Erwin</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Philibert</surname>
              <given-names>WJ</given-names>
            </name>
            <name>
              <surname>Beach</surname>
              <given-names>SR</given-names>
            </name>
            <name>
              <surname>Brody</surname>
              <given-names>GH</given-names>
            </name>
          </person-group>
          <article-title>Methylation array data can simultaneously identify individuals and convey protected health information: an unrecognized ethical concern</article-title>
          <source>Clin Epigenetics</source>
          <year>2014</year>
          <volume>6</volume>
          <fpage>28</fpage>
        </citation>
      </ref>
      <ref id="R61">
        <label>61</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Qiu</surname>
              <given-names>YL</given-names>
            </name>
            <name>
              <surname>Zheng</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Gevaert</surname>
              <given-names>O</given-names>
            </name>
          </person-group>
          <article-title>Genomic data imputation with variational auto-encoders</article-title>
          <source>Gigascience</source>
          <year>2020</year>
          <volume>9</volume>
          <issue>8</issue>
          <fpage>giaa082</fpage>
        </citation>
      </ref>
      <ref id="R62">
        <label>62</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Rieke</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Hancox</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Li</surname>
              <given-names>W</given-names>
            </name>
            <name>
              <surname>Milletar&#xEC;</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Roth</surname>
              <given-names>HR</given-names>
            </name>
            <name>
              <surname>Albarqouni</surname>
              <given-names>S</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>The future of digital health with federated learning</article-title>
          <source>npj Digital Med</source>
          <year>2020</year>
          <volume>3</volume>
          <fpage>119</fpage>
        </citation>
      </ref>
      <ref id="R63">
        <label>63</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Robert</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Pelletier</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Exploring the impact of single-nucleotide polymorphisms on translation</article-title>
          <source>Front Genet</source>
          <year>2018</year>
          <volume>9</volume>
          <fpage>507</fpage>
        </citation>
      </ref>
      <ref id="R64">
        <label>64</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Sadat</surname>
              <given-names>MN</given-names>
            </name>
            <name>
              <surname>Al Aziz</surname>
              <given-names>MM</given-names>
            </name>
            <name>
              <surname>Mohammed</surname>
              <given-names>N</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Jiang</surname>
              <given-names>X</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>SAFETY: Secure gwAs in federated environment through a hYbrid solution</article-title>
          <source>IEEE&#x2F;ACM Trans Comput Biol Bioinform</source>
          <year>2019</year>
          <volume>16</volume>
          <fpage>93–102</fpage>
        </citation>
      </ref>
      <ref id="R65">
        <label>65</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Schadt</surname>
              <given-names>EE</given-names>
            </name>
            <name>
              <surname>Woo</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Hao</surname>
              <given-names>K</given-names>
            </name>
          </person-group>
          <article-title>Bayesian method to predict individual SNP genotypes from gene expression data</article-title>
          <source>Nat Genet</source>
          <year>2012</year>
          <volume>44</volume>
          <fpage>603–8</fpage>
        </citation>
      </ref>
      <ref id="R66">
        <label>66</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Shabani</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Borry</surname>
              <given-names>P</given-names>
            </name>
          </person-group>
          <article-title>Rules for processing genetic data for research purposes in view of the new EU General Data Protection Regulation</article-title>
          <source>Eur J Hum Genet</source>
          <year>2018</year>
          <volume>26</volume>
          <fpage>149–56</fpage>
        </citation>
      </ref>
      <ref id="R67">
        <label>67</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Sheller</surname>
              <given-names>MJ</given-names>
            </name>
            <name>
              <surname>Reina</surname>
              <given-names>GA</given-names>
            </name>
            <name>
              <surname>Edwards</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Martin</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Bakas</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <article-title>Multi-institutional deep learning modeling without sharing patient data: A feasibility study on brain tumor segmentation</article-title>
          <source>Brainlesion</source>
          <year>2019</year>
          <volume>11383</volume>
          <fpage>92–104</fpage>
        </citation>
      </ref>
      <ref id="R68">
        <label>68</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Suwinski</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Ong</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Ling</surname>
              <given-names>MHT</given-names>
            </name>
            <name>
              <surname>Poh</surname>
              <given-names>YM</given-names>
            </name>
            <name>
              <surname>Khan</surname>
              <given-names>AM</given-names>
            </name>
            <name>
              <surname>Ong</surname>
              <given-names>HS</given-names>
            </name>
          </person-group>
          <article-title>Advancing personalized medicine through the application of whole exome sequencing and big data analytics</article-title>
          <source>Front Genet</source>
          <year>2019</year>
          <volume>10</volume>
          <fpage>49</fpage>
        </citation>
      </ref>
      <ref id="R69">
        <label>69</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Sweeney</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Abu</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Winn</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Identifying participants in the personal genome project by name. Harvard: Data Privacy Lab, IQSS, Harvard University. White paper</article-title>
          <year>2013</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1304.7605">https://arxiv.org/abs/1304.7605</ext-link></comment>
        </citation>
      </ref>
      <ref id="R70">
        <label>70</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Tang</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Barbacioru</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>Y</given-names>
            </name>
            <name>
              <surname>Nordman</surname>
              <given-names>E</given-names>
            </name>
            <name>
              <surname>Lee</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Xu</surname>
              <given-names>N</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>mRNA-Seq whole-transcriptome analysis of a single cell</article-title>
          <source>Nat Methods</source>
          <year>2009</year>
          <volume>6</volume>
          <fpage>377–82</fpage>
        </citation>
      </ref>
      <ref id="R71">
        <label>71</label>
        <citation citation-type="web">
          <collab>The White House</collab>
          <article-title>Precision Medicine Initiative &#x7C; The White House</article-title>
          <year>2015</year>
          <access-date>2021 Jun 19</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://obamawhitehouse.archives.gov/precision-medicine">https://obamawhitehouse.archives.gov/precision-medicine</ext-link></comment>
        </citation>
      </ref>
      <ref id="R72">
        <label>72</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Torkzadehmahani</surname>
              <given-names>R</given-names>
            </name>
            <name>
              <surname>Kairouz</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Paten</surname>
              <given-names>B</given-names>
            </name>
          </person-group>
          <article-title>DP-CGAN: Differentially private synthetic data and label generation</article-title>
          <year>2019</year>
          <conf-name>2019 IEEE&#x2F;CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp 98-104)</conf-name>
          <publisher-loc>New York, NY</publisher-loc>
          <publisher-name>IEEE</publisher-name>
        </citation>
      </ref>
      <ref id="R73">
        <label>73</label>
        <citation citation-type="web">
          <collab>U.S. Equal Employment Opportunity Commission</collab>
          <article-title>The Genetic Information Nondiscrimination Act of 2008 &#x7C; U.S. Equal Employment Opportunity Commission</article-title>
          <year>2008</year>
          <access-date>2021 Jun 9</access-date>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://www.eeoc.gov/statutes/genetic-information-nondiscrimination-act-2008">https://www.eeoc.gov/statutes/genetic-information-nondiscrimination-act-2008</ext-link></comment>
        </citation>
      </ref>
      <ref id="R74">
        <label>74</label>
        <citation citation-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Van Bulck</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Minkin</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Weisse</surname>
              <given-names>O</given-names>
            </name>
            <name>
              <surname>Genkin</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Kasikci</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Piessens</surname>
              <given-names>F</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Foreshadow: Extracting the keys to the Intel SGX Kingdom with transient out-of-order execution</article-title>
          <year>2018</year>
          <conf-name>SEC&#x27;18: Proceedings of the 27th USENIX Conference on Security Symposium, August 2018 (pp 991&#x2013;1008)</conf-name>
          <publisher-name>USENIX Association</publisher-name>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://www.usenix.org/system/files/conference/usenixsecurity18/sec18-van_bulck.pdf">https://www.usenix.org/system/files/conference/usenixsecurity18/sec18-van_bulck.pdf</ext-link></comment>
        </citation>
      </ref>
      <ref id="R75">
        <label>75</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Ward</surname>
              <given-names>ET</given-names>
            </name>
            <name>
              <surname>Kostick</surname>
              <given-names>KM</given-names>
            </name>
            <name>
              <surname>L&#xE1;zaro-Mu&#xF1;oz</surname>
              <given-names>G</given-names>
            </name>
          </person-group>
          <article-title>Integrating genomics into psychiatric practice: Ethical and legal challenges for clinicians</article-title>
          <source>Harv Rev Psychiatry</source>
          <year>2019</year>
          <volume>27</volume>
          <fpage>53–64</fpage>
        </citation>
      </ref>
      <ref id="R76">
        <label>76</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Warnat-Herresthal</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Schultze</surname>
              <given-names>H</given-names>
            </name>
            <name>
              <surname>Shastry</surname>
              <given-names>KL</given-names>
            </name>
            <name>
              <surname>Manamohan</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Mukherjee</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Garg</surname>
              <given-names>V</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Swarm Learning for decentralized and confidential clinical machine learning</article-title>
          <source>Nature</source>
          <year>2021</year>
          <volume>594</volume>
          <issue>7862</issue>
          <fpage>265</fpage>
          <lpage>270</lpage>
        </citation>
      </ref>
      <ref id="R77">
        <label>77</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Weinhold</surname>
              <given-names>B</given-names>
            </name>
          </person-group>
          <article-title>Epigenetics: the science of change</article-title>
          <source>Environ Health Perspect</source>
          <year>2006</year>
          <volume>114</volume>
          <fpage>A160</fpage>
          <lpage>A167</lpage>
        </citation>
      </ref>
      <ref id="R78">
        <label>78</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Wilkinson</surname>
              <given-names>MD</given-names>
            </name>
            <name>
              <surname>Dumontier</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Aalbersberg</surname>
              <given-names>IJJ</given-names>
            </name>
            <name>
              <surname>Appleton</surname>
              <given-names>G</given-names>
            </name>
            <name>
              <surname>Axton</surname>
              <given-names>M</given-names>
            </name>
            <name>
              <surname>Baak</surname>
              <given-names>A</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>The FAIR Guiding Principles for scientific data management and stewardship</article-title>
          <source>Sci Data</source>
          <year>2016</year>
          <volume>3</volume>
          <fpage>160018</fpage>
        </citation>
      </ref>
      <ref id="R79">
        <label>79</label>
        <citation citation-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Xie</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Lin</surname>
              <given-names>K</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>S</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>F</given-names>
            </name>
            <name>
              <surname>Zhou</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>Differentially private generative adversarial network</article-title>
          <year>2018</year>
          <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1802.06739">https://arxiv.org/abs/1802.06739</ext-link></comment>
        </citation>
      </ref>
      <ref id="R80">
        <label>80</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Xu</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Glicksberg</surname>
              <given-names>BS</given-names>
            </name>
            <name>
              <surname>Su</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Walker</surname>
              <given-names>P</given-names>
            </name>
            <name>
              <surname>Bian</surname>
              <given-names>J</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>F</given-names>
            </name>
          </person-group>
          <article-title>Federated learning for healthcare informatics</article-title>
          <source>J Healthc Inform Res</source>
          <year>2021</year>
          <volume>5</volume>
          <fpage>1–19</fpage>
        </citation>
      </ref>
      <ref id="R81">
        <label>81</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Yelmen</surname>
              <given-names>B</given-names>
            </name>
            <name>
              <surname>Decelle</surname>
              <given-names>A</given-names>
            </name>
            <name>
              <surname>Ongaro</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Marnetto</surname>
              <given-names>D</given-names>
            </name>
            <name>
              <surname>Tallec</surname>
              <given-names>C</given-names>
            </name>
            <name>
              <surname>Montinaro</surname>
              <given-names>F</given-names>
            </name>
            <etal />
          </person-group>
          <article-title>Creating artificial human genomes using generative neural networks</article-title>
          <source>PLoS Genet</source>
          <year>2021</year>
          <volume>17</volume>
          <fpage>e1009303</fpage>
        </citation>
      </ref>
      <ref id="R82">
        <label>82</label>
        <citation citation-type="journal">
          <person-group>
            <name>
              <surname>Zabel</surname>
              <given-names>J</given-names>
            </name>
          </person-group>
          <article-title>The killer inside us: Law, ethics, and the forensic use of family genetics</article-title>
          <source>Berkeley J Crim Law</source>
          <year>2019</year>
          <volume>24</volume>
          <issue>2</issue>
          <fpage>47</fpage>
          <lpage>100</lpage>
        </citation>
      </ref>
      <ref id="R83">
        <label>83</label>
        <citation citation-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Zhu</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Han</surname>
              <given-names>S</given-names>
            </name>
          </person-group>
          <person-group person-group-type="editor">
            <name>
              <surname>Yang</surname>
              <given-names>Q</given-names>
            </name>
            <name>
              <surname>Fan</surname>
              <given-names>L</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>H</given-names>
            </name>
          </person-group>
          <article-title>Deep leakage from gradients</article-title>
          <source>Federated learning: Privacy and incentive</source>
          <year>2020</year>
          <publisher-loc>Cham</publisher-loc>
          <publisher-name>Springer International Publishing</publisher-name>
          <fpage>17</fpage>
          <lpage>31</lpage>
        </citation>
      </ref>
    </ref-list>
  </back>
  <floats-wrap>
    <fig id="F1" position="float">
      <label>Figure 1</label>
      <caption><title>Brief overview over the contents of this review. a: The three different types of genomic data that are covered in this work. A: Adenine; C: Cytosine; G: Guanine; T: Thymine; me: methyl group. b: Shown are a selection of applications that encourage data sharing. From left to right: Genomic data sharing is often required when building machine learning models in order to increase the available sample size required for training. Collecting and enriching data on minorities can reduce subpopulation bias in a trained model. Data often needs joining in multiparty studies when it is collected at different sites. Other motivators are sharing genomic data to allow the reproducibility of results or to reuse the data for new scientific questions. c: The subject re-identification is the core concern in genomic data privacy. The ability to produce uniquely identifying Single-Nucleotide-Polymorphism(SNP)-barcodes from the data allows an adversary to cross-reference these with public databases, often containing meta information that give away sensitive medical information. d: A timeline of selected laws that were introduced in several countries to protect citizens from discrimination based on genome-related data. e: Displayed are a selection of commonly used data sharing methods, colour-coded based on the maximum level of security they can provide. f: Selection of upcoming sharing techniques that are subject of ongoing research. Also shown as a necessary future step is the invocation of globally valid laws to protect subjects from discrimination in the case of a security breach. GAN: Generative Adversarial Network; RBM: Restricted Boltzmann Machine; VAE: Variational Autoencoder.</title></caption>
      <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="EXCLI-20-1243-g-001" />
    </fig>
  </floats-wrap>
</article>