🧬 Part 5: Complex Inheritance and Population Genetics English

← Back to Index
⚠️ For personal study use only. Commercial use is prohibited.
0%
0 / 53 listened Reset
Part 1 Part 2 Part 3 Part 4 Part 5 Part 6 Part 7 Part 8 Part 9 Part 10

Chapter 9: Complex Inheritance of Common Multifactorial Disorders

Ch9 · Pt1 chapter 9 Complex Inheritance of Common Multifactorial Disorders Cristen J. Willer
Gonçalo R. Abecasis Common diseases such as heart disease, cancer, diabetes, neuropsychiatric disease, and asthma cause morbidity and premature mortality in nearly two of every three individuals (Tabl...
Ch9 · Pt2 150 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE affected cases or unaffected controls) but can also be summarized using continuous quantitative traits or even discrete ordinal scales. At first glance, a binary classification strategy is the simpler approach; a disease, such as asthma, obesity, or hearing loss, is classified as present or absent in each individual being studied. Distinguishing between individuals who have a disease and those who do not may not always be straightforward and may require detailed examination, specialized testing, or even arbitrary distinctions and cutoff points. As an alternative to detailed examination of each study participant, many contemporary studies use automated algorithms for assigning disease states to individuals based on electronically encoded information in their medical records. A popular set of strategies in this class is the use of Phe Codes, a disease state definition based on presence or absence of one or more billing, diagnostic, or procedure codes in the medical record. It’s worth noting that while these strategies can be extremely practical and have enabled many successful gene-mapping experiments, they can also result in arbitrary or inconsistent classification of individual individuals (e.g., depending on the choices their health care provider might have made in diagnosis, testing, billing, or treatment). In some cases, multiple instances of a code in the electronic health record are required for a more confident diagnosis, and individuals with unclear phenotypes are excluded from both the case and control groups. Quantitative traits can often avoid arbitrary boundaries between cases and controls. Instead of classifying individuals as asthmatic cases and nonasthmatic controls, we might measure their lung capacity or the number of emergency room visits per year. Instead of classifying individuals as hard-of-hearing cases or normal-hearing controls, we might use a quantitative measure to quantify the loudness or pitch of sounds each individual can hear. Instead of classifying individuals as obese or nonobese, we might measure their weight or body mass index. Quantitative traits are also commonly used to summarize disease-related measurable physiologic or biochemical quantities such as blood pressure, serum cholesterol concentration, or activity levels that vary among individuals within a population. The Normal Distribution As is often the case with physiologic quantities, such as systolic blood pressure, a graph of the number (or the fraction) of individuals in the population (y-axis) having a particular quantitative value (x-axis) approximates the familiar, bell-shaped curve known as the normal (or gaussian) distribution (Fig. 9.1A). The position of the peak and the width of the curve of the normal distribution are governed by two quantities, the mean (µ) and the variance (σ2), respectively. The mean is the arithmetic average of the values, and because – for many traits – more people have values for the trait near the average, the curve ordinarily has its peak at the mean value. The variance (or its square root, σ, the standard deviation [SD]) is a measure of how much spread there is in the values to either side of the mean and therefore determines the breadth of the curve. Any physiologic quantity that can be measured in a sample of a population is a quantitative phenotype, and the mean and variance for that sample can be calculated and used to approximate the underlying mean and variance of the population from which the sample was drawn. For example, the systolic blood pressure of thousands of men in two different age groups is shown in Fig. 9.1B. The systolic blood pressure of the younger cohort is nearly symmetric; in the older age group, however, the curve becomes more skewed (asymmetric), with more individuals with systolic blood pressures above the mean than below. The normal distribution provides guidelines for setting the limits of the normal range. A normal range is often defined as the values of a quantitative trait that are seen in ~95% of the population. Statistical theory states that when the values of a quantitative trait in a population follow the bell-shaped normal curve (i.e., are normally distributed), ~5% of the population will have measurements more than 2 SD above or below the population mean. It is important to note that an individual may be perfectly healthy (i.e., “normal”) despite having a trait value outside the normal range. Furthermore, since 5% of the population will, by definition, be outside the normal range, when many traits are measured and thus classified, each individual will typically be extreme in several traits. It’s important to note that this concept of normal range should not be confused with health and disease. For example, because body mass is typically high in industrialized societies, many individuals within the normal range (e.g., within 2 SD of the mean) might be considered clinically obese. As another example, because the human body can physiologically tolerate variation in platelet levels, the extreme platelet levels used to diagnose clinical conditions like thrombocytosis TABLE 9.1 Frequency of Different Types of Genetic Disease Type Incidence at Birth (per 1000) Prevalence at Age 25 (per 1000) Population Prevalence (per 1000) Disorders due to genome and chromosome alterations 6 1.8 3.8 Disorders due to single-gene variants 10 3.6 20 Disorders with multifactorial inheritance ≈50 ≈50 ≈600 Data from Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin's principles and practice of medical genetics, ed 3, Edinburgh, 1997, Churchill Livingstone.
150 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE affected cases or unaffected controls) but can also be summarized using continuous quantitative traits or even discrete ordinal scales. At f...
Ch9 · Pt3 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 151 or thrombocytopenia are far outside the normal range. In general the normal distribution and the normal range are a statistical convenience but cannot be used directly to diagnose health and disease. For convenience, many genetic analyses assume that traits follow this bell-shaped normal distribution in the population. Although this is approximately true for many traits (such as height, weight, total cholesterol, and blood pressure in younger individuals), it is clearly not the case in other cases (such as triglyceride levels, number of moles in skin, or number of children in a family). For convenience, many genetic studies map the original measurements, which may not be normally distributed, to a new measurement scale that is normally distributed. This can provide much more flexibility in the choice of analysis strategy. The process typically works by rank-ordering the quantitative measurements (first highest among 100, second highest among 100, etc.) and then mapping these ranks to corresponding values in a simulated normal distribution (2.58 SD, 2.17 SD, etc.). The inverse normal distribution function is used to tabulate the expected values in the simulated normal distribution. This popular strategy handles outlier values or nonnormality in the original measured values and provides genetic results in SD units allowing for easier comparison between different studies. FAMILIAL AGGREGATION AND CORRELATION Allele Sharing Among Relatives The more closely related two individuals are, the more alleles they share in common, on average (see Chapter 7). The most extreme example of allele sharing is identical (monozygotic [MZ]) twins (see later in this chapter), who have the same alleles at every locus, with possibly a few small exceptions arising from somatic variants. The next most closely related individuals are typically first-degree relatives, such as a parent offspring or sibling pairs, including fraternal (dizygotic [DZ]) twins. In a parent-child pair, the child shares at least one allele out of two (50% of alleles) in common with the parent at every genetic location. This shared allele is on the chromosome the child inherited from that parent; sharing on the chromosome inherited from the other parent can result from consanguinity, very distant relatedness, or chance. Siblings (including DZ twins) also share 50% or more of their alleles on average, but this can vary along the genome. This is because, at a given locus, a pair of siblings inherits the same two chromosomes from their parents ¼ of the time, inherits one chromosome in common ½ the time, and inherits no chromosomes in common the remaining ¼ of the time (Fig. 9.2). At any one locus, in the absence of consanguinity, the average number of chromosomes that a sibling pair is expected to share identical by descent (i.e., alleles that are identical because they are copies of the same ancestral chromosome) is: 14 14 12 () () () 2alleles 0allele 1allele 0.5 0.5 0 1allele + + = + + = The more distantly related two members of a family are, the fewer alleles they are expected to inherit from a common ancestor. Familial Aggregation in Binary Traits If genetic variants modify disease risk, relatives of an affected individual will have a greater-than-expected rate of the disease than unrelated individuals with similar nongenetic risk profiles (familial aggregation of Percent of population Quantity being measured Systolic blood pressure 30 20 10 0 –2 SD –1 SD MEAN +1 SD +2 SD 100 120 140 160 180 200 134±40 127±34 A B Figure 9.1 (A) The normal gaussian distribution, with mean (average) and standard deviation (SD) indicated. For many traits, the “normal” range is considered the mean ±2 SD, as indicated by the shaded region. (B) Distribution of systolic blood pressure in ~3300 men aged 40–45 (solid line) and ~2200 men aged 50–55 (dotted line). The mean and ±2 SD are shown above double-headed arrows. (B, Data from Sive PH, Medalie JH, Kahn HA, et al: Distribution and multiple regression analysis of blood pressure in 10,000 Israeli men, Am J Epidemiol 93:317–327, 1971.)
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 151 or thrombocytopenia are far outside the normal range. In general the normal distribution and the normal range are a statistical c...
Ch9 · Pt4 152 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE disease). This is because closely related family members are expected to share, on average, some disease predisposing alleles. Next, we will discuss two approaches to measuring familial aggregation: relative risk ratios and family history case-control studies. Relative Risk Ratio One way to measure familial aggregation of a disease is by comparing the frequency of the disease in the relatives of an affected proband with its disease frequency (prevalence) in the general population. The relative risk ratio λr (where the subscript r refers to relatives) is defined as: λr Prevalence of the disease in the relatives of an affecte = d person Prevalence of the disease in the general population The value of λr as a measure of familial aggregation depends both on how frequently a disease occurs in relatives of an affected individual (the numerator) and on the population prevalence (the denominator); the larger λr is, the greater the familial aggregation. This estimate is typically performed within a specific type of close relative (e.g., siblings, offspring, identical twins). The population prevalence enters into the calculation because the more common a disease is, the greater the likelihood that apparent aggregation may be a coincidence. A value of λr = 1 indicates that a relative is no more likely to develop the disease than is any individual in the population, whereas a value greater than 1 indicates that a relative is more likely to develop the disease. Examples of relative risk ratios determined for various diseases in samples of siblings (λs where s is for siblings) are shown in Table 9.2. Since many diseases have sexand/or age-specific prevalences, it is important to ensure relatives and reference populations are appropriately matched with respect to these factors. Family History Case-Control Studies Another approach to estimating familial aggregation is the case-control study, in which individuals with a disease (the cases) are compared with suitably chosen individuals without the disease (the controls), with respect to family history of disease (as well as other factors, such as environmental exposures, occupation, geographic location, parity, and previous illnesses). To assess a possible genetic contribution to familial aggregation of a disease, the frequency with which the disease is found in the extended families of the cases (positive family history) is compared with the frequency of positive family history among suitable controls, matched for age and ancestry. Spouses can be used as controls in this situation because they usually match the cases in age and ancestry and share the same household environment (provided the disease is of similar prevalence in males and females). Other frequently used controls are individuals with unrelated diseases matched for age, sex, and ancestry. Thus, for example, in a study of multiple sclerosis (MS), ~3.5% of first-degree relatives of patients with MS also had MS, a prevalence that was much higher than among first-degree relatives of matched controls without A1A3 A1A2 A3A4 A1A3 A1A4 A1A4 A2A3 A2A3 A2A4 A2A4 2 1 1 0 1 2 0 1 1 0 2 1 0 1 1 2 Sib #1 Sib #2 Genotype of sib #1 Genotype of sib #2 Figure 9.2 Allele sharing at an arbitrary locus between sibs concordant for a disease. The parents’ genotypes are shown as A1A2 for the father and A3A4 for the mother. All four possible genotypes for sib #1 are given across the top of the table, and all four possible genotypes for sib #2 are given along the left side of the table. The numbers inside the boxes represent the number of alleles both sibs have in common for all 16 different combinations of genotypes for both sibs. For example, the upper left-hand corner has the number 2 because sib #1 and sib #2 both have the genotype A1A3 and so have both A1 and A3 alleles in common. The bottom left-hand corner contains the number 0 because sib #1 has genotype A1A3, whereas sib #2 has genotype A2A4, so there are no alleles in common. TABLE 9.2 Risk Ratios λs for Siblings of Probands with Diseases with Familial Aggregation and Complex Inheritance Disease Relationship λs Schizophrenia Siblings 12 Autism Siblings 150 Manic-depressive (bipolar) disorder Siblings 7 Type 1 diabetes mellitus Siblings 35 Crohn disease Siblings 25 Multiple sclerosis Siblings 24 Data from Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin's principles and practice of medical genetics, ed 3, Edinburgh, 1997, Churchill Livingstone; King RA, Rotter JI, Motulsky AG: The genetic basis of common diseases, ed 2, Oxford, England, 2002, Oxford University Press.
152 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE disease). This is because closely related family members are expected to share, on average, some disease predisposing alleles. Next, we will...
Ch9 · Pt5 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 153 MS (0.2%). That is, the odds of having a first-degree relative with MS were 18 times higher among, people with MS than among controls. (In Chapter 11, we will discuss how one calculates odds ratios in case-control studies.) One can conclude therefore that substantial familial aggregation is occurring in MS, thereby providing evidence of a genetic predisposition to this disease. These types of studies are somewhat vulnerable to recall bias, since diseased individuals are more likely to be aware of the disease status of similarly affected relatives. Measuring the Genetic Contribution to Quantitative Traits Sharing of alleles that govern a particular quantitative trait affects the distribution of values of that trait in family members. The more sharing of alleles that govern a quantitative trait there is among relatives, the more similar values of the trait are expected to be. The effect of genetic variation on quantitative traits is often measured and reported in two related ways: correlation between relatives and heritability. Familial Correlation Just like the relative risk ratios are used to summarize aggregation of disease within families, there are analogous strategies to summarize whether quantitative traits aggregate within families. Checking for these patterns of familial aggregation provides important clues about the role of genetic variation in each trait. Prior to studying the genetic factors underlying a trait, it is important to first establish that the trait has a genetic component. Geneticists answer this question in a few different ways. The tendency for the values of a physiologic measurement to be more similar among relatives is summarized through the correlation of these physiologic quantities among relatives. The coefficient of correlation (symbolized by the letter r) is a statistical measure of correlation applied to a pair of measurements, such as a child’s serum cholesterol level and that of a parent. A positive correlation between the cholesterol measurements in two groups of relatives exists if it is found that a higher level in the first individual (e.g., the child) predicts a proportionately higher level in the relative (e.g., the parent). When a correlation exists, a graph of values in the proband and his or her relatives, in which each point represents a proband-relative pair of values, will tend to cluster around a straight line. The value of r can range from 0 when there is no correlation to +1 for perfect positive correlation. In the example of serum cholesterol, Fig. 9.3 shows a modest positive correlation (r = 0.294) between serum cholesterol level of mothers 350 250 150 50 Cholesterol of son, mg% Cholesterol of mother, mg% 50 150 250 350 450 Figure 9.3 Plot of serum cholesterol levels in a group of mothers aged 30–39 and in their male children aged 4–9 years. Each dot represents a mother-son pair of measurements. The straight line is a “best fit” through the data points. (Data from Johnson BC, Epstein FH, Kjelsberg MO: Distributions and familial studies of blood pressure and serum cholesterol levels in a total community – Tecumseh, Michigan. J Chronic Dis 18:147–160, 1965.)
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 153 MS (0.2%). That is, the odds of having a first-degree relative with MS were 18 times higher among, people with MS than among cont...
Ch9 · Pt6 154 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE aged 30 to 39 and those of their male children aged 4 to 9. A negative correlation exists when an increase in the first individual’s measurement predicts a lower measurement in relatives. The measurements are still correlated but in the opposite direction. In such a case, the value of r can range from 0 to −1 for a perfect negative correlation. This is relatively rare in genetic studies of relatives, but it can occur in other settings where correlations are used to summarize relationships between measurements (e.g., an individual’s activity levels might be negatively correlated with body mass). Heritability The concept of heritability of a quantitative trait (symbolized as H2) was developed to describe how much the genetic differences between individuals in a population contribute to variability of that trait in the population. H2 is defined as the fraction of the variation in a quantitative trait that is due to genetic variation in the broadest sense, regardless of the mechanism by which the various alleles affect the phenotype. The higher the heritability, the greater the contribution of genetic differences among people to the variability of the trait in the population. The value of H2 varies from 0, if genotype contributes nothing to the total phenotypic variance in a population, to 1, if genotype is totally responsible for the phenotypic variance in that population. Heritability of a human trait is a theoretical quantity that is usually estimated from the correlation between measurements of that trait among relatives of known degrees of relatedness, such as parents and children, siblings, or, as we shall see later in this chapter, twins. DETERMINING THE RELATIVE CONTRIBUTIONS OF GENES AND ENVIRONMENT TO COMPLEX DISEASE Distinguishing Between Genetic and Environmental Influences Using Family Studies For both qualitative and quantitative traits, similarities among family members are most likely the result of shared genetics and shared environment, such as socioeconomic status, local environment, dietary habits, or cultural behaviors, all of which are frequently shared among family members but are generally considered to be of nongenetic origin. Given evidence of familial aggregation of a disease or correlation of a quantitative trait, geneticists attempt to separate the relative contributions of genotype and environment to the phenotype using a variety of approaches. Historically, when sample sizes were smaller and genetic data were less available, researchers would collect pedigree information from cases (probands) and their family members. This would enable estimates of disease risk in different relatives, grouped by their shared genetics (e.g., 50% genetic sharing in parent-offspring pairs or full siblings, 25% genetic sharing in grandparent-grandchild pairs, avuncular or half-sibling pairs) and an evaluation of whether the degree of similarity between individuals attenuates in proportion to their genetic relatedness. Another common attempt to control for shared environment is to compare the concordance of disease status in MZ or identical twins with that in same-sex DZ or fraternal twins. Since twins share much of their environment in utero and early childhood whether identical or fraternal, it is convenient to hypothesize that any difference in concordance is due to the difference in genetic sharing (100% sharing for identical twins and ~50% sharing for fraternal twins). More recently, with the development of biobanks (i.e., very large-scale collections of research participants, coupled with genetic data and large-scale electronic health records), heritability of many phenotypes and diseases can be efficiently examined using very distant relative pairs identified by comparing their genetic data directly. For example, we might hypothesize that for a genetic trait, concordance of disease status or correlation between quantitative measurements will be higher between individuals who share 4% of their genetic material than between individuals who share 2% of their genetic material. We next discuss some possibilities in more detail. Attenuation of Risk for Progressively More Distant Relatives One approach is to compare λr measurements or quantitative trait correlations between relatives of varying degrees of relatedness to the proband. For example, if genes predispose to a disease, one would expect λr to be greatest for MZ twins, to be somewhat smaller for firstdegree relatives such as sibs or parent-child pairs, and to continue to decrease as allele-sharing decreases among the more distant relatives in a family (see Fig. 7.3). To illustrate this approach, consider cleft lip with or without cleft palate, or CL(P), one of the most common congenital malformations, affecting 1.4 per 1000 newborns worldwide. CL(P) originates as a failure of fusion of embryonic tissues that will go to make up the upper lip and the hard palate at approximately day 35 of gestation. It is a multifactorial disorder with complex inheritance; for reasons that are not well understood, ~60% to 80% of those affected with CL(P) are males. Despite the similarity in names, CL(P) is usually etiologically distinct from isolated cleft palate (i.e., without cleft lip). CL(P) is heterogeneous and includes forms in which the clefting is only one feature of a syndrome that includes other features, known as syndromic CL(P), as well as forms that are not associated with other birth defects, which are known as nonsyndromic CL(P). Syndromic CL(P) can be inherited as a mendelian single-gene disorder or can be caused by chromosome disorders (especially trisomy 13 and 4p− deletion syndrome) (see Chapter 6) or teratogenic exposure (rubella embryopathy, thalidomide, or anticonvulsants) (see
154 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE aged 30 to 39 and those of their male children aged 4 to 9. A negative correlation exists when an increase in the first individual’s measure...
Ch9 · Pt7 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 155 Chapter 15). Nonsyndromic CL(P) can also be inherited as a single-gene disorder but more commonly is a sporadic occurrence and demonstrates some degree of familial aggregation without an obvious mendelian inheritance pattern. The risk for CL(P) in a child increases as a function of the number of relatives the child has who are affected with CL(P) and the more closely related they are to the child (Table 9.3). The simplest explanation for this is that the more closely related one is to the proband and, the more probands there are in the family, the more likely one is to inherit disease-susceptibility alleles, thus increasing risk for the disorder. Another approach is to compare the disease relative risk ratio in biological relatives of the proband with that in biologically unrelated family members (e.g., adoptees or spouses), all living in the same household environment. Returning to MS, for example, λr is 190 for MZ twins and 20 to 40 for first-degree biological relatives (parents, children, and sibs). In contrast, λr is 1 for the adopted siblings of an affected individual, suggesting that most of the familial aggregation in MS is genetic rather than the result of a shared environment. A similar analysis can be carried out for quantitative traits such as blood pressure: No correlation exists between a child’s blood pressure and that of the child’s adopted siblings, in contrast to the positive correlation for blood pressure of biological siblings, all living in the same household. Distinguishing Between Genetic and Environmental Influences Using Twin Studies Of all methods used to separate genetic and environmental influences, geneticists have historically relied most heavily on twin studies. Twinning MZ and DZ twins are “experiments of nature” that provide an excellent opportunity to separate environmental and genetic influences on phenotypes in humans. MZ twins arise from the cleavage of a single fertilized zygote into two separate zygotes early in embryogenesis (see Chapter 14). They occur in ~0.3% of all births, without large differences among different populations. At the time the zygote cleaves in two, MZ twins start out with identical genotypes at every locus and are therefore often thought of as having identical genomes. In contrast, DZ twins arise from the simultaneous fertilization of two eggs by two sperm; genetically, DZ twins are siblings who share a womb and, like all siblings, share, on average, 50% of the alleles at all loci. DZ twins are of the same sex half the time and of opposite sex the other half. In contrast to MZ twins, DZ twins occur with a frequency that varies as much as fivefold in individuals with different ancestries - e.g., 0.2% among individuals with Asian ancestry and 1% of births in parts of Africa. Disease Concordance in MZ and DZ Twins When twins have the same disease, they are said to be concordant for that disorder. Conversely, when only one member of the pair of twins is affected and the other is not, the relatives are discordant for the disease. An examination of how frequently MZ twins are concordant for a disease is a powerful method for determining whether genotype alone is sufficient to produce a particular disease. The differences between a disease that is mendelian from one that shows complex inheritance are immediately evident. Using sickle cell disease (Case 42) as an example of a mendelian disorder, if one MZ twin has sickle cell disease, the other twin will always have the disease as well. In contrast, as an example of a multifactorial disorder, when one MZ twin has type 1 diabetes mellitus (previously known as insulin-dependent or juvenile diabetes) the other twin will also have type 1 diabetes only ~40% of such twin pairs. Disease concordance of less than 100% in MZ twins is strong evidence that nongenetic factors play a role in the disease. Such factors could include environmental influences, such as exposure to infection or diet, as well as other effects, such as somatic variation, effects of aging, or epigenetic changes in gene expression in one twin compared with the other. MZ and same-sex DZ twins share a common intrauterine environment and biological sex and are usually reared together in the same household by the same parents. Thus a comparison of concordance for a disease between MZ and same-sex DZ twins shows how frequently disease occurs when relatives who experience the same prenatal and often the same postnatal environment have the same alleles at every locus (MZ twins) and how often it occurs when they share 50% of alleles in common (DZ twins). Greater concordance in MZ versus DZ twins is strong evidence of a genetic component to the disease, as shown in Table 9.4 for a number of disorders. TABLE 9.3 Risk for Cleft Lip with or without Cleft Palate in a Child Depending on the Number of Affected Parents and Other Relatives Affected Relatives Risk for CL(P) (%) No. of Affected Parents 0 1 2 None 0.1 3 34 One sibling 3 11 40 Two siblings 8 19 45 One sibling and one second-degree relative 6 16 43 One sibling and one third-degree relative 4 14 44 CL(P), Cleft lip with or without cleft palate.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 155 Chapter 15). Nonsyndromic CL(P) can also be inherited as a single-gene disorder but more commonly is a sporadic occurrence and de...
Ch9 · Pt8 156 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Estimating Heritability from Twin Studies Just as twins are be used to separate the roles of genes and environment in qualitative disease traits, twins are also used to estimate the heritability of a quantitative trait using the correlation in the values of a physiologic measurement in MZ and DZ twins. If one assumes that the alleles affecting the trait exert their effect additively (which is certainly overly simplistic and probably incorrect in many cases), MZ twins, who share 100% of their alleles, have twice the amount of allele sharing compared to DZ twins, who share 50% of their alleles on average. H2, introduced earlier in this chapter, can therefore be approximated by taking twice the difference in the correlation coefficient r for a quantitative trait between MZ twins (r MZ) and r between same-sex DZ twins (r DZ) (as given by Falconer’s formula): H r r MZ DZ 2 2
() If the variability of the trait is determined chiefly by environment, the correlation within pairs of DZ twins will be similar to that seen between pairs of MZ twins; there will be little differenc...
Ch9 · Pt9 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 157 also a problem for twin studies. Many studies rely on asking one twin with a particular disease to recruit the other twin to participate in a study (volunteer-based ascertainment), rather than ascertaining them first as twins through a twin registry and only then examining their health status (population-based ascertainment). Volunteer-based ascertainment can give biased results because twins, particularly MZ twins who may be emotionally close, are more likely to volunteer if they are concordant than if they are not, which inflates the concordance rate. Similarly, because case-control studies of family history often rely, for practical reasons, on taking a history from the proband rather than examining all the relatives directly, there may be recall bias, in which a proband may be more likely to know of family members with the same or similar disease. Such biases will inflate the level of familial aggregation. Other difficulties arise in measuring and interpreting heritability. The same trait may yield different measurements of heritability in different populations because of different allele frequencies or diverse environmental conditions. For example, heritability measurements of height would be lower when measured in a population with widespread famine that stunts growth in childhood as compared to the same population after food becomes plentiful. Heritability of a trait should therefore not be thought of as an intrinsic, universally applicable measure of “how genetic” the trait is, because it depends on the population and environment in which the estimate is being made. Although heritability estimates are still made in genetic research, most geneticists consider them to be only preliminary but convenient estimates of the role of genetic variation in phenotypic variation. Potential Genetic or Epigenetic Differences Despite the evident power of twin studies, one must caution against thinking of such studies as perfectly controlled experiments that compare individuals who share either half or all of their genetic variation and are exposed either to the same or to different environments. Studies of MZ twins assume the twins are genetically identical. Although this is mostly true, genotype and gene expression patterns may come to differ between MZ twins because of genetic or epigenetic changes that occur after the cleavage event that produced the MZ twin embryos. For example, genotype may differ due to somatic rearrangements and/or rare somatic mutations that occur after the cleavage event (see Chapter 3). Epigenetic changes may occur in response to environmental or stochastic factors, thus leading to differences in gene expression between MZ twins. (Female MZ twins have an additional source of variability because of the stochastic nature of X inactivation patterns in various tissues, as presented in Chapter 6.) Other Limitations Another problem may arise when assuming that the environmental exposure of MZ and DZ twins has been held constant when they are reared together but not when twins are reared apart. Environmental exposures, including even intrauterine environment, may vary for twins reared in the same family. For example, MZ twins frequently share a placenta, and there may be a disparity between the twins in blood supply, intrauterine development, and birthweight. For late-onset diseases, such as neurodegenerative disease of late adulthood, the assumption that MZ and DZ twins are exposed to similar environments throughout their adult lives becomes less and less valid, and thus a difference in concordance provides less strong evidence for genetic factors in disease causation. Conversely, one assumes that by determining disease concordance in MZ twins reared apart, one is measuring the effect of different environments on the same genotype. However, the environment of twins reared apart may actually not be as different as one might suppose. Thus no twin study is a perfectly controlled assessment of genetic versus environmental influence. Finally, caution is necessary when generalizing from twin studies. The most extreme situation would be when the phenotype being studied is only sometimes genetic in origin; that is, nongenetic phenocopies may exist. If genotype alone causes the disease in half the pairs of twins (MZ twin concordance of 100%) in your sample and a nongenetic phenocopy affects only one twin of the other half of twin pairs in your sample (MZ twin concordance of 0%), twin studies will show an intermediate level of 50% concordance that really applies to neither form of the disease. Genome-wide Association Studies Lei Sun and Wei Deng Complex diseases and traits are influenced by a combination of genetic and environmental risk factors. In the last decade, genome-wide association studies (GWAS) have been particularly fruitful in identifying disease susceptibility loci and biologic pathways, elucidating the genetic architecture of many complex diseases. The concept of GWAS is elegantly simple: examining each individual genetic variant – usually a biallelic locus (which is typically either a single nucleotide polymorphism [SNP] or an insertion deletion variant or indel) – for evidence of association with a disease phenotype or quantitative trait. Individual variant results can point to specific regions of the genome that are associated with a disease or trait. Sometimes the association signal implicates a protein-altering variant, which has a simple interpretation in terms of biologic and functional consequences, but most of the time the association signal falls among the noncoding regions of the genome. Many approaches have been developed to map associated genetic variants
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 157 also a problem for twin studies. Many studies rely on asking one twin with a particular disease to recruit the other twin to part...
Ch9 · Pt10 158 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE to their effector genes, but for complex disease the most sophisticated approaches only guess the correct gene a little over half of the time. These individual association results, over the entire genome, reveal the landscape of phenotype-genotype associations, where the association evidence is often summarized pictorially by the famous Manhattan plot of –log 10 p-values (Fig. 9.4). More recently, some have proposed a Brisbane plot enumerating the number of independently associated variants within a specified window size (useful for GWAS with many findings), or a Miami plot showing results from two similar GWASs above and below the horizontal axis (i.e., a reflection) (Figs. 9.5 and 9.6). Polygenic Risk Scores Lei Sun and Wei Deng In addition to discovery of novel association signals, GWAS have enabled the estimation of genetic effects associated with millions of individual SNPs across a range of complex diseases. The magnitude of these effects for each genetic variant is typically small compared to established clinical risk factors. As the higher risk allele at each SNP locus only moderately increases the risk of disease, it is natural to construct a composite score that can potentially capture the overall genetic burden across the genome. This is indeed the idea behind the PRS: a weighted sum of the numbers of the risk alleles of all associated SNPs, where the weights are derived from effect size estimates from GWAS. The earlier polygenic risk research focused on association, and one of the most prominent early examples of PRS was in schizophrenia. There, a large number of SNPs individually below the genome-wide detection threshold (of p-value <5 × 10−8, which is the threshold typically used to allow for millions of tests for common variants distributed throughout the genome) was used to construct a PRS highly associated with schizophrenia status. More recent PRS research has focused on the utility of PRS for prediction of individuals at risk of disease, on the ability of PRS to help refine diagnoses by separating different types of cases, and on the ability of PRS to aid in the selection of optimal treatment regimes. For example, a higher PRS is capable of identifying individuals with a higher disease risk: Individuals with PRS values in the top 1% of the distribution have at least a threefold increase in risk of developing coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, or breast cancer. Indeed, PRS of this magnitude could have many clinical utilities: They might inform screening strategies, motivate health-related behavior changes, and potentially help identify treatment targets for precision medicine. PRS for breast cancer was the first to be implemented in the Can Risk/BOADICEA risk algorithm, which was used clinically to predict individuals at higher risk of breast cancer, and became available for prediction of coronary artery disease by at least one private company in 2022 (Color). In the case of diagnostic refinement, PRS was found to distinguish between type 1 and type 2 diabetes, which had previously been assigned primarily based on age of onset. The construction of a good PRS, predictive of the risk of a disease, requires a systematic approach to (1) powerfully identify associated SNPs, (2) accurately estimate their genetic effects on the disease, and (3) ensure the PRS is applied to the appropriate people (e.g., whose age, ancestry, and clinical risk profiles are similar to those in whom the genetic effects were originally estimated). The ideal PRS construction should include all disease-associated SNPs and no others; a realistic PRS typically not only misses some of the truly associated SNPs but includes false positives due to limited power of GWAS. The low power of GWAS is a direct result of moderate-to-weak effects of individual SNPs on a complex disease with polygenic inheritance. The estimation of effect can be challenging, as it is often biased toward discovery – otherwise known as the winner’s curse – where the estimated genetic effect of an SNP is inflated relative to the true value. Finally, to avoid data dredging (also known as p-hacking), for both (1) and (2) above, a (discovery) sample of individuals is used, which is independent of the (target) sample of individuals where the PRS is evaluated for prediction accuracy. While statistical methodology in capturing the polygenic risk associated with diseases has progressed significantly, for many complex diseases and quantitative traits the amount of phenotypic variance explained by PRS derived from GWASs is modest. Thus for most common diseases we are still evaluating the value of PRS in the clinic and in other settings where they might influence the health and wellbeing of individuals. With the availability of biobank-scale studies and more sophisticated PRS methods, we expect incremental improvement in the predictiveness of PRS, particularly for diseases with few clinical predictors. A major methodologic issue in the construction of a PRS is population heterogeneity, also known as the transportability issue. Because genetic effect sizes may vary according to the ancestral background of the study population, the PRS weights derived in one population might not translate well into another. Another major gap is the lack of methods to include SNPs from the X chromosome, thus missing 5% of the haploid genome. Finally, the genetic risk for a disease, predicted using PRS, has a theoretical upper bound that depends on the population-specific heritability of the disease. This is not to say, however, that heritability limits the absolute disease risk for an individual. Consider the well-known example of familial breast cancer, whereby individuals with pathogenic variants in the BRCA1 gene can have a substantially increased risk of early-onset cancer, even though these relatively highly penetrant variants only account for a small number of breast cancer cases overall and only explain a small proportion of cancer heritability in the populations where they occur.
158 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE to their effector genes, but for complex disease the most sophisticated approaches only guess the correct gene a little over half of the tim...
Ch9 · Pt11 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 159 70 60 50 40 30 IL1RL1 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 20 -log 10(p-value) 16 12 8 4 0 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 Figure 9.4 Manhattan plot example for a genome-wide association analysis of asthma in UK Biobank. The association analysis compares genotypes at millions of genetic variants between asthma cases, defined on the basis of diagnosis codes in their individual medical records, and controls. The results of each comparison are summarized in a p-value which is plotted in the graph. The top few association signals have been labeled with the name of the nearest gene. (The figure is based on the UK Biobank association analyses reported in Taliun D, Harris DN, Kessler MD, et al: Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program, Nature 590: 290–299, 2021. https:// doi.org/10.1038/s 41586-021-03205-y.) Meanwhile, the clinical use of PRS is not without caveats. A main controversy is the potential health disparity resulting from the fact that most large GWAS and PRS have been conducted and tested only in large samples of European ancestry. Second, the interpretation of a PRS score (e.g., relative vs. absolute risk) can sometimes be a barrier in clinical settings, requiring continued education of clinicians to help them to better disclose results to families. Third, it remains an open question how best to combine PRS and modifiable clinical risk factors to inform disease prevention. Finally, standardized reporting of PRS results will facilitate knowledge translation. EXAMPLES OF COMMON MULTIFACTORIAL DISEASES WITH A GENETIC CONTRIBUTION In this section and the next, we turn to considering examples of several common conditions that illustrate general concepts of multifactorial disorders and their complex inheritance, as summarized in the accompanying box (see Box 9.1). BOX 9.1 CHARACTERISTICS OF INHERITANCE OF COMPLEX DISEASES Genetic variation contributes to diseases with complex inheritance, but these diseases are not single-gene disorders and do not demonstrate a simple mendelian pattern of inheritance. Diseases with complex inheritance often demonstrate familial aggregation because close relatives of an affected individual are likely to also share disease-predisposing alleles. Diseases with complex inheritance are more common among the close relatives of a proband and become less common in relatives who are less closely related and therefore share fewer predisposing alleles. Greater concordance for disease is expected among monozygotic versus dizygotic twins. However, pairs of relatives who share disease-­ predisposing genotypes at relevant loci may still be discordant for phenotype (show lack of penetrance) because of the crucial role of nongenetic factors in disease causation. The most extreme examples of lack of penetrance despite identical genotypes are discordant monozygotic twins.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 159 70 60 50 40 30 IL1RL1 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 20 -log 10(p-value) 16 12 8 4 0 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8...
Ch9 · Pt12 160 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE IL1RL1 0 0 8 16 24 32 40 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 4 8 12 16 20 30 40 50 60 70 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 -log 10(p-value) -log 10(p-value) Figure 9.5 Miami plot example for a genome-wide association analysis of asthma and nasal polyps in UK Biobank. The association analysis on the top panel compares genotypes at millions of genetic variants between asthma cases, defined on the basis of diagnosis codes in their individual medical records, and controls. The analysis on the bottom panel compares genotypes at the same variants between individuals with nasal polyps and controls. Note the many shared signals between the two analyses. The results of each comparison are summarized in a p-value, which is plotted in the graph. The top few association signals have been labeled with the name of the nearest gene in the asthma portion of the plot. (The figure is based on the UK biobank association analyses reported in Taliun D, Harris DN, Kessler MD, et al: Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program, Nature 590: 290–299, 2021. https://doi.org/10.1038/s 41586-021-03205-y.)
160 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE IL1RL1 0 0 8 16 24 32 40 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 4 8 12 16 20 30 40 50 60 70 D2HGDH TSLP HLA-DQB1 IL33 GA...
Ch9 · Pt13 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 161 Figure 9.6 Brisbane plot example for a genome-wide association analysis of height in 5.3-m individuals. Each dot shows one of the 12,111 independent genome-wide significant variants associated with height. Density was calculated as the number of other independent associated variants within 100 kb. (Figure from height GWAS in Yengo L, Vedantam S, Marouli E, et al: A saturated map of common genetic variants associated with human height, Nature 610(7933):704–712, 2022. https:doi.org/10.1038/ s 41586-022-05275-y.) Multifactorial Congenital Malformations Many common congenital malformations, occurring as isolated defects and not as part of a syndrome, are multifactorial and demonstrate complex inheritance (Table 9.6). Among these, congenital heart malformations are some of the most common and serve to illustrate the current state of understanding of other categories of congenital malformation. Congenital heart defects (CHDs) occur at a frequency of ~4 to 8 per 1000 births. They are a heterogeneous group, caused in some cases by single-gene or chromosomal mechanisms and in others by exposure to teratogens, such as rubella infection or maternal diabetes. The cause is usually unknown, however, and the majority of cases are believed to be multifactorial in origin. There are many types of CHDs, with different population incidences and empirical risks. It is known that when heart defects recur in a family, however, the affected children do not necessarily have exactly the same anatomic defect but instead show recurrence of lesions that are similar with regard to developmental mechanisms (see Chapter 14). By using developmental mechanisms as a classification scheme, five main groups of CHDs can be distinguished: Flow lesions Defects in cell migration TABLE 9.6 Some Common Congenital Malformations With Multifactorial Inheritance Malformation Approximate Population Incidence (per 1000) Cleft lip with or without cleft palate 0.4–1.7 Cleft palate 0.4 Congenital dislocation of hip 2* Congenital heart defects 4–8 Ventricular septal defect 1.7 Patent ductus arteriosus 0.5 Atrial septal defect 1.0 Aortic stenosis 0.5 Neural tube defects 2–10 Spina bifida and anencephaly Variable Pyloric stenosis 1,† 5* *Per 1000 males. †Per 1000 females. Data from Carter CO: Genetics of common single malformations, Br Med Bull 32:21–26, 1976; Nora JJ: Multifactorial inheritance hypothesis for the etiology of congenital heart diseases: The genetic environmental interaction, Circulation 38:604–617, 1968; Lin AE, Garver KL: Genetic counseling for congenital heart defects, J Pediatr 113:1105–1109, 1988.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 161 Figure 9.6 Brisbane plot example for a genome-wide association analysis of height in 5.3-m individuals. Each dot shows one of the...
Ch9 · Pt14 162 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Defects in cell death Abnormalities in extracellular matrix Defects in targeted growth The subtype of congenital heart malformations known as flow lesions illustrates the familial aggregation and elevated risk for recurrence in relatives of an affected individual, all characteristic of a complex trait (Table 9.7). Flow lesions, which constitute ~50% of all CHDs, include hypoplastic left heart syndrome, coarctation of the aorta, atrial septal defect of the secundum type, pulmonary valve stenosis, a common type of ventricular septal defect, and other forms (Fig. 9.7). Up to 25% of individuals with flow lesions, particularly tetralogy of Fallot, may have the deletion of chromosome region 22q11 seen in the velocardiofacial syndrome (see Chapter 6). Certain isolated CHDs are inherited as multifactorial traits. Until more is known, the figures shown in Table 9.7 can be used as estimates of the recurrence risk for flow lesions in first-degree relatives. There is, however, a rapid falloff in risk (to levels not much higher than the population risk) in second- and third-degree relatives of index patients with flow lesions. Similarly, relatives of index patients with types of CHDs other than flow lesions can be offered reassurance that their risk is not much greater than that of the general population. For further reassurance, many CHDs can now be assessed prenatally by ultrasonography (see Chapter 18). TABLE 9.7 Population Incidence and Recurrence Risks for Various Flow Lesions Defect Population Incidence (%) Frequency in Sibs (%) λs Ventricular septal defect 0.17 4.3 25 Patent ductus arteriosus 0.083 3.2 38 Atrial septal defect 0.066 3.2 48 Aortic stenosis 0.044 2.6 59 Normal Atrial septal defect Coarctation of the aorta Tetralogy of Fallot Patent ductus arteriosus (PDA) Hypoplastic left heart AO PA LA RA RV LV AO PA LA LV RA RV RV RA LA PA AO LV AO PA LA RA RV LV AO PA PDA LA RA RV LV AO PA LA LV RV RA PDA Figure 9.7 Diagram of various flow lesions seen in congenital heart disease. Blood on the left side of the circulation is shown in red, on the right side in blue. Abnormal admixture of oxygenated and deoxygenated blood is purple. AO, Aorta; LA, left atrium; LV, left ventricle; PA, pulmonary artery; RA, right atrium; RV, right ventricle. Neuropsychiatric Disorders Mental illnesses are some of the most common and perplexing of human diseases, affecting 4% of the human population worldwide. As of 2020, the annual cost in medical care and social services exceeds $200 billion in the United States alone. Among the most severe of the mental illnesses are schizophrenia and bipolar disease (manic-depressive illness). Schizophrenia affects 1% of the world’s population. It is a devastating psychiatric illness, with onset commonly in late adolescence or young adulthood, and is characterized by abnormalities in thought, emotion, and social relationships, often associated with delusional thinking and disordered mood. A genetic contribution
162 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Defects in cell death Abnormalities in extracellular matrix Defects in targeted growth The subtype of congenital heart malformations known a...
Ch9 · Pt15 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 163 TABLE 9.8 Recurrence Risks and Relative Risk Ratios in Schizophrenia Families Relation to Individual Affected by Schizophrenia Recurrence Risk (%) λr Child of two schizophrenic parents 46 23 Child 9–16 11.5 Sibling 8–14 11 Nephew or niece 1–4 2.5 Uncle or aunt 2 2 First cousin 2–6 4 Grandchild 2–8 5 to schizophrenia is supported by both twin and family aggregation studies. MZ concordance in schizophrenia is estimated to be 40% to 60%; DZ concordance is 10% to 16%. The recurrence risk ratio is elevated in first- and second-degree relatives of individuals with schizophrenia (Table 9.8). Although there is considerable evidence of a genetic contribution to schizophrenia, only a subset of the genes and alleles that predispose to the disease has been identified to date. A major exception is the small percentage (<2%) of all schizophrenia that is found in individuals with interstitial deletions of particular chromosomes, such as the 22q11 deletion responsible for the velocardiofacial syndrome. It is estimated that 25% of individuals with 22q11 deletions develop schizophrenia, even in the absence of many or most of the other physical signs of the syndrome. The mechanism by which a deletion of 3 Mb of DNA on 22q11 (see Fig. 6.7) causes mental illness in individuals with this syndrome is unknown. Chromosomal microarrays have been used to scan the entire genome for other deletions and duplications, many too small to be detectable by standard cytogenetic approaches, as introduced in Chapter 5. These studies have revealed numerous deletions and duplications (copy number variants) throughout the genome in both normal individuals and individuals with a variety of psychiatric and neurodevelopmental disorders (see Chapter 6). In particular, small (1–1.5 Mb) interstitial deletions at 1q21.1, 15q11.2, and 15q13.3 have been implicated repeatedly in a small fraction of individuals with schizophrenia. For the vast majority of people with schizophrenia, however, genetic lesions are not known, and counseling therefore relies on empirical risk figures (see Table 9.8). Bipolar disease is predominantly a mood disorder in which episodes of mood elevation, grandiosity, high-risk dangerous behavior, and inflated self-esteem (mania) alternate with periods of depression, decreased interest in what are normally pleasurable activities, feelings of worthlessness, and suicidal thinking. The prevalence of bipolar disease is 0.8%, approximately equal to that of schizophrenia, with a similar age at onset. The ­seriousness of this condition is underscored by the high rate of suicide in affected individuals. A genetic contribution to bipolar disease is strongly supported by twin and family aggregation studies. MZ twin concordance is 40% to 60%; DZ twin concordance is 4% to 8%. Disease risk is also elevated in relatives of affected individuals (Table 9.9). One striking aspect of bipolar disease in families is that the condition has variable expressivity; some members of the same family demonstrate classic bipolar illness, others have depression alone (unipolar disorder), and others carry a diagnosis of a psychiatric syndrome that involves both thought and mood (schizoaffective disorder). Even less is known about genes and alleles that predispose to bipolar disease than is known for schizophrenia; in particular, although an increase in de novo deletions or duplications has been identified in bipolar psychosis, recurrent copy number variants involving particular regions of the genome have not been identified. Counseling therefore typically relies on empirical risk figures (see Table 9.9). Coronary Artery Disease Coronary artery disease (CAD) kills ~500,000 individuals in the United States yearly and is one of the most frequent causes of morbidity and mortality in the developed world. CAD due to atherosclerosis is the major cause of the nearly 1.5 million cases of myocardial infarction (MI) and the more than 200,000 deaths from acute MI occurring annually. In the aggregate, CAD costs more than $143 billion in health care expenses alone each year in the United States, not including lost productivity. For unknown reasons, males are at higher risk for CAD both in the general population and within affected families. Family studies have repeatedly supported a role for heredity in CAD, particularly when it occurs in relatively young individuals. The pattern of increased risk suggests that when the proband is female or young there is likely to be a greater genetic contribution to MI in the family, thereby increasing the risk for disease in the proband’s relatives. For example, the recurrence risk (Table 9.10) in male first-degree relatives of a female proband is sevenfold greater than that in the general population, compared with the 2.5-fold increased risk in female relatives of a male proband. TABLE 9.9 Recurrence Risks and Relative Risk Ratios in Bipolar Disorder Families Relation to Individual Affected With Bipolar Disease Recurrence Risk (%)* λr Child of two parents with bipolar disease 50–70 75 Child 27 34 Sibling 20–30 31 Second-degree relative 5 6 *Recurrence of bipolar, unipolar, or schizoaffective disorder.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 163 TABLE 9.8 Recurrence Risks and Relative Risk Ratios in Schizophrenia Families Relation to Individual Affected by Schizophrenia Re...
Ch9 · Pt16 164 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE A few mendelian disorders leading to CAD are known. Familial hypercholesterolemia (Case 16), an autosomal dominant defect of the low-density lipoprotein (LDL) receptor discussed in Chapter 13, is one of the most common of these but accounts for only ~5% of survivors of MI. Most cases of CAD show multifactorial inheritance, with both nongenetic and genetic predisposing factors. There are many stages in the evolution of atherosclerotic lesions in the coronary artery. What begins as a fatty streak in the intima of the artery evolves into a fibrous plaque containing smooth muscle, lipid, and fibrous tissue. These intimal plaques become vascular and may bleed, ulcerate, and calcify, thereby causing severe vessel narrowing as well as providing fertile ground for thrombosis, resulting in sudden, complete occlusion and MI. Given the many stages in the evolution of atherosclerotic lesions in the coronary artery, it is not surprising that many genetic differences affecting the various pathologic processes involved could predispose to or protect from CAD (Fig. 9.8; also see box). Additional risk factors for CAD include other disorders that are themselves multifactorial with genetic components, such as hypertension, TABLE 9.10 Risk for Coronary Artery Disease in Relatives of a Proband Proband Increased Risk for CAD in a Family Member* Male 3-fold in male first-degree relatives 2.5-fold in female first-degree relatives Female 7-fold in male first-degree relatives Female <55 yr 11.4-fold in male first-degree relatives Two male relatives <55 yr 13-fold in first-degree relatives *Relative to the risk in the general population. CAD, Coronary artery disease. Data from Silberberg JS: Risk associated with various definitions of family history of coronary heart disease, Am J Epidemiol 147:1133–1139, 1998. TABLE 9.11 Twin Concordance Rates and Relative Risks for Fatal Myocardial Infarction When Proband Had Early Fatal Myocardial Infarction* Sex of the Twins Concordance MZ Twins Increased Risk† in an MZ Twin Concordance DZ Twins Increased Risk† in a DZ Twin Male 0.39 6–8-fold 0.26 3-fold Female 0.44 15-fold 0.14 2.6-fold *Early myocardial infarction defined as age <55 years in males, age <65 years in females. †Relative to the risk in the general population. DZ, Dizygotic; MZ, monozygotic. Data from Marenberg ME: Genetic susceptibility to death from coronary heart disease in a study of twins, NEJM 330:1041–1046, 1994. Normal Fatty streaks White blood cells White blood cells Platelets and fibrin Thrombus Calcium Red blood cells Lipid-rich plaque Inflammation and calcification Scar development with calcification Scar Early Lipid rich Internal rupture Calcified shell Calcified plaque Vulnerable Rupture Thrombus Myocardial infarction Obstruction Figure 9.8 Sections of coronary artery demonstrating the steps leading to coronary artery disease. Genetic and environmental factors operating at any or all of the steps in this pathway can contribute to the development of this complex, common disease. (Modified from an original figure courtesy Larry Almonte, with permission.) When the proband is young (<55 years) and female, the risk for CAD is more than 11 times greater than that of the general population. Having multiple relatives affected at a young age increases risk substantially as well. Twin studies also support a role for genetic variants in CAD (Table 9.11).
164 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE A few mendelian disorders leading to CAD are known. Familial hypercholesterolemia (Case 16), an autosomal dominant defect of the low-density...
Ch9 · Pt17 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 165 obesity, and diabetes mellitus. The metabolic and physiologic derangements represented by these disorders also contribute to enhancing the risk for CAD. Finally, diet, physical activity, systemic inflammation, and smoking are environmental factors that also play a major role in influencing the risk for CAD. Given all the different processes, metabolic derangements, and environmental factors that contribute to the development of CAD, it is easy to imagine that genetic susceptibility to CAD could be a complex multifactorial condition (see Box 9.2). or how a particular genotype and set of environmental influences interact to cause a disease or to determine the value of a particular physiologic measurement. In most cases, all we can show is that there is some genetic contribution and estimate its magnitude. These genetic contributions reflect the aggregate effects of many variants that might number in hundreds or thousands. There are, however, a few multifactorial diseases with complex inheritance for which we have begun to identify the genetic and, in some cases, environmental factors responsible for increasing disease susceptibility. We give a few examples in the next part of this chapter, illustrating increasing levels of complexity. Modifier Genes in Mendelian Disorders As discussed in Chapter 7, allelic variation at a single locus can explain variation in the phenotype in many single-gene disorders. However, even for well-characterized mendelian disorders known to be due to defects in a single gene, variation at other gene loci may impact some aspect of the phenotype, illustrating features of complex inheritance. In cystic fibrosis (CF) (Case 12), for example, whether an individual has pancreatic insufficiency requiring enzyme replacement can be explained largely by which genetic changes are present in the CFTR gene (see Chapter 13). The correlation is imperfect, however, for other phenotypes. For example, the variation in the degree of pulmonary disease seen in CF patients remains unexplained by allelic heterogeneity. It has been proposed that the genotype at other genetic loci could act as genetic modifiers, that is, genes whose alleles have an effect on the severity of pulmonary disease seen in CF. For example, reduction in forced expiratory volume after 1 second (FEV1), calculated as a percentage of the value expected for CF patients (a CF-specific FEV1 percent), is a quantitative trait commonly used to measure deterioration in pulmonary function in CF. A comparison of CF-specific FEV1 percent in affected MZ versus affected DZ twins provides an estimate of the heritability of the severity of lung disease in CF patients of ~50%. This value is independent of the specific CFTR allele(s) (because both kinds of twins will have the same pathogenic CF variants). Two loci harboring alleles responsible for modifying the severity of pulmonary disease in CF are known: MBL2, a gene that encodes a serum protein called mannose-binding lectin; and the TGFB1 locus encoding the cytokine transforming growth factor β (TGFβ). Mannose-binding lectin is a plasma protein in the innate immune system that binds to many pathogenic organisms and aids in their destruction by phagocytosis and complement activation. A number of common alleles that result in reduced blood levels of the lectin exist at the MBL2 locus in European populations. Lower levels of mannose-binding lectin appear BOX 9.2 GENES AND GENE PRODUCTS INVOLVED IN THE STEPWISE PROCESS OF CORONARY ARTERY DISEASE A large number of genes and gene products have been suggested and, in some cases, implicated in promoting one or more of the developmental stages of coronary artery disease. These include genes involved in the following: Serum lipid transport and metabolism – cholesterol, apolipoprotein E, apolipoprotein C-III, the low-density lipoprotein (LDL) receptor, and lipoprotein(a) – as well as total cholesterol level. Elevated LDL cholesterol level and triglyceride levels, both of which elevate the risk for coronary artery disease, are themselves quantitative traits with significant heritabilities. Vasoactivity, such as angiotensin-converting enzyme. Blood coagulation, platelet adhesion, and fibrinolysis, such as plasminogen activator inhibitor 1, and the platelet surface glycoproteins Ib and IIIa. Inflammatory and immune pathways. Arterial wall components.
CAD is often an incidental finding in family histories of individuals with other genetic diseases. In view of the high recurrence risk, physicians and genetic counselors may need to consider whether f...
Ch9 · Pt18 166 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE associated with worse outcomes for CF lung disease, perhaps because low levels of lectin result in difficulties with containing respiratory pathogens, particularly Pseudomonas. Alleles at the TGFB1 locus that result in higher TGFβ production are also associated with worse outcome, perhaps because TGFβ promotes lung scarring and fibrosis after inflammation. Thus both MBL2 and TGFB1 are modifier genes, variants at which – while they do not cause CF – can modify the clinical phenotype associated with disease-causing alleles at the CFTR locus. Digenic Inheritance The next level of complexity is a disorder determined by the additive effect of the genotypes at two or more loci. One clear example of such a disease phenotype has been found in a few families of patients with a form of retinal degeneration called retinitis pigmentosa (RP) (Fig. 9.9). Affected individuals in these families are heterozygous for pathogenic alleles at two different loci (double heterozygotes). One locus encodes the photoreceptor membrane protein peripherin and the other encodes a related photoreceptor membrane protein called Rom 1. Heterozygotes for only one or the other of these mutations are unaffected. Thus the RP in this family is caused by the simplest form of multigenic inheritance, inheritance due to the effect of variant alleles at two loci, without any known environmental factors that greatly influence disease occurrence or severity. The proteins encoded by these two genes are likely to have overlapping physiologic function because they are both located in the stacks of membranous disks found in retinal photoreceptors. It is the additive effect of having an abnormality in two proteins with overlapping function that produces disease. A multigenic model has also been proposed in a few families with Bardet-Biedl syndrome, a rare birth defect characterized by obesity, variable degrees of intellectual disability, retinal degeneration, polydactyly, and genitourinary malformations. Fourteen different genes have been found in which pathogenic variants cause the syndrome. Although inheritance is clearly autosomal recessive in most families, a few families appear to demonstrate digenic inheritance, in which the disease occurs only when an individual is homozygous or compound heterozygous for mutations at one of these 14 loci and is heterozygous for a variant at another of the loci. Gene-Environment Interactions in Venous Thrombosis Another example of gene-gene interaction predisposing to disease is found in the group of conditions referred to as hypercoagulability states, in which venous or arterial clots form inappropriately and cause life-­threatening complications of thrombophilia (Case 46). With hypercoagulability, however, there is a third factor, an environmental influence that in the presence of the predisposing genetic factors increases the risk for disease even more. One such disorder is idiopathic cerebral vein thrombosis, a disease in which clots form in the venous system of the brain, causing catastrophic occlusion of cerebral veins in the absence of an inciting event such as infection or tumor. It affects young adults, and although quite rare (<1 per 100,000 in the population), it carries a high mortality rate (5–30%). Three relatively common factors – two genetic and one environmental – that lead to abnormal coagulability of the clotting system are each known to individually increase the risk for cerebral vein thrombosis (Fig. 9.10): 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/1 Genotype: peripherin: 1 or mut ROM1: 1 or mut 1/1 1/mut 1/mut 1/mut Figure 9.9 Pedigree of a family with retinitis pigmentosa due to digenic inheritance. Dark blue symbols are affected individuals. Each individual’s genotypes at the peripherin locus (first line) and ROM1 locus (second line) are written below each symbol. The normal allele is 1; the allele carrying a mutation is mut. Light blue symbols are unaffected, despite carrying a pathogenic variant in one or the other gene. (Redrawn from Kajiwara K, Berson EL, Dryja TP: Digenic retinitis pigmentosa due to pathogenic variants at the unlinked peripherin/RDS and ROM1 loci, Science 264:1604–1608, 1994.)
166 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE associated with worse outcomes for CF lung disease, perhaps because low levels of lectin result in difficulties with containing respiratory...
Ch9 · Pt19 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 167 A missense variant in the gene for the clotting factor, factor V A variant in the 3′ untranslated region (UTR) of the gene for the clotting factor prothrombin The use of oral contraceptives A common allele of factor V, factor V Leiden (FVL) (Case 46), in which arginine is replaced by glutamine at position 506 (Arg 506Gln), has a frequency of ~2.5% in populations of European origin but is rarer in other population groups. This alteration affects a cleavage site used to degrade factor V, thereby making the protein more stable and able to exert its procoagulant effect for a longer duration. Heterozygous carriers of FVL, ~5% of the European population, have a risk for cerebral vein thrombosis that, although still quite low, is 7-fold higher than that in the general population; homozygotes have a risk that is 80-fold higher. The second genetic risk factor, a pathogenic variant in the prothrombin gene, changes a G to an A at position 20210 in the 3′ UTR of the gene (prothrombin g.20210 G>A). Approximately 2.4% of individuals of European ancestry are heterozygotes, but it is rare in other groups. This change appears to increase the level of prothrombin mRNA, resulting in increased translation and elevated levels of the protein. Being heterozygous for the prothrombin 20210 G>A allele raises the risk for cerebral vein thrombosis three- to sixfold. The use of oral contraceptives containing synthetic estrogen increases the risk for thrombosis 14- to 22-fold, independent of genotype at the factor V and prothrombin loci, probably by increasing the levels of many clotting factors in the blood. Although using oral contraceptives and being heterozygous for FVL cause only a modest increase in risk compared with either factor alone, oral contraceptive use in a heterozygote for prothrombin 20210 G>A raises the relative risk for cerebral vein thrombosis 30- to 150-fold! There is also interest in the role of FVL and prothrombin 20210 G>A alleles in deep venous thrombosis (DVT) of the lower extremities, a condition that occurs in ~1 in 1000 individuals per year, far more common than idiopathic cerebral venous thrombosis. Mortality due to DVT (primarily due to pulmonary embolus) can be up to 10%, depending on age and the presence of other medical conditions. Many environmental factors are known to increase the risk for DVT and include trauma, surgery (particularly orthopedic surgery), malignant disease, prolonged periods of immobility, oral contraceptive use, and advanced age. The FVL allele increases the relative risk for a first episode of DVT 7-fold in heterozygotes; heterozygotes who use oral contraceptives see their risk increased 30-fold compared with controls. Heterozygotes for prothrombin 20210 G>A also have an increase in their relative risk for DVT of two- to threefold. Notably, double heterozygotes for FVL and prothrombin 20210 G>A have a relative increased risk of 20-fold – a risk approaching a few percent of the population. Thus each of these three factors, two genetic and one environmental, on its own increases the risk for an abnormal hypercoagulable state; having two or all three of these factors at the same time raises the risk even more, to the point that thrombophilia screening programs for selected populations may be indicated in the future. Multiple Coding and Noncoding Elements in Hirschsprung Disease A more complicated set of interacting genetic factors has been described in the pathogenesis of a developmental abnormality of the enteric nervous system in the gut known as Hirschsprung disease (HSCR). In HSCR, there is complete absence of some or all of the intrinsic ganglion cells in the myenteric and submucosal plexuses of the colon. An aganglionic colon is incapable of peristalsis, resulting in severe constipation, symptoms of intestinal obstruction, and massive dilatation of the colon (megacolon) proximal to the aganglionic segment. The disorder affects ~1 in 5000 newborns of European ancestry but is twice as common among Asian infants. HSCR occurs as an isolated birth defect 70% of the time, as part of a chromosomal syndrome 12% of the time, and as one element of a broad constellation of congenital abnormalities in the remainder of cases. Among individuals with HSCR as an isolated birth defect, 80% have only a single, short aganglionic segment of colon at the level of the rectum (hence, HSCR-S), whereas 20% Factor Xa Factor Va Factor V Factor X (↑OC) Thrombin Prothrombin (↑OC) Fibrin Clot Fibrinogen Intrinsic pathway Extrinsic pathway Figure 9.10 The clotting cascade relevant to factor V Leiden and prothrombin variants. Once factor X is activated, through either the intrinsic or extrinsic pathway, activated factor V promotes the production of the coagulant protein thrombin from prothrombin, which in turn cleaves fibrinogen to generate fibrin required for clot formation. Oral contraceptives (OC) increase blood levels of prothrombin and factor X as well as a number of other coagulation factors. The hypercoagulable state can be explained as a synergistic interaction of genetic and environmental factors that increase the levels of factor V, prothrombin, factor X, and others to promote clotting. Activated forms of coagulation proteins are indicated by the letter a. Solid arrows are pathways; dashed arrows are stimulators.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 167 A missense variant in the gene for the clotting factor, factor V A variant in the 3′ untranslated region (UTR) of the gene for th...
Ch9 · Pt20 168 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE have aganglionosis of a long segment of colon, the entire colon or, occasionally, the entire colon plus the ileum (hence, HSCR-L). Familial HSCR-L is often characterized by patterns of inheritance that suggest dominant or recessive inheritance, but consistently with reduced penetrance. HSCR-L is most commonly caused by loss-of-function missense or nonsense mutations in the RET gene, which encodes RET, a receptor tyrosine kinase. A small minority of families have pathogenic variants in genes encoding ligands that bind to RET, but with even lower penetrance than those families with RET variants. HSCR-S is the more common type of HSCR and has many of the characteristics of a disorder with complex genetics. The relative risk ratio for sibs, λs, is very high (~200), but MZ twins do not show perfect concordance, and families do not show any obvious mendelian inheritance pattern for the disorder. When pairs of siblings concordant for HSCR-S were analyzed genome wide to see which loci and which sets of alleles at these loci each sib had in common with an affected brother or sister, alleles at three loci (including RET) were found to be significantly shared, suggesting gene-gene interactions and/or multigenic inheritance; indeed, most of the concordant sibpairs were found to share alleles at all three loci. Although the non-RET loci have yet to be identified, Fig. 9.11 illustrates the range of interactions necessary to account for much of the penetrance of HSCR in even this small cohort of patients. Pathogenic variants of HSCR mutations have now been described at over a dozen loci, with RET mutations being by far the most common. The current data suggest that the RET gene is implicated in nearly all individuals with HSCR and, in particular, have pointed to two interacting noncoding regulatory variants near the RET gene, one in a potent gut enhancer with a binding site for the relevant transcription factor SOX10 and the other at an even more distant noncoding site some 125 kb upstream of the RET transcription start site. Thus HSCR-S is a multifactorial disease that results from mutations in or near the RET locus, perturbing the normally tightly controlled process of enteric nervous system development, combined with pathogenic variants at a number of other loci, both known, such as EDNRB and GDNF, and still unknown. Current genomic approaches of the type discussed in Chapter 11 suggest the possibility that many dozens of additional genes could be involved. The identification of common, low-penetrant variants in noncoding elements serves to illustrate that the gene variants responsible for modifying expression of a multifactorial trait may be subtle in how they exert their effects on gene expression and, as a consequence, on disease penetrance and expressivity. It is also sobering to realize that the underlying genetic mechanisms for this relatively well-defined congenital malformation have turned out to be so surprisingly complex; still, they are likely to be far simpler than are the mechanisms involved in the more common complex diseases, such as diabetes. Type 1 Diabetes Mellitus A common complex disease for which some of the underlying genetic architecture is being delineated is diabetes mellitus. Diabetes occurs in two major forms: type 1 (T1D) (sometimes referred to as insulin dependent [IDDM]) and type 2 (T2D) (sometimes referred to as non–insulin dependent [NIDDM]), representing ~10% Figure 9.11 Current understanding of the genetic factors underlying inheritance of HSCR. As with most complex traits, genetic susceptibility can be explained by common variants with relatively small impacts on individual risk, and a spectrum towards very rare alleles with a substantial impact on risk. HSCR shows phenotypic severity differences - whereas syndromic HSCR and TCA/L-HSCR are more rare and severe, whereas sporadic S-HSCR is more common. (From Karim A, Tang CS, Tam PK. The Emerging Genetic Landscape of Hirschsprung Disease and Its Potential Clinical Applications. Front Pediatr 2021;9:638093.)
168 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE have aganglionosis of a long segment of colon, the entire colon or, occasionally, the entire colon plus the ileum (hence, HSCR-L). Familial...
Ch9 · Pt21 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 169 and 88% of all cases, respectively. Familial aggregation is seen in both types of diabetes, but in any given family usually only T1D or T2D is present. They differ in typical onset age, MZ twin concordance, and association with particular genetic variants at particular loci. Here, we focus on T1D to illustrate the major features of complex inheritance in diabetes. T1D has an incidence in the population of European ancestry of ~2 per 1000 (0.2%), but this is lower in African and Asian ancestry populations. It usually manifests in childhood or adolescence. It results from autoimmune destruction of the β cells of the pancreas, which normally produce insulin. A large majority of children who will go on to have T1D develop multiple autoantibodies early in childhood against a variety of endogenous proteins, including insulin, well before they develop overt disease. There is strong evidence for genetic factors in T1D: concordance among MZ twins is ~40%, which far exceeds the 5% concordance in DZ twins. The lifetime risk for T1D in siblings of an affected proband is ~7%, resulting in an estimated λs of ≈35. However, the earlier the age of onset of the T1D in the proband, the greater is λs. The Major Histocompatibility Complex The major genetic factor in T1D is the major histocompatibility complex (MHC) locus, which spans some 3 Mb on chromosome 6 and is the most highly polymorphic locus in the human genome, with over 200 known genes (many involved in immune functions) and well over 2000 alleles known in populations around the globe (Fig. 9.12). On the basis of structural and functional differences, two major subclasses, the class I and class II genes, correspond to the human leukocyte antigen (HLA) genes, originally discovered by virtue of their importance in tissue transplantation between unrelated individuals. The HLA class I (HLA-A, HLA-B, HLA-C) and class II (HLA-DR, HLA-DQ, HLA-DP) genes encode cell surface proteins that play a critical role in the presentation of antigen to lymphocytes, which cannot recognize and respond to an antigen unless it is complexed with an HLA molecule on the surface of an antigen-presenting cell. Within the MHC, the HLA class I and class II genes are by far the most highly polymorphic loci (see Fig. 9.12). The original studies showing an association between T1D and alleles designated as HLA-DR3 and HLA-DR4 relied on a serologic method in use at that time for distinguishing between different HLA alleles, one that was based on immunologic reactions in a test tube. This method has long been superseded by direct determination of the DNA sequence of different alleles, and sequencing of the MHC in a large number of individuals has revealed that the serologically determined “alleles” associated with T1D are not single alleles at all (see Box 9.3). Both DR3 and DR4 can be subdivided into a dozen or more alleles located at a locus now termed HLA-DRB1. The set of HLA alleles at the different class I and class II loci on a given chromosome together form a haplotype. Within any one ancestral group, some HLA alleles and haplotypes are found commonly; others are rare or never seen. The differences in the distribution and frequency of the alleles and haplotypes within the MHC are the result of complex genetic, environmental, and historical factors at play in each of the different populations. The extreme levels of genetic variation at HLA loci and their resulting haplotypes have been extraordinarily useful for identifying associations of particular variants with specific diseases (see Chapter 11), many of which (as one might predict) are autoimmune disorders, associated with an abnormal immune response apparently directed against one or more self-antigens resulting from polymorphism in immune response genes. Furthermore, it is now clear that the association between certain DRB1 alleles and T1D is due, in part, to alleles at two other class II loci, DQA1 and DQB1, located ~80 kb away from DRB1, that form a particular combination of alleles with each other (i.e., a haplotype) that is typically inherited as a unit (due to linkage disequilibrium; see Chapter 11). DQA1 and DQB1 encode the α and β chains of the class II DQ protein. Certain combinations of alleles at these three loci form a haplotype that increases the risk for T1D more than 11-fold over that for the general population, whereas other combinations of alleles reduce the risk 50-fold. The DQB1*0303 allele contained in this protective haplotype results in the amino acid aspartic acid at position 57 of the DQB1 product, whereas other amino acids at this position (alanine, valine, or serine) confer susceptibility. In fact, ~90% of individuals with T1D are homozygous for DQB1 alleles that do not encode aspartic acid at position 57. It is likely that differences in antigen binding, determined by which amino acid is at position BOX 9.3 HUMAN ANTIGEN ALLELES AND HAPLOTYPES The human leukocyte antigen (HLA) system can be confusing at first because the nomenclature used to define and describe different HLA alleles has undergone a fundamental change with the advent of widespread DNA sequencing of the major histocompatibility complex (MHC). According to the older system of HLA nomenclature, the different alleles were distinguished from one another serologically. However, as the genes responsible for encoding the class I and class II MHC chains were identified and sequenced (see Fig. 9.12), single HLA alleles initially defined serologically were shown to consist of multiple alleles defined by different DNA sequence variants even within the same serological allele. The 100 serologic specificities at HLA-A, B, C, DR, DQ, and DP loci now comprise more than 1300 alleles defined at the DNA sequence level! For example, what used to be a single B27 allele defined serologically is now referred to as HLA-B*2701, HLA-B*2702, and so on, based on DNA-based genotyping.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 169 and 88% of all cases, respectively. Familial aggregation is seen in both types of diabetes, but in any given family usually only...
Ch9 · Pt22 170 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Gene annotations Simple nucleotide polymorphisms Strand db SNP>1% Base pairs 30,000,000 30,500,000 31,000,000 31,500,000 32,000,000 32,500,000 33,000,000 6 ZFP57 ZNRD1 AS1 RNF39 HLA-F HLA-G HLA-A HCG9 ZNRD1 PPP1R11 TRIM40 TRIM15 TRIM39 RPP21 HLA-E TRIM31 TRIM10 TRIM26 HCG17 HCG18 PRR3 ABCF1 MRPS188 ATAT1 GNL1 PPP1R10 DHX16 NRM MDC1 IER3 TIGD1L SFTA2 HCG21 HLA-C HLA-B CDSN PSORS1C2 CCHCR1 POU5F1 PSORS1C3 MICB HCP5 MICA HCG27 TUBB HCG20 DPCR1 MUC21 HCG22 PSORS1C1 TCF19 DDR1 GTF2H4 VARS2 HCG23 HLA-DRA HLA-DQA1 HLA-DQA2 PSM89 BRD2 HLA-DPB1 HCG24 HLA-DPA1 HLA-DOA HLA-DMB HLA-DMA C6orf 10 BTNL2 HLA-DR65 HLA-DR61 HLA-DQB1 HLA-DQ62 HLA-DOB TAP2 PSM68 TAP1 PPT2 EGFL8 RNF5 HSPA1A HSPA1B C2 CFB STK19 C4A C4B CVP21A2 AJF1 APOM CSNK2B LY6G6B LY6G6F LY6G6D MSH5 NFKBIL1 TNF LTA LST1 BAT1 ATP6V1G2 LTB NCR3 BAT3 BAT4 BAT5 LY6G6C DDAH2 CLIC1 VARS LSM2 HSPA1L NEU1 SLC44A4 EHMT2 RDBP DOM3Z TNX5 ATF68 PRRT1 AGPAT1 AGER PBX2 NOTCH4 Figure 9.12 Genomic landscape of the major histocompatibility complex (MHC). The classic MHC is shown on the short arm of chromosome 6, comprising the class I region (yellow) and class II region (blue), both enriched in human leukocyte antigen (HLA) genes. Sequence-level variation is shown for single nucleotide polymorphisms (SNPs) found with at least 1% frequency. Remarkably high levels of genetic variation are seen in regions containing the classic HLA genes where variation is enriched in coding exons involved in defining the antigen-binding cleft. Other genes (pink) in the MHC region show lower levels of genetic variation. db SNP, Minor allele frequency in the SNP database. (Modified from Trowsdale J, Knight JC: Major histocompatibility complex genomics and human disease, Annu Rev Genom Hum Genet 14:301–323, 2013.) TABLE 9.12 Empirical Risks for Counseling in Type 1 Diabetes Relationship to Affected Individual Risk for Development of Type 1 Diabetes (%) None 0.2 MZ twin 40 Sibling 7 Sibling with no DR haplotypes in common 1 Sibling with 1 DR haplotype in common 5 Sibling with 2 DR haplotypes in common 17* Child 4 Child of affected mother 3 Child of affected father 5 *20–25% for particular shared haplotypes. MZ, Monozygotic. 57, contribute directly to the autoimmune response that destroys the insulin-producing cells of the pancreas. Other loci and alleles in the MHC, however, are also important, as can be seen from the fact that some individuals with T1D have an aspartic acid at this position. Genes Other Than Class II Major Histocompatibility Complex Loci in Type 1 Diabetes The MHC haplotype alone accounts for only a portion of the genetic contribution to the risk for T1D in siblings of a proband. Family studies in T1D (Table 9.12) suggest that even when siblings share the same MHC class II haplotypes, the risk for disease is only ~17%, still well below the MZ twin concordance rate of ~40%. Thus there must be other genes elsewhere in the genome that contribute to the development of T1D (assuming that MZ twins and sibs have similar environmental exposures). Indeed, genetic association studies (to be described in Chapter 11) indicate that variation at more than 50 different loci around the genome can increase susceptibility to T1D, although most have very small effects on increasing disease susceptibility.
170 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Gene annotations Simple nucleotide polymorphisms Strand db SNP>1% Base pairs 30,000,000 30,500,000 31,000,000 31,500,000 32,000,000 32,500,0...
Ch9 · Pt23 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 171 It is important to stress, however, that genetic factors alone do not cause T1D because the MZ twin concordance rate is only ~40%, not 100%. Until a more complete picture develops of the genetic and nongenetic factors that cause T1D, risk counseling using HLA haplotyping must remain empirical (see Table 9.12). Alzheimer Disease Alzheimer disease (AD) (Case 4) is a fatal neurodegenerative disease that affects 1% to 2% of the US population. It is the most common cause of dementia in older adults and is responsible for more than half of all cases of dementia. As with other dementias, patients experience a chronic, progressive loss of memory and other cognitive functions, associated with loss of certain types of cortical neurons. Age, sex, and family history are the most significant risk factors for AD. Once a person reaches 65 years of age, the risk for any dementia, and AD in particular, increases substantially with age and female sex (Table 9.13). AD can be diagnosed definitively only postmortem, on the basis of neuropathologic findings of characteristic protein aggregates (β-amyloid plaques and neurofibrillary tangles; see Chapter 13). The most important constituent of the plaques is a small (39–42 amino acid) peptide, Aβ, derived from cleavage of a normal neuronal protein, the amyloid protein precursor. The secondary structure of Aβ gives the plaques the staining characteristics of amyloid proteins. In addition to three rare autosomal dominant forms of the disease (see Chapter 13), in which disease onset is in the third to fifth decade, there is a common form of AD with onset after the age of 60 years (late onset). This form has no obvious mendelian inheritance pattern but shows familial aggregation and an elevated relative risk ratio (λs = ≈4) typical of disorders with complex inheritance. Twin studies have been inconsistent but suggest MZ concordance of ~50% and DZ concordance of ~18%. The ε4 Allele of Apolipoprotein E The major locus with alleles found to be significantly associated with common late-onset AD is APOE, which encodes apolipoprotein E. Apolipoprotein E is a protein component of the LDL particle and is involved in clearing LDL through an interaction with high-affinity receptors in the liver. Apolipoprotein E is also a constituent of amyloid plaques in AD and is known to bind the Aβ peptide. The APOE gene has three alleles, ε2, ε3, and ε4, due to substitutions of arginine for two different cysteine residues in the protein (see Chapter 13). When the genotypes at the APOE locus were analyzed in individuals with AD and controls, a genotype with at least one ε4 allele was found two to three times more frequently among patients compared with controls in both the general US and Japanese populations (Table 9.14), with much less of an association in Hispanic and African populations. Even more striking is that the risk for AD appears to increase further if both APOE alleles are ε4, through an effect on the age at onset of AD; individuals with two ε4 alleles have an earlier onset of disease than those with only one. In a study of people with AD and unaffected controls, the age at which AD developed in the affected individuals was earliest for ε4/ε4 homozygotes, next for ε4/ε3 heterozygotes, and significantly less for the other genotypes (Fig. 9.13). In the population in general, the risk for developing AD by age 80 is approaching 10%. The ε4 allele is clearly a predisposing factor that increases the risk for development of AD by shifting the age at onset to an earlier age, such that ε3/ε4 heterozygotes have a 40% risk for developing the disease, and ε4/ε4 have a 60% risk by age 85. Despite this increased risk, other genetic and environmental factors must be important because a significant proportion of ε3/ε4 and ε4/ε4 individuals live to extreme old age with no evidence of AD. There are also reports of association between the presence of the ε4 allele and neurodegenerative disease after traumatic head injury (as seen in professional boxers, football players, and soldiers who have suffered blast injuries), indicating that at least one environmental factor, brain trauma, can interact with the ε4 allele in the pathogenesis of AD. TABLE 9.13 Cumulative Age- and Sex-Specific Risks for Alzheimer Disease and Dementia Time Interval Past Age 65 Yr Risk for Development of AD (%) Risk for Development of Any Dementia (%) 65–80 yr Male 6.3 10.9 Female 12 19 65–100 yr Male 25 32.8 Female 28.1 45 AD, Alzheimer disease. Data from Seshadri S, Wolf PA, Beiser A, et al: Lifetime risk of dementia and Alzheimer’s disease. The impact of mortality on risk estimates in the Framingham Study, Neurology 49:1498–1504, 1997. TABLE 9.14 Association of Apolipoprotein E ε4 Allele With Alzheimer Disease* Frequency Genotype United States Japan AD Control AD Control ε4/ε4; ε4/ε3; or ε4/ε2 0.64 0.31 0.47 0.17 ε3/ε3; ε2/ε3; or ε2/ε2 0.36 0.69 0.53 0.83 *Frequency of genotypes with and without the ε4 allele among Alzheimer disease (AD) patients and controls from the United States and Japan.
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 171 It is important to stress, however, that genetic factors alone do not cause T1D because the MZ twin concordance rate is only ~40%...
Ch9 · Pt24 172 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE The ε4 variant of APOE represents a prime example of a predisposing allele: It predisposes to a complex trait in a powerful way but does not predestine any individual carrying the allele to the disease. Additional genes as well as environmental effects are also clearly involved; although several of these appear to have a significant effect, most remain to be identified. In general, testing of asymptomatic people for the APOE ε4 allele remains controversial because knowing that one is a heterozygote or homozygote for the ε4 allele does not mean one will develop AD, and there are no preventative interventions to reduce disease risk in more susceptible individuals (see Chapter 19). THE CHALLENGE OF MULTIFACTORIAL DISEASE WITH COMPLEX INHERITANCE The greatest challenge facing medical genetics and genomic medicine going forward is unraveling the complex interactions between the variants at multiple loci and the relevant environmental factors that underlie the susceptibility to common multifactorial disease. This area of research is the central focus of the field of population-based genetic epidemiology (to be discussed more fully in Chapter 10). The field is developing rapidly, and it is clear that the genetic contribution to many more complex diseases in humans will be elucidated in the coming years. Such understanding will, in time, allow the development of novel preventive and therapeutic measures for the common disorders that cause such significant morbidity and mortality in the population. GENERAL REFERENCES Chakravarti A, Clark AG, Mootha VK: Distilling pathophysiology from complex disease genetics, Cell 155:21–26, 2013. Rimoin DL, Pyeritz RE, Korf BR: Emery and Rimoin's essential medical genetics, Waltham, MA, 2020, Academic Press (Elsevier). Scott W, Ritchie M: Genetic analysis of complex disease, ed 3, Hoboken, NJ, 2022, John Wiley and Sons. REFERENCES FOR SPECIFIC TOPICS Baylis RA, Smith NL, Klarin D, et al: Epidemiology and genetics of venous thromboembolism and chronic venous disease, Circ Res 128:1988–2002, 2021. Bellenguez C, Küçükali F, Jansen IE, et al: New insights into the genetic etiology of Alzheimer’s disease and related dementias, Nat Genet 54:412–436, 2022. Bigdeli TB, Fanous AH, Li Y, et al: Genome-wide association studies of schizophrenia and bipolar disorder in a diverse cohort of US veterans, Schizophr Bull 47(2):517–529, 2021. Grant SFA, Wells AD, Rich SS: Next steps in the identification of gene targets for type 1 diabetes, Diabetologia 63:2260–2269, 2020. Karim A, Tang CS, Tam PK: The emerging genetic landscape of Hirschsprung disease and its potential clinical applications, Front Pediatr 9:638093, 2021. https://doi.org/10.3389/fped.2021.638093 Khera AV, Chaffin M, Aragam KG, et al: Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations, Nat Genet 50:1219–1224, 2018. https:// www.nature.com/articles/s 41588-018-0183-z Matzaraki M, Kumar V, Wijmenga C, et al: The MHC locus and genetic susceptibility to autoimmune and infectious diseases, Genome Bio 18:76, 2017. Shoaib M, Ye Q, Iglay Reger H, et al: Evaluation of polygenic risk scores to differentiate between type 1 and type 2 diabetes, Genet Epidemiol. 2023. https://doi.org/10.1002/gepi.22521 Uffelmann E, Huang QQ, Munung NS, et al: Genome-wide association studies, Nat Rev Methods Primers 1:59, 2021. Wu P, Gifford A, Meng X, et al: Mapping ICD-10 and ICD-10-CM codes to phecodes: Workflow development and initial evaluation, JMIR Med Inform 7(4):e 14325, 2019. https://doi.org/10.2196/14325 Risk for Alzheimer disease 70 60 50 40 30 20 10 0 50 55 60 65 70 75 80 85 Age Population male Population female ε3/ε3 male ε3/ε3 female ε3/ε4 male ε3/ε4 female ε4/ε4 male ε4/ε4 female Figure 9.13 Chance of developing Alzheimer disease as a function of age for different APOE genotypes for each sex. At one extreme is the ε4/ε4 homozygote, who has a ≈40% chance of remaining free of the disease by the age of 85 years, whereas an ε3/ε3 homozygote has ≈70% to ≈90% chance of remaining disease free at the age of 85 years, depending on the sex. General population risk is also shown for comparison. (Modified from Roberts JS, Cupples LA, Relkin NR, et al: Genetic risk assessment for adult children of people with Alzheimer’s disease: the Risk Evaluation and Education for Alzheimer’s Disease (REVEAL) study, J Geriatr Psychiatry Neurol 18:250– 255, 2005.)
172 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE The ε4 variant of APOE represents a prime example of a predisposing allele: It predisposes to a complex trait in a powerful way but does not...
Ch9 · Pt25 CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 173 PROBLEMS 1. For a specific disease, the concordance rate in monozygotic (MZ) twins is 80% and the concordance rate in dizygotic (DZ) twins is 20%. What does this tell us about whether there are genetic or environmental contributions to susceptibility for this disease? 2. Match the genetic mapping technique that would be most cost-effective to find genetic predictors in each scenario: a. Searching for the diseasecausing variant in a large pedigree with 4-5 related individuals who have a rare disease i. Genome-wide association study using array-based genotyping and imputation b. Searching for common variants that increase susceptibility to a disease in a large case-control sample with well-studied ancestry ii. Exome sequencing with linkage analysis c. Searching for specific effector genes for a large number of diseases in a biobank sample iii. Exome or genome sequencing with genebased burden testing 3. One of the most important challenges in human genetics is keeping up with the rapidly evolving literature about each trait. Pick a complex disease of your choice and find the largest genome wide association study published in the last 5 years and compare it to one published 5 years before that. If you don’t have a favorite complex disease, consider atrial fibrillation as a potential example. a. How many disease loci were identified in each of the two studies? b. What was the strongest signal in each of the studies, defined by p-value? Summarize disease allele frequency, odds ratio and nearest gene of all genomewide hits. c. What was the strongest signal in each of the studies, defined by odds ratio? Summarize disease allele frequency, odds ratio and nearest gene. d. Did the two studies focus on similar questions? What new questions did the new study consider?
CHAPTER 9 — Complex Inheritance of Common Multifactorial Disorders 173 PROBLEMS 1. For a specific disease, the concordance rate in monozygotic (MZ) twins is 80% and the concordance rate in dizygotic...

Chapter 10: Population Genetics for Genomic Medicine

Ch10 · Pt1 chapter 10 Population Genetics for Genomic Medicine Alice B. Popejoy INTRODUCTION TO POPULATION GENETICS FOR GENOMIC MEDICINE We have explored in previous chapters the nature of genetic and genomic variation, mechanisms and types of mutation, and the inheritance of alleles (genetic variants) in families. Throughout, we have alluded to observed differences in allele frequencies across the globe, whether assessed by examining single nucleotide variants (SNVs), insertions and deletions (indels), or copy number variants (CNVs) in the genomes of many thousands of individuals. We have also discussed how the incidence and prevalence of some genetic disorders may differ among populations, and how discoveries have been made by selectively sampling individuals with specific phenotypes. Here we take a deeper look into the assumptions and limitations of how populations are defined in genomics research and medicine, considering in greater detail the underlying forces that shift or maintain allele frequencies over time. Identifying genetic susceptibilities to diseases is a key objective of medical genetics, with a critical role in clinical diagnosis and genetic counseling. Concepts and observations from population genetics inform our understanding of the genetic architecture of health and disease by observing how variants may differ in frequency, effect size, and phenotypic expression in human populations. When we refer to a population frequency, we consider a hypothetical gene pool as a collection of all the alleles at a particular locus for the entire population. The frequency of an allele is thus its proportion among all alleles at the same locus in a population. In this chapter, we describe a central organizing concept of population genetics, Hardy-Weinberg equilibrium (HWE), and its utility for genomics and clinical genetics in understanding the relationship between allele and genotype frequencies. Assumptions of the HardyWeinberg principle are considered in the context of factors that cause true or apparent deviation from equilibrium in real, as opposed to idealized, populations. Clinical genetics is primarily concerned with rare, often de novo mutations that cause genetic conditions and may severely impact function, with similar incidences across populations. Genetic variants that impact function and reproduction are rare in nearly all human populations because they are eliminated from the gene pool through natural selection. These variants are most often de novo – not inherited – so there is typically no link between genetic ancestral origins and the incidence of these pathogenic mutations. Most human genomic variation is shared among all groups, so there are exceedingly rare cases of clinically relevant variants that are found exclusively in one category or group of patients. The exception to this rule is when a defined ancestral group has experienced a bottleneck that reduces and then regrows the population with a subset of its original genetic variants. This increases the chance of maintaining a pathogenic variant in the population, as selective pressure is weaker than in a more genetically diverse population. Population genetics is the quantitative study of the distribution of genetic variation in populations, including trends in the frequencies of genes and genotypes over time. Differences in the frequencies of alleles that cause genetic disease are of particular interest to the medical geneticist and genetic counselor because they contribute to differences in disease risk among populations. Nevertheless, it is important to consider what we do not know and examine our baseline assumptions about genetic etiologies of health and disease. We present clinical examples to illustrate how the creation of genomic knowledge, as in all fields, depends on who is included in the underlying research, how their attributes are represented, and what categories are used to classify or stratify people into groups. The distributions of genetic variants and attributes in families, communities, and geographic regions are driven by a combination of social, environmental, and biological factors. Social scientists, anthropologists, and evolutionary biologists employ mathematical descriptions of shifting allele frequencies across geographic regions and time to reconstruct our evolutionary histories. Knowing about differences in allele frequencies across populations may be important for physicians to predict an increased likelihood of certain conditions in patients with specific ancestral origins linked to a pathogenic
chapter 10 Population Genetics for Genomic Medicine Alice B. Popejoy INTRODUCTION TO POPULATION GENETICS FOR GENOMIC MEDICINE We have explored in previous chapters the nature of genetic and genomic va...
Ch10 · Pt2 176 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE variant. However, it is important to keep in mind that most genetic variation that is used to describe ancestral origins is selectively neutral or has an unknown biological function. With today’s access to genome sequencing technology, considering all types of variation, we know that all humans share an average of 99% of our DNA sequences across the genome. The relatively small proportion of our genomes that do vary (polymorphism) do not differentiate us into discrete social categories; on the contrary, greater genomic variation has been observed within broad groupings (e.g., based on “race” or “ethnicity”) than between them. Nevertheless, humans tend to make meaning out of anecdotal observations, attributing physical or health-related differences we observe to underlying attributes of broad semantic categories that reinforce social stereotypes. This is one mechanism by which biological or scientific racism is enacted in genetics. Data subjects are often stratified into groups to make statistical comparisons, which (with sufficient sample sizes) may yield average differences that offer a biased confirmation of natural classes. Some have used these findings to argue for the inevitability of systemic inequities. This chapter touches on the history and impact of classifying humans into nominal groupings, making the case for genetics researchers and clinicians to understand the origins and assumptions underlying conceptual and functional models used to create and apply knowledge. Statistical methods have been used to identify variants that can differentiate many individuals into broad “continental ancestry” groupings based on allele frequencies observed across geographic regions. Such ancestry informative markers (AIMs) often lack functional relevance, but they have been used as a proxy for genomic background, which may inflate estimated genetic distances between groups. When such variants are associated with complex, multifactorial traits (with both genetic and environmental influences) such as obesity, diabetes, heart disease, and asthma, it is difficult to determine the true causal factors driving statistically significant associations, because social determinants of health influence the environment differently across sociocultural groups. A central challenge for human genomics research is thus teasing apart causal factors that drive phenotypic variation from confounders that are associated with patterns of both genotype and phenotype variation. Stratification is not inherently problematic as a statistical strategy to reduce the complexity of genomic and environmental backgrounds, but there are deep scientific and ethical flaws in approaches that collapse these nuanced dimensions of variation into static population descriptors such as race, ethnicity, and ancestry. European colonization, slavery, and eugenics have institutionalized conceptual frameworks tied to categories of difference, preserving hierarchies of power. These frameworks have persisted in colonized societies and science, with implications for analytic approaches still used today in biomedical research. Racial categories are constructed using poorly defined criteria that subdivide humankind using physical appearance (i.e., skin color, hair texture, and facial features) combined with characteristic identities that have their origins in the geographical, historical, cultural, religious, and linguistic backgrounds of the communities in which an individual was born and raised. Physical traits linked to racial and ethnic stereotypes may be influenced by genetics. However, genetic variation and diverse phenotypes exist across all populations, so attempts to equate continental origins with race are misguided. DNA alone cannot be used to assign someone to a social identity group, and it is primarily nongenetic, social and systemic factors that create the conditions for observed group-level differences in health outcomes. Concept of race and racism are important in discussions of social and health policy, tracking disparities in outcomes and social determinants of health. Longstanding systemic and structural factors create differences in health outcomes among racial groups, which can contribute to confounding and disparities in the quality of care, such as timely screening and test referrals. Thus, race may be used as a proxy for the effects of racism, but not as a proxy for genetic background or to calculate genetic disease risk. Building on themes from previous chapters, we paint a detailed picture of global genomic diversity that is largely shared among all human populations, and the factors that influence shifts in relative frequencies of genetic variants over time. We demonstrate the importance of African genomic diversity for all populations, highlighting the need for clinically relevant data on genetic heterogeneity, structural variation, and haplotype diversity across the globe. These data will inform our understanding of genetic etiologies of health and disease in everyone and must be included in public databases. Finally, we seek to clarify common misconceptions about human populations and diversity that could exacerbate health disparities in genomics and medicine. HUMAN ORIGINS AND FOUNDER EFFECTS OF SERIAL MIGRATION Our species (Homo sapiens) is comprised of more than 7 billion members, all of whom share a common ancestral lineage on the African continent dating back ~200,000 years ago. Fig. 10.1 illustrates the origin of modern humans in Southern Africa and suggests a series of migrations north through the continent beginning ~100,000 years ago. Subsequent dispersals of subsets of these migratory modern human ancestors led our species out of Africa to populate the other continents, such that we are thought to have reached South America ~15,000 years ago. Migration can change allele frequencies through the process of gene flow, defined as the slow diffusion
176 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE variant. However, it is important to keep in mind that most genetic variation that is used to describe ancestral origins is selectively neut...
Ch10 · Pt3 CHAPTER 10 — Population Genetics for Genomic Medicine 177 of variants, or alleles, across a barrier. Gene flow usually involves a large population and a gradual change in allele frequencies. The allele frequencies of migrant populations gradually merge into the gene pool of the population into which they have migrated, a process referred to as genetic admixture. The term migration is used here in the broad sense of crossing a reproductive barrier, which may be social or cultural (not necessarily geographic) and does not require physical movement from one region to another. During serial migration events a subset of the genomic variation from the original population travels with the migratory group into the new population. These events create founder effects, which are marked by a reduction in genomic diversity from the original, ancestral population to the new subpopulation. Founder events such as bottlenecks, the reduction in size and subsequent regrowth of a population whereby it loses a large portion of its original diversity, create so-called founder populations. In populations that have experienced a bottleneck (e.g., European ancestry groups), we observe greater genetic homogeneity (lower heterozygosity) than would be expected, based on the overall estimated population size. Founder populations thus have lower effective population sizes (Ne) relative to their census populations, and various estimators for genetic diversity use this concept to model human population history. Founder populations with small Ne are relevant for clinical genetics, because individuals with these ancestries are generally at a greater risk of accumulating genetic diseases that are rare in other populations. Dominant deleterious mutations can rise to high frequency more easily when there are fewer potential reproductive partners, as this reduces competition in the form of alternative alleles. For example, the high incidence of Huntington disease (Case 24) among inhabitants around Lake Maracaibo, Venezuela, resulted from genetic isolation following the introduction of a (likely European) genetic variant that causes Huntington disease. Similarly, French-Canadians in the Saguenay-LacSaint-Jean region of Quebec are at greater risk for the autosomal recessive condition hereditary type I tyrosinemia. If untreated, this condition causes hepatic failure and renal tubular dysfunction due to deficiency of fumarylacetoacetase, an enzyme in the degradative pathway of tyrosine. The disease incidence was estimated (in 1994) at 1 in 1846, with the pathogenic allele carrier frequency estimated at 1 in 22. Nearly all pathogenic alleles observed in Saguenay-Lac-Saint-Jean patients are due to the same variant inherited through a shared lineage of recent common ancestor(s). This observation is consistent with the history of this population, which is a known founder population or genetic isolate, meaning there was reduced genetic variation among individuals who originally founded the population, followed by population growth and a paucity of new alleles flowing into it. As a result of serial founder events throughout our history, nearly all human genomic variation in modern populations is only a subset of the variation that exists in Africa, the ancestral homeland of our species, and the current home of modern human populations with 50-60Kya 15Kya 60-100Kya 45Kya 45Kya Founder effect Source of founder effect Migration path 35-40Kya Figure 10.1 Ancient dispersal patterns of modern humans during the past 100,000 years. This map highlights demic events that began with a source population in southern Africa 60 to 100 kya and concluded with the settlement of South America approximately 12 to 14 kya. Wide arrows indicate major founder events during the demographic expansion into different continental regions. Colored arcs indicate the putative source for each of these founder events. Thin arrows indicate potential migration paths. Many additional migrations occurred during the Holocene. From Henn BM, Cavalli-Sforza LL, Feldman MW. The great human expansion. Proceedings of the National Academy of Sciences of the United States of America. 109:17758–64. PMID 23077256. https://doi.org/10.1007/S12045-019-0830-4
CHAPTER 10 — Population Genetics for Genomic Medicine 177 of variants, or alleles, across a barrier. Gene flow usually involves a large population and a gradual change in allele frequencies. The allel...
Ch10 · Pt4 178 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE more genetic diversity compared to anywhere else in the world. African populations and individuals with recent African ancestries have the largest effective population sizes, with genomic evidence pointing to Indigenous South Africans having the most diverse and anciently diverged genomes. It is a common error to think of these ancestral genomes as an artifact of the past; they have in fact continued to accumulate variation and to experience effects of natural selection over time. In addition to standing ancestral variation, new mutations arise – albeit slowly – and other sources of variation have been introduced in different ancestral lineages. Evidence (mostly from studies of European ancestry populations) suggests Homo sapiens encountered other early hominid species such as Neanderthals and Denisovans – whose genetic fragments were preserved in cold, dry climates. Segments of these archaic hominids’ chromosomes appear to have flowed into human genomes through introgression approximately ~50 to 90k years ago. By comparing human genomes to DNA extracted from excavated bone fragments, ancient DNA researchers estimate that 1% to 4% of the DNA of modern humans may be derived from these species. While there is little evidence of archaic hominid introgression into human populations in regions of the world marked by hot and humid climates, (which are not conducive to the preservation of DNA in ancient remains), it is likely that such introgression events occurred wherever other hominid species met and mingled with ancient Homo sapiens. DEFINING HUMAN POPULATIONS TO CHARACTERIZE DIVERSITY There are many ways to define a population – whether criteria are based on geography, societal and cultural factors, geopolitical or national borders, or any other features that may be considered characteristic of individuals within a group. Whether an individual or group is designated as belonging to a “population” depends on the classification framework, and how it is implemented in practice. Many designated “populations” have nothing to do with genetics – such as the census population of a large, urban metropolis. The specific criteria that are used to define a population often flow from certain research or clinical objectives and may also be influenced by social, cultural, and political factors that privilege certain approaches. Who has the authority to determine such classification criteria also plays an important role in defining human populations. Decisions made by researchers and clinicians about how to classify participants or patients into groups strongly influence population-level estimates such as allele frequencies and disease prevalence. Methods used to detect and interpret combinations and mixtures of these alleles, and to investigate how they impact our health, play an important role in what is generally accepted about average group-level differences. For example, population-level statistics such as point estimates that serve to represent averages across an entire group (e.g., allele frequencies) are influenced by who is included and excluded from the group, and their level of genetic relatedness. The magnitude of genetic similarities and differences among individuals defined by some category or population descriptor thus impacts the accuracy of any given point estimate for that category. It also influences the precision with which such a point estimate can be used to make predictions about individuals in the group. For example, population categories such as “Asian” and “African” are so broad that they describe more than half of the world population and land mass. The vast genetic and environmental heterogeneity within such groups call into question the reliability of point estimates based on such groupings. In contrast, populations with less total variability are more likely to yield group-level estimates that reflect a greater number of individuals in the category. For example, populations that have experienced bottlenecks have less overall diversity in the underlying gene pool and population mean and median values may be more informative. This is important because it means that research methods or approaches to estimating disease risk based on point estimates have differential validity and utility across population groups, benefiting those with less variability (i.e., European ancestries). Many clinical case studies present population allele frequencies as point estimates corresponding to different race, ethnicity, or ancestry categories. Unfortunately, these terms are conceptualized and used in vastly different ways among clinical genetics professionals and researchers, so summary statistics that are based on such poorly or broadly defined groupings may be unreliable and inconsistent, depending on the underlying characteristics of the groups. Genetic ancestry and ethnicity represent different kinds of information, so direct comparisons cannot be made between them (despite misguided attempts to ‘valide’ self-identified race or ethnicity using estimated proportions of DNA shared with ‘known’ population reference datasets). Most people are not limited to one ancestry or ethnic identity, and these identity concepts can change over time. As such, ontological frameworks and methodologies in population genetics that force a single-group assignment are conceptually inconsistent, insufficient for characterizing the spectrum of human cultural and genomic diversity, and are becoming increasingly obsolete. For example, while “Ashkenazi Jewish” represents a cultural or ethnic designation without specific geographic constraints, its related definition of having “ancestry from Africa, the Middle East, or the Mediterranean” is geographically diverse and includes
178 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE more genetic diversity compared to anywhere else in the world. African populations and individuals with recent African ancestries have the l...
Ch10 · Pt5 CHAPTER 10 — Population Genetics for Genomic Medicine 179 a vast array of social and cultural identities. For many years, carrier screening for autosomal recessive diseases such as for Tay-Sachs disease used a risk-based strategy that relied on self-described ethnicity; this approach is now known to introduce inaccuracy into the screening process. Thus, recent recommendations from the American College of Medical Genetics and Genomics (ACMG) note that carrier screening paradigms should be agnostic to race, ethnicity, and ancestry. Carrier screening, referrals for genetic testing, and reporting risk predictions to patients or other care providers have relied on patient self-reported information about race, ethnicity, or ancestry. Certain populations may be at higher risk for conditions that are not routinely screened for, but only individuals who have self-identified with (or for whom a provider has determined their membership in) these populations would have access to those screenings. In some circumstances, physicians may not think to refer patients for genetic testing unless they fit the stereotype of the population known to be at high risk, and insurance companies may not cover genetic testing for certain conditions unless a patient has disclosed information about their background that is consistent with population-specific elevated risk of disease. Thus, whether and how patients are classified into population groups or categories is of great importance here. Expanded carrier screening has received much attention and has the potential to alleviate some of the pitfalls of relying on proxy measures and poorly defined population groups to guide clinical decision-making. However, once the results of genetic tests are received, there are additional uses for population-level information about patients that inform the interpretation and curation of findings. A survey of clinical genetics professionals and researchers, conducted mainly in the United States, revealed that patient self-reported race and ethnicity information may be used to make decisions about ordering tests, interpreting results, and reporting findings to patients. At least 18% of respondents reported that this information may have been entered into a patient’s medical record without asking them directly. Fewer than 5% of clinical genetics professionals who responded to the survey reported that they conduct ancestry analyses or have access to ancestry estimates based on patients’ DNA for the purpose of clinical variant interpretation and gene curation. This means that social categories are used as a proxy for genomic background in clinical genetics. Cultural and political contexts shape the way we collect and use information about patients’ social/cultural identities and ancestral backgrounds – and asking for a person’s race or ethnicity is illegal in some parts of the world. It is important for clinical genetics professionals to know about the risks of relying too heavily on this information to make predictions or decisions in the clinical setting. Box 10.1 offers descriptions of “race,” “ethnicity,” and “ancestry” that have been useful for genomics researchers. BOX 10.1 DESCRIPTIONS OF RACE, ETHNICITY, AND ANCESTRY “Race,” “ethnicity,” and “ancestry” are often used interchangeably, yet they have no universal definitions. We provide brief descriptions of our usage below. For extensive discussion in context of genomics, including recommendations from professional organizations, see Banda et al. (2015); Mersha and Abebe (2015); Race, Ethnicity, and Genetics Working Group (2005). Race: A culturally and politically charge term, for which definitions and meaning are context-specific. Race is related to individual and/or group identity and is often linked to stereotypes of visible physical attributes such as skin and hair pigmentation. The concept of race is tightly linked to social power dynamics and has historically been used to justify hierarchies of power, discrimination, and oppression in an unequal society. Social and cultural conditions may differ among racial groups, on average, and these differences may lead to environmental effects such as chronic stress and unequal access to goods and services, including healthcare and nutrition. These inequities can affect environmental risk for complex diseases and/or potentially interact with genetics to affect risk. Ethnicity: Describes people as belonging to cultural groups, usually on the basis of shared language, traditions, foods, etc. Ethnicity has often been used interchangeably with race and is similarly ambiguous. To the extent that traits are affected by social and environmental differences, ethnicity has previously served as a proxy for health and disease risk at the population level as a result of social, cultural and community effects described above. There is no universal agreement on a system of “ethnic” groupings worldwide. Some ethnic groups may share genetic factors due to similar ancestral origins; other groups may be more social and cultural in nature. Ancestry: Meaning varies by context. Here, we use the term to denote genetic ancestry, a description of the population(s) from which an individual’s recent biological ancestors originated, as reflected in the DNA inherited from those ancestors. Genetic ancestry can be estimated via comparison of participant’s genotypes to global reference populations, so incomplete availability of these references can create biased estimates. We note that different methods of calculating genetic ancestry can yield different results. Thus, discreet labeling of ancestral populations oversimplifies the complexity of human genetic variation and demography. Nevertheless, accounting for systemic differences in allele frequencies and linkage disequilibrium is necessary for genetic analyses. In this paper, diversity in genomics is described primarily in terms of ancestry. From Peterson RE, Kuchenbaecker K, Walters RK, et al: Genome-wide association studies in ancestrally diverse populations: Opportunities, methods, pitfalls, and recommendations, Cell 179(3):589-603, 2019. https://doi.org/10.1016/j.cell.2019.08.051
CHAPTER 10 — Population Genetics for Genomic Medicine 179 a vast array of social and cultural identities. For many years, carrier screening for autosomal recessive diseases such as for Tay-Sachs disea...
Ch10 · Pt6 180 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE History and Influence of Eugenics on Population Genetics Countless atrocities have been committed in the name of perceived differences among human beings, from oppression, discrimination, and displacement to slavery and genocide. In the United States, Jim Crow laws were enacted upon the abolition of slavery and persisted overtly through the 1960s, which forced segregation of “Black” and “white” people to preserve exploitative power dynamics and justify economic and social injustice. The ideological underpinning of segregating hospitals and clinical care was that white and Black people needed different medical care due to perceived differences in biology. However, there are no unique, fundamental genetic differences between groups of people who identify as “white” or “Black,” and the persistence of this racial binary in genomics research has contributed to misconceptions and harm. Many in human genomics may not realize how the field’s history is rooted in the American Eugenics Movement, which provided a false sense of scientific legitimacy to Nazi propaganda during WWII. Indeed, the Annals of Human Genetics was originally called the Annals of Eugenics, founded in 1925 by an English eugenics thought leader, Francis Galton. Mentees of Galton’s became the journal’s editors, including Karl Pearson and R. A. Fisher, who developed statistical methods we still use today. They include the chi-square test, p-values, analysis of variance (ANOVA), and principal component analysis (PCA). Another of Galton’s mentees, Charles Davenport, used taxonomic frameworks classify humans, the scientific endeavor of biological racism. The idea that racial groups were genetically distinct predates eugenics, back to early European slave traders who devized categories that dehumanized people with darker skin to justify their enslavement and exploitation. While it may feel uncomfortable to engage the history of these ideologies researchers and physicians must understand how they persist in our scientific and clinical methodologies so that we can heal the wounds of the past, moving forward with greater care and precision. Aristotle’s Scala Naturae laid the foundation for Carl Linnaeus’s 1737 Systema Naturae, a taxonomy of humans subdivided into four continental groups based on skin color: “whitish European,” “reddish American,” “tawny Asian,” and “blackish African”; a fifth category described “wild and monstrous humans, unknown groups, and more or less abnormal people.” Such taxonomic classes strike a shuddering resemblance to continent- or race-based categories that persist in human genomics. It is everyone’s responsibility to reflect on this to ensure our studies are robust to the conceptual and mathematical influence of constructed social hierarchies rooted in biological racism. It is a powerful and pervasive myth that genetics are primarily responsible for differences we observe across geographic regions, cultural contexts, and social or political identities. Eugenic scientists invoked natural class theory, but also relied on genetic essentialism–the notion that genetics are necessary and sufficient for all phenotypes. Early in the field’s history, statistical geneticist R. A. Fisher was commissioned by Leonard Darwin to develop mathematical models that could provide a biological explanation for trait differences between groups of individuals. It is widely accepted today that the social and culture-bound traits that eugenicists focused on (e.g., imbecility and promiscuity) are not caused by genetics. However, concepts and statistical methods that Fisher developed to explain how such traits occurred more frequently within families are still used today. Twin studies of inherited contributions to disease – using Fisher’s Heritability – were conducted on children imprisoned in concentration camps during WWII. Heritability estimates from twin studies have motivated GWAS of social phenomena (e.g., educational attainment) despite their problematic assumptions and historical abuses. Heritability estimates have curiously remained central to certain branches of human genetics and social science. Causal relationships between exposures and outcomes can rarely be proven, because there is no way to disentangle the impact of shared family or cultural environment (twins reared apart are often raised by relatives or other community members) and the impact of shared genetic variation within those extended family communities. In particular, for complex, or multifactorial, traits (in which the environment plays a role), it is important to acknowledge the limitations of this heuristic. Heritability estimates cannot truly distinguish between genetic and nongenetic effects, and there must be a proposed mechanism justify claims that genetics could influence purely social outcomes. There is no scientific basis for using broad social categories or sweeping statistical assumptions to infer anyone’s precise genomic variation, so we need to develop better methods and conceptual frameworks to bolster our interpretation of genomics. Researchers, clinicians, professional societies, and funding agencies are dealing with problems introduced by the legacy of using race, ethnicity, and ancestry in biomedical research and medicine. Eliminating race-based clinical algorithms may reduce health disparities, but in the absence of robust data from diverse populations, the standard of care is biased toward white people of European descent. Collecting this information is illegal in some countries; thus rendering invisible the influence of social stigma and discrimination on health outcomes. We must be able to track health and healthcare disparities, while avoiding unscientific and unethical applications of broad social categories. Complex Disease and Confounding It is important to recognize the role of confounding in the causal pathway from exposures to health outcomes.
180 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE History and Influence of Eugenics on Population Genetics Countless atrocities have been committed in the name of perceived differences among...
Ch10 · Pt7 CHAPTER 10 — Population Genetics for Genomic Medicine 181 Confounding can create spurious associations (e.g., between continental ancestry and disease risk), even if there is no genetic basis for the condition. Statistical associations observed between outcomes and exposures are confounded when there is another, unmeasured or hidden factor that is both causally linked to the outcome of interest and associated with the exposure. This generates statistical correlations between measured exposure(s) and outcome(s) despite the absence of a causal relationship. Fig. 10.2 (B) illustrates confounding by social determinants of health (SDOH) and group categories (race, ethnicity, and ancestry) in a causal pathway for complex traits. The true cause(s) of these traits could be nongenetic, environmental factors at a higher prevalence in the population, and/or genetic factors that are difficult to identify due to small effect sizes across multiple loci. Linkage disequilibrium (LD), or the non-independence of allele frequencies along segments of chromosomes that are inherited together, creates additional statistical challenges. In an association study, the true causal variant cannot be distinguished from other variants that are in LD with it. If there is confounding by SDOH associated with genetic ancestry and an outcome of interest, ancestral patterns of LD may create group-specific genetic signals that are not causal. One example of uncontrolled confounding leading to harmful and false conceptions is the belief that genetics primarily drives racial or ethnic differences in health outcomes, or other complex traits with environmental components. Because some racial and ethnic groupings are associated with genetic ancestry at the global, or continental, level (i.e., African, Asian, European, Latin American, etc.), genetic variation that happens to be more prevalent in some groups than others (by chance) may be mistaken for causal genetic factors contributing to intergroup differences in traits. Instead, these are caused by environmental factors that also differ among groups. In these cases, nongenetic factors confound the relationship between ancestry-associated traits or outcomes and genetic variation shared in ancestry-associated cultural groups. Identity by Descent Versus Allele Sharing in Unrelated Individuals Throughout recorded history and still today, humans have migrated around the world and exchanged DNA with one another such that our ancestral roots live in our genomes as mixtures of chromosomal segments (haplotypes) that have been inherited through our maternal and paternal lineages. When these lineages are related, segments of maternal and paternal chromosomes in an individual will be shared in the manner of being identical by descent (IBD), meaning they descend from a recent common ancestor. In contrast, genomic loci that have shared variation in a large population with multiple ancestries, or between ancestral groups, are considered identical by state (IBS) and do not necessarily share a recent common ancestral origin. IBD sharing between maternal and paternal chromosomes is a result of consanguinity, or reproduction between genetically related individuals. This, like founder effects, leads to an increase in the frequency of autosomal recessive traits because carriers of the same recessive allele are more likely to meet and reproduce. The kinds of recessive disorders seen in the offspring of related parents may be very rare and unusual in the general population. This is because consanguineous mating allows an uncommon allele inherited from a heterozygous common ancestor to become homozygous. A B Mendelian trait: Autosomal dominant, monogenic, rare Pathogenic variant: Rare, de novo, highly penetrant Social determinant(s) of health (SDOH) Complex trait: Multifactorial, polygenic, highly prevalent Candidate variant(s): Common, inherited, variable penetrance Correlation observed between ancestry and complex trait variation as a result of confounding by SDOH and social categories linked to ancestry. No correlation between ancestry and mendelian trait variation because causal variant(s) are very rare, and often de novo. PATHWAY n A = 10,000 n B = 10,000 n A = 100 n B = 100 Ancestry A Ancestry A Ancestry B Ancestry B CONFOUNDING Complex trait variation Social category with ancestries A and B (racial or ethnic group) CAUSAL Figure 10.2 (A) Mendelian traits are typically rare, due to deleterious effects of de novo variants that are unlikely to be inherited by subsequent generations so individuals with different ancestries have a similar risk profile. (B) Complex traits are common, caused by combined genetic, social, and other environmental factors. Ancestry may appear to influence disease risk if it is associated with the trait of interest and social determinants of health.
CHAPTER 10 — Population Genetics for Genomic Medicine 181 Confounding can create spurious associations (e.g., between continental ancestry and disease risk), even if there is no genetic basis for the...
Ch10 · Pt8 182 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE IBD sharing is more prevalent in genetic isolates, small populations derived from a limited number of common ancestors who tended to mate only among themselves. Reproduction between two apparently “unrelated” individuals in a genetic isolate may have the same risk for certain recessive conditions as that observed in consanguineous reproduction because the individuals are both carriers via inheritance from common ancestors within the isolate. Fig. 10.3A illustrates a simplified path of alleles shared IBD, inherited from parents who are related or have recent, closely shared ancestry (cryptic relatedness); in contrast, Fig. 10.3B shows alleles IBS that are de novo mutations and/or inherited independently. The concept of a coalescent describes the joining together of distinct lineages or alleles backward in time (t) to a most recent common ancestor (MRCA). Coalescent theory is used to estimate t for two copies of an allele or haplotype (collection of linked alleles) that are shared IBD (Fig. 10.3A). In contrast, identical alleles are IBS (Fig. 10.3B) when they arise through independent processes such as mutation or inheritance through different ancestral lineages. The relative number of alleles (or average lengths of stretches of chromosomes) shared IBD among individuals in a population is proportional to the effective population size. Smaller Ne leads to greater IBD sharing, on average, because there are fewer haplotypes circulating in the population. This leads to a greater prior probability that two recessive deleterious alleles will be inherited by chance through closely or distantly related ancestral lineages. Assortative mating describes a phenomenon in which certain groups of individuals have – for a variety of historical, cultural, or religious reasons – remained relatively genetically separate during modern times. When mate selection in a population is restricted for any reason to members of a particular group, and that group happens to have a variant with a higher frequency than the population as a whole, the result will be an apparent excess of homozygotes in the overall population beyond what one would predict. A clinically important aspect of assortative mating is the tendency to choose partners with similar traits, such as congenital deafness or blindness. In such cases, the genotypes of two reproductive partners at loci influencing the trait are not predicted from population allele frequencies. For example, consider achondroplasia (Case 2), an autosomal dominant form of skeletal dysplasia with a population incidence of 1 per 15,000 to 1 per 40,000 live births. Offspring homozygous for the achondroplasia variant have a severe, lethal form of skeletal dysplasia that is almost never seen unless both parents have achondroplasia and are thus heterozygous for the variant. This would be highly unlikely to occur by chance, except for assortative mating among those with achondroplasia. When reproductive partners have autosomal recessive disorders caused by the same pathogenic variant or by allelic variants in the same gene, all their offspring will also have the disease. Even if there is locus heterogeneity with assortative mating, the chance that two individuals carry pathogenic variants in the same locus is increased over what it would be under true random mating, and therefore likelihood of the trait in their offspring is also increased. Genetic Ancestry and Population Structure The concept of genetic ancestry may seem more scientifically valid and concrete than self-reported measures like race and ethnicity because it is based on genetic information; however, it also lacks a firm “ground truth.” Genetic ancestry is a dynamic and relative measure; estimates can change over time depending on the reference data and methods used, all of which have inherent assumptions and limitations. Often reported as percentages of the genome that can be traced back to ancestral populations, ancestry estimates are usually based on the MRCA A B t generations cryptic relatedness unrelated lineages mutation inheritance mutation IBD IBS - recent ancestry - consanguinity Figure 10.3 (A) Coalescence of an allele shared identical by descent (IBD) between two individuals traced back to their most recent common ancestor (MRCA) t generations ago; and (B) de novo or recently inherited alleles identical by state (IBS) in unrelated individuals. Blue diamonds are individuals and the edges connecting them indicate their relations in a pedigree; orange circles and arrows represent pathogenic variants, their origins traced forward being inheritance and backward being coalescence in lineages.
182 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE IBD sharing is more prevalent in genetic isolates, small populations derived from a limited number of common ancestors who tended to mate on...
Ch10 · Pt9 CHAPTER 10 — Population Genetics for Genomic Medicine 183 average proportion of an individual’s DNA that most closely matches a given reference dataset with assigned population labels (compared to all other reference data available for the analysis). This means that the more robust and geographically specific reference datasets are, the higher the resolution achieved for reported proportions of population-specific ancestry. For example, most direct-to-consumer (DTC) genetic ancestry companies can predict specific geographic origins for individuals’ European ancestry components (e.g., a small village in Northern Ireland) but often report ancestry components at the continental level for regions that are underrepresented among reference datasets, such as “Sub-Saharan Africa.” Representation in genomic reference datasets disproportionately excludes the most genetically diverse ancestries while including mostly Europeans. Since genetic ancestry is estimated using reference data, individuals with ancestries from regions of the world that have yet to be broadly included in genomics research may not receive accurate, detailed, or informative ancestry proportion results. Similarly, imputation (the process of filling in of alleles that are missing from genotype data using some reference dataset), is less accurate in groups with greater genetic diversity that is missing from reference resources. Alleles that differ in their frequencies among preconstructed ancestry groupings are referred to as ancestry informative markers (AIMs). AIMs have been identified to differentiate among broad geographical groupings (e.g., African, East Asian, South Asian, European, Middle Eastern, Native American, and Pacific Islanders). Such markers have been used for charting human migration patterns, for documenting historical admixture between or among populations, and for determining the degree of genetic diversity among ancestral population groups. Studies of hundreds of thousands of AIMs from across the genome have been used to distinguish and determine the genome-wide relationships among many different populations. Since AIMs are selected to maximize differences between predefined groups, they should not be considered representative of genome-wide variation among those groups. Polymorphisms that happen to occur at higher frequencies in certain groups of individuals may be identified as AIMs by chance, even if the grouping scheme is not otherwise biologically meaningful. Predefined groups are often based on sociocultural categories or their respective broad, continental groupings; they influence the loci that are selected, then those AIMS may be treated as a proxy for genomic background. This has downstream implications for research and quality control measures that rely on AIMs (e.g., to identify population-specific reference panels, impute missing genomic information, or test the accuracy of various analytic methods). In 2008, population geneticist John Novembre and colleagues famously published results of a principal component analysis (PCA) that appear to reconstruct the geography of Europe using genotype data sampled from countries across the continent (Fig. 10.4). Subsequently, many genomics researchers have used PCA and other statistical clustering methods like admixture analysis that reduce the complexity of data to visualize population structure. While these approaches may provide some insight, they also have limitations and may be misleading. Novembre and colleagues have published concerns about the potential for confusion about global genetic population structure because of these methods. For example, geographic clusters that appear distinct in the PCA plot of Europe are genetically very similar, but the figure may mislead people into thinking they have large overall genetic differences. In the case of Europe, clines or gradients in allele frequencies are roughly aligned with latitude and longitude, as humans migrated EastWest and South-North across the continent. This pattern of allele frequencies mirroring geography is unique to Europe, and the structure shown only emerges after >100k loci are included in the analysis because allele frequency differences are so small. Genetic variation among populations may be falsely perceived as divided among regional groups, though few alleles are restricted to just one region of the world. Data from the US Population Architecture using Genomics and Epidemiology (PAGE) study were used to visualize population structure with PCA (Fig. 10.5). Each colored dot in the plot of principal components (PCs) 1 and 2 represents an individual research participant, and its color corresponds to the self-reported race or ethnicity category selected by the participant. The dots are positioned relative to one another according to similarities and differences in genotypes across the genome. We can see from the spread of these individuals across PCs that there is a complex spectrum of shared variation that cannot be adequately represented by categorical data structures. This illustrates why sociocultural categories used in the US Census cannot be considered genetically differentiated; there is a continuous distribution of genomic variation within and among groups such that most variants are shared and their frequency distributions overlap. Misclassification results from individuals being assigned to the wrong analytic group in a study (e.g., case being classified incorrectly as a control, or vice versa). In clinical algorithms that rely on racial and ethnic classification, misclassification could be considered the incorrect attribution of a category-based mean value to an individual patient. Furthermore, classification error is almost certain due to genetic heterogeneity within racial and ethnic groups that violate baseline assumptions and motivations for applying population- or group-level adjustments. If genetic ancestry were used instead of self-reported measures of racial or ethnic identity, the preselected nature of
CHAPTER 10 — Population Genetics for Genomic Medicine 183 average proportion of an individual’s DNA that most closely matches a given reference dataset with assigned population labels (compared to all...
Ch10 · Pt10 184 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE AIMs and then using those to classify individuals into discrete groupings may create the exact same problems as race and ethnicity categories. That is, ancestry categories that are semantic variations of racial and ethnic groupings do no better at describing genomic variation because human genetic diversity is a narrow spectrum of [mostly] shared alleles at gradually changing, relative frequencies. Humans from different regions of the world have genetic similarities and differences that may have nothing to do with geography. Whereas there may be geographic trends in aggregate, an allele frequency cannot be used to determine an individual’s genotype. Even Duffy blood group alleles, which are classically known for differences in frequencies by geography, can be seen on every continent, with clines or gradients showing clear variation within those regions. Earlier versions of this very textbook have described blood group variation according to geographic, racial, and ethnic categories (which change over time and across cultural contexts) but today, we have enough genetic data to establish that neither is genomic variation restricted to such categories nor is it easily characterized by any categorical framework. To illustrate this point, Fig. 10.6 showcases continuous global frequencies of the three common Duffy blood group alleles (FY*A, FY*B, and FY*BES), which have historically been used as the canonical example of geographic differentiation. We highlight here that not all parts of Africa have a high frequency of the FY*BES allele that confers protection against malarial infection via the Plasmodium parasites, and some other parts of the world (not in Africa) have elevated frequencies as well. In geographic locations where Plasmodium species causing malaria in humans are endemic, the protective allele is highly prevalent. All three common Duffy alleles are present at a range of frequencies across the Americas, and none are restricted to a single continent or large geographic region that corresponds to continental ancestry groupings. There are many different factors that can explain how differences in disease incidence, prevalence, and allele frequencies arise among biogeographic populations. In cases where genetic variants are contributing Figure 10.4 A statistical summary of genetic data from 1387 Europeans based on principal component axis one (PC1) and axis two (PC2). Small colored labels represent individuals and large colored points represent median PC1 and PC2 values for each country. The inset map provides a key to the labels. The PC axes are rotated to emphasize the similarity to the geographic map of Europe. AL, Albania; AT, Austria; BA, Bosnia-Herzegovina; BE, Belgium; BG, Bulgaria; CH, Switzerland; CY, Cyprus; CZ, Czech Republic; DE, Germany; DK, Denmark; ES, Spain; FI, Finland; FR, France; GB, Great Britain; GR, Greece; HR, Croatia; HU, Hungary; IE, Ireland; IT, Italy; KS, Kosovo; LV, Latvia; MK, Macedonia; NO, Norway; NL, Netherlands; PL, Poland; PT, Portugal; RO, Romania; RS, Serbia and Montenegro; RU, Russia; Sct, Scotland; SE, Sweden; SI, Slovenia; SK, Slovakia; TR, Turkey; UA, Ukraine; YG, Yugoslavia. From Novembre J, Johnson T, Bryc K. et al. Genes mirror geography within Europe. Nature 456:98–101, 2008. https://doi.org/10.1038/nature 07331
184 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE AIMs and then using those to classify individuals into discrete groupings may create the exact same problems as race and ethnicity categorie...
Ch10 · Pt11 CHAPTER 10 — Population Genetics for Genomic Medicine 185 to disease etiology, it is possible that inheritance of an ancestral haplotype containing a pathogenic variant is more likely, given reported or estimated ancestry of an individual patient. However, disease-causing alleles that reduce the fitness of an individual tend to be rare in populations that are sufficiently large (such that other haplotypes are frequent enough to outcompete one that is pathogenic). In smaller populations that have been genetically isolated from others, or in situations where environmental conditions enhance the fitness of carriers of pathogenic variants and create a heterozygote advantage, it is possible to see disease-causing alleles at frequencies higher than would be expected given disease prevalence. Other factors include genetic drift, which applies to benign variants that rise to high frequency due to physical proximity to fitness-enhancing alleles, without having direct impact on the phenotype. If the overall population is sufficiently large, the frequency of an allele in a small subset of the population will not change the total population allele frequency. However, the larger the subset of carriers relative to the total population, the greater the chance that it will alter the population allele frequency. Therefore, classification of individuals into population categories important because the relationship between the numerator (number of observations) and the denominator (total population under investigation) determines the allele frequency. In practice, estimates of allele frequencies are used in combination with disease prevalence and incidence to determine genotype frequencies, given their modes of inheritance. Figure 10.5 Inclusion of multiethnic samples enables discovery and replication in GWAS. The population substructure that is present in the multiethnic sample of PAGE (n = 49,839) reveals complex patterns preventing meaningful stratification. PC1 and PC2 show major patterns of variation, stratified by self-identified race/ethnicity. Individuals denoted by orange self-identified as “Other.” From Wojcik GL, Graff M, Nishimura KK, et al. Genetic analyses of diverse populations improves discovery for complex traits, Nature 570:514–518, 2019. https://doi.org/10.1038/s 41586-019-1310-4
CHAPTER 10 — Population Genetics for Genomic Medicine 185 to disease etiology, it is possible that inheritance of an ancestral haplotype containing a pathogenic variant is more likely, given reported...
Ch10 · Pt12 186 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Most sections in this chapter have thus far dealt with the complexity of defining human populations, characterizing ancestry and genetic population structure, and accounting for nongenetic factors associated with systemic inequities that may create erroneous correlations among group-level phenotypic differences, ancestry estimates, and sociocultural categories. While it is important to recognize limitations of point estimates and other discrete measures to characterize continuous variables (e.g., genomic variant distributions in human populations), there is utility in calculating allele and genotype frequencies to inform genomic research and medicine. In the next section, we describe how to calculate these frequencies and offer clinically relevant examples. ALLELE AND GENOTYPE FREQUENCIES For autosomal loci, the size of the gene pool at one locus is twice the number of individuals (2 N) in the population because each autosomal genotype consists of two alleles. Consider a population that is structured by recent ancestry or migration from another population, such that 10% of the total population contains a group in which the minor allele frequency (MAF) for a biallelic, autosomal recessive disease is 5% (q = 0.05). Since the trait has two alleles whose frequencies sum to 1 in the population (p + q = 1), we can infer the other allele’s frequency to be p = 0.95 in this group by subtracting the frequency of q from 1. In the remaining 90% of the total population, the frequency of this pathogenic variant is so small that it is not observed and presumed to be absent such that q≈0andp≈1. Let us continue with the Duffy blood group example to illustrate the relationship between allele and genotype frequencies in populations. Consider the gene ACKR1 (atypical chemokine receptor 1) on chromosome 1 (1q23.2), which encodes the major subunit of the Duffy blood group system and serves as an entry point for Plasmodium vivax, an important parasite causing malaria. Hundreds of variants have been observed in this gene with varying functional consequences. We will focus on a single nucleotide variant (SNV) in the promoter region of ACKR1 (rs 2814778 in Ensembl), which confers protection against malarial infection and exhibits allele frequency differences across global populations. We draw on data from the 1000 Genomes (1KG) Project to calculate allele frequencies from observed genotype FY*A frequency FY*B frequency FY*BES frequency FY*A IQR FY*B IQR 0–5% FY*BES IQR 50–70% 0–5% 5–10% 10–20% 20–30% 50–60% 0–5% 5–10% 10–20% 20–30% 30–50% 50–70% 30–50% 70–80% 80–90% 90–95% 95–100% 70–85% 50–70% 30–50% 30–50% 20–30% 10–20% 5–10% 0–5% 50–70% 70–80% 80–90% 90–95% 95–100% 0–5% 5–10% 10–20% 20–30% 30–50% 50–70% 70–90% 20–30% 10–20% 5–10% 0–5% 5–10% 10–20% 20–30% 30–50% B C A D E F Figure 10.6 Global Duffy blood group allele frequencies and uncertainty maps. A, B, and C correspond to FY*A, FY*B, and FY*BES allele frequency maps, respectively (median values of the prediction posterior distributions); D–F show the respective interquartile ranges (IQR) of each allele frequency map (25–75% interval). Predictions are made on a 5 × 5 km grid in Africa and a 10 × 10 km grid elsewhere. From Howes R, Patil A, Piel F. et al. The global distribution of the Duffy blood group. Nat Commun 2:266, 2011. https://doi.org/10.1038/ ncomms 1265
186 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Most sections in this chapter have thus far dealt with the complexity of defining human populations, characterizing ancestry and genetic pop...
Ch10 · Pt13 CHAPTER 10 — Population Genetics for Genomic Medicine 187 frequencies. Table 10.1 reports the genotype frequencies for homozygous (T|T and C|C) and heterozygous (C|T) individuals. Because each homozygous individual has two copies of the same allele, and heterozygous individuals have one copy of each allele, the frequency of each allele is twice the number of individuals in the population homozygous for the allele, plus the number of heterozygotes, divided by the total number of alleles in the population or 2N:
T: 2 1793 1 88 2504 2 0.734 × () + × () × = C: 2 623 1 88 2504 2 0.266 × () + × () × = Rather than calculating the frequency of each allele independently, the calculated frequency of one allele can s...
Ch10 · Pt14 188 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Parent Generation (P1,2) Possible Mating Pairs P1 Genotype: AA P2 Genotype: AA A AA AA AA Aa Aa Aa AA Aa Aa Aa AA AA AA AA AA Aa Aa Aa Aa aa aa aa Aa Aa Aa Aa Aa aa aa aa Aa aa aa aa Aa Aa A A A a a a A A a a a P2 Genotype: Aa P2 Genotype: aa P1 Genotype: Aa P1 Genotype: aa F1 F1 F1 Predicted offspring (F1) genotypes, given genotypes of possible parent mating pairs (P1,2) in a population F1 F1 F1 F1 F1 F1 The table above uses Punnett squares to illustrate all possible genotypes present in a generation of offspring (F1) resulting from all possible combinations of genotypes in the parent generation (P), given two ACB ASW BEB CDX CEU CHB CHS CLM ESN FIN GBR GIH GWD IBS ITU JPT KHV LWK MSL MXL PEL PJL PUR STU TSI YRI Private to population Private to continent Shared across continents Shared across all continents 18 million 12 million 24 million Individual Variant sites per genome (million) 3.8 4 4.2 4.4 4.6 4.8 5 MSL ESN LWK YRI GWD ACB ASW PUR CLM MXL PEL BEB ITU PJL STU GIH KHV JPT CHB CDX CHS TSI IBS CEU GBR FIN Singletons per genome (×1,000) 0 2 4 6 8 10 12 14 16 18 20 LWK GWD MSL ACB ASW YRI ESN BEB STU ITU PJL GIH CHB KHV CHS JPT CDX TSI CEU IBS GBR FIN PEL MXL CLM PUR A B c Figure 10.7 Population sampling in the 1000 Genomes (1KG) Project. (A) Polymorphic variants within sampled populations. The area of each pie is proportional to the number of polymorphisms within a population. Pies are divided into four slices, representing variants private to a population (darker color unique to population), private to a continental area (lighter color shared across continental group), shared across continental areas (light gray), and shared across all continents (dark gray). Dashed lines indicate populations sampled outside of their ancestral continental region. (B) The number of variant sites per genome. (C) The average number of singletons per genome. From The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature 526:68–74, 2015. https:// doi.org/10.1038/nature 15393 alleles A and a: AA, Aa, and aa. Assuming that there is no mutation or other violations of HWE conditions, the proportions of genotypes in F1 can be calculated by plugging in the P1 genotype pairings as such:
188 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Parent Generation (P1,2) Possible Mating Pairs P1 Genotype: AA P2 Genotype: AA A AA AA AA Aa Aa Aa AA Aa Aa Aa AA AA AA AA AA Aa Aa Aa Aa aa...
Ch10 · Pt15 CHAPTER 10 — Population Genetics for Genomic Medicine 189 Suppose p is the frequency of allele A, and q is the frequency of allele a in the gene pool, such that p + q = 1 for a hypothetical biallelic trait. Substituting in the frequency variables p and q (below) for their Aa F1 (AA x aa)p(1.0) Aa F1 (AA x Aa)p(0.5) AAF1 (AA x Aa)p(0.5) Aa F1 (Aa x aa)p(0.5) Aa F1 (Aa x Aa)p(0.25) AAF1 (Aa x Aa)p(0.25) AAF1 (AA x AA)p(1.0) AAF1 (Aa x AA)p(0.5) aa F1 (Aa x aa)p(0.5) aa F1 (Aa x Aa)p(0.25) Aa F1 (Aa x Aa)p(0.25) Aa F1 (Aa x AA)p(0.5) aa F1 (aa x aa)p(1.0) aa F1 (aa x Aa)p(0.5) Aa F1 (aa x Aa)p(0.5) Aa F1 (aa x AA)p(1.0) AA Aa aa AA Parent (p) Genotypes Offspring genotype frequencies given parental genotypes, assuming Hardy-Weinberg conditions (no mutation) Aa aa Now let us assume alleles combine into genotypes randomly; that is, mating in the population is completely random with respect to the genotypes at this locus. The chance that two A alleles will pair up to give the AA genotype is then p 2; the chance that two a alleles will come together to give the aa genotype is q 2; and the chance of having one A and one a pair, resulting in the Aa genotype, is 2pq (the factor 2 comes from the fact that the A allele could be inherited from one parent and the a allele from the other, or vice versa). This applies to all autosomal loci and to the X chromosome in females, but not to X-linked loci in males who have just one X chromosome. respective alleles in F1 genotype proportion equations (above), the table below expresses offspring genotype proportions for all possible mating pairs in P in terms of p and q. Genotype frequencies in offspring generation (F1), given parental genotype frequencies; assuming two alleles (A and a) are in HWE with population allele frequencies Freq[A] = p and Freq[a] = q; such that p + q = 1. Parent (P) Genotype Frequencies AA = p 2 AAF1 (p 2 x p 2)(1.0) = p 4 AAF1 (2pq x p 2)(0.5) = p 3q Aa F1 (2pq x p 2)(0.5) = p 3q Aa F1 (q 2 x p 2)(1.0) = p 2q2 AAF1 (p 2 x 2pq)(0.5) = p 3q AAF1 (2pq x 2pq)(0.25) = p 2q2 Aa F1 (2pq x 2pq)(0.25) = p 2q2 Aa F1 (q 2 x 2pq)(0.5) = pq 3 Aa F1 (p 2 x 2pq)(0.5) = p 3q Aa F1 (2pq x 2pq)(0.25) = p 2q2 aa F1 (2pq x 2pq)(0.25) = p 2q2 aa F1 (q 2 x 2pq)(0.5) = pq 3 Aa F1 (p 2 x q 2)(1.0) = p 2q2 Aa F1 (2pq x q 2)(0.5) = pq 3 aa F1 (2pq x q 2)(0.5) = pq 3 aa F1 (q 2 x q 2)(1.0) = q 4 AA = p 2 Aa = 2pq Aa = 2pq aa = q 2 aa = q 2 The Hardy-Weinberg principle states that the frequency of the three genotypes AA, Aa, and aa is given by the terms of the binomial expansion of (p + q)2 = p 2 + 2pq + q 2 = 1. If allele frequencies do not change from generation to generation, the proportion of genotypes will not change either; that is, the genotype frequencies from generation to generation will remain constant (at equilibrium) in the population if the allele frequencies (p and q) remain constant. When there is random mating in a population at HWE equilibrium, and genotypes AA, Aa, and aa are present in the proportions p 2:2pq:q 2, then genotype frequencies in the next generation will remain in the same relative proportions, p 2:2pq:q 2. This is proven as follows:
CHAPTER 10 — Population Genetics for Genomic Medicine 189 Suppose p is the frequency of allele A, and q is the frequency of allele a in the gene pool, such that p + q = 1 for a hypothetical biallelic...
Ch10 · Pt16 190 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Genotype frequencies in the parent generation (P) can be used to predict genotype frequencies in the first generation of offspring (F1) by adding up the resulting genotype frequencies from all possible mating pairs (P1,2). Constant Genotype frequency Proportions for a Population in Hardy-Weinberg Equilibrium (HWE) Genotype frequency Proportions remain constant in a population under Hardy-Weinberg conditions. AA x AA: AA x Aa: AA x aa: Aa x Aa: Aa x aa: = p 2(p 2 + 2pq + q 2) = p 2(p + q)2 = p 2(1)2 aa x aa: F1 AA p 2p2 = p 4 p 2q + p 3q = 2p3q 0 p 2q2 p 4 + 2p3p + p 2p2 = 2pq(p 2 + 2pq + q 2) = 2pq(p + q)2 = 2pq(1)2 2p3q + 2p2q2 + 2p2q2 + 2pq 3 = q 2(p 2 + 2pq + q 2) = q 2(p + q)2 = q 2(1)2 p 2p2 + 2pq 3 + q 4 0 0 0 p 3q + p 3q = 2p3q p 2q2 + p 2q2 = 2p2q2 p 2q2 + p 2q2 = 2p2q2 pq 3 + pq 3 = 2pq 3 0 0 0 0 p 2q2 pq 3 + pq 3 = 2pq 3 q 4:::::::::::::::: F2 Aa aa AA P Aa aa p 2 AA = p 2 Aa = 2pq aa = q 2 q 2 2pq AA p 2 Aa + + 2pq aa = 1 q 2 BOX 10.2 HARDY-WEINBERG EQUILIBRIUM ASSUMPTIONS The principle of Hardy-Weinberg equilibrium rests on the assumption that genotype frequency proportions remain constant over time because of the following: Random mating. Reproductive pairings are random with respect to the locus in question. No genetic drift. The population under study is sufficiently large such that alleles are not likely to dramatically rise or drop in frequency by random chance. No mutation. Rate of mutation is low such that allele frequencies are not impacted. No selection. Individuals are equally capable of passing on their genes, regardless of genotype, preserving equality between each allele frequency and its chance of being inherited. No gene flow. There has been no significant migration of individuals between populations with significantly different allele frequencies. A population that appears to meet these assumptions is in Hardy-Weinberg equilibrium. It is important to note that HWE does not require any particular values for p and q; whatever allele frequencies happen to be present in the population will result in genotype frequencies of p 2:2pq:q 2, and these relative genotype frequencies will remain constant from generation to generation as long as the allele frequencies remain constant and the other conditions introduced in Box 10.2 are met. This principle can be adapted for genes with more than two alleles. For example, if a locus has three alleles, with frequencies p, q, and r, the genotypic distribution can be determined from (p + q + r)2 = 1. In general terms, the genotype frequencies for any known number of alleles an with allele frequencies p 1, p 2, … pn can be derived from the terms of the expansion of (p 1 + p 2 + … pn)2. Applying HWE to the ACKR1 (Duffy blood group) example given earlier, with relative frequencies of the two alleles in the 1KG Project dataset 0.734 (for the T allele) and 0.266 (for the C allele), the relative proportions of the three combinations of alleles (genotypes) are Pr[T|T] = p 2 = 0.734 × 0.734 = 0.539 (for an individual having two T alleles), Pr[C|C] = q 2 = 0.266 × 0.266 = 0.071 (for two C alleles), and Pr[C|T] = 2pq = (0.734 × 0.266) + (0.734 × 0.266) = 0.39 (for individuals with one T and one C allele). These genotype frequencies were calculated assuming HWE; so, when applied to a population of 2504 individuals, the derived numbers of people with the three different genotypes (TT:CT:CC) should be equivalent to the observed genotype frequencies from the 1KG dataset. However, when we do this calculation (total population size × genotype frequency), the proportions of individuals with each genotype are 1350:977:177. This is very different from the actual observed proportions in Table 10.1 (1793:88:623), which indicates that the assumptions of HWE do not hold for the “population” defined as all individuals in the 1KG dataset. This
190 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Genotype frequencies in the parent generation (P) can be used to predict genotype frequencies in the first generation of offspring (F1) by a...
Ch10 · Pt17 CHAPTER 10 — Population Genetics for Genomic Medicine 191 makes sense, because the sampling scheme of the project was meant to identify individuals from different parts of the world to enable comparisons of average frequencies across the globe and cannot be considered a single population in HWE. Two key HWE assumptions that are violated in this scenario are: (1) Random mating, because we would not expect individuals to meet and reproduce with people from across the globe with equal chance compared to those close by; and (2) selection, since we know this allele is under positive selective pressure in locations that have a high incidence of malaria. Now, let us try this exercise again with a different variant in the same gene (e.g., ACKR1 rs 36007769; Ensembl) that is synonymous and therefore not predicted to change the protein, such that its frequency is not under the influence of natural selection (unless it is in strong LD with a functionally relevant variant that is under selection and carries it along). We can again use HWE to calculate predicted genotype frequencies from observed allele frequencies in 1KG and compare those results to the observed genotype frequencies reported for this dataset in Ensembl. Table 10.3 reports observed allele counts and frequencies for the synonymous SNV rs 36007769 in the ACKR1 (Duffy blood group) gene, as reported in Ensembl for the 1KG Project (Phase 3) dataset. Recall that the HWE equation p 2 + 2pq + q 2 = 1 can be used to calculate expected genotype frequencies, given observed allele frequencies such that the first and third terms (p 2 and q 2) are the expected frequencies of homozygous genotypes for alleles p and q, respectively, and the second term (2pq) is the expected frequency of heterozygotes. Table 10.4 reports the results of these calculations as derived genotype frequencies. Comparing these derived frequencies (0.992:0.008: 0.0) to observed genotype frequencies in the 1KG dataset (0.993:0.007:0.0), we can see that they are roughly equivalent. Given that the assumptions of HWE hold, we would expect these genotype frequencies to remain constant generation after generation. Although the 1KG dataset is not a genetically homogeneous population and it includes samples from many different parts of the world, the absence of selection acting on this variant and the fact that it is rare in a relatively large population (defined as the dataset) mean that genotype frequency calculations based on HWE are still useful. Upon further inspection of the incidence of this variant, 5 occurrences are in Central and South American ancestry (AMR) populations and 13 occurrences are in European ancestry (EUR) populations. Using the reported allele frequencies in those populations from 1KG, AMR (G: 0.993, A: 0.007), and EUR (G: 0.987, A: 0.013) to calculate expected genotype frequencies with HWE, they are equivalent to reported frequency proportions in AMR (0.986:0.014:0) and EUR (0.974:0.026:0). This suggests that HWE holds for each of these populations for this specific variant, and thus their frequency proportions should persist in subsequent generations. Applying Hardy-Weinberg Equilibrium to Autosomal Recessive Traits The major practical application of the Hardy-Weinberg principle in medical genetics is in genetic counseling for autosomal recessive conditions. For a disease such as phenylketonuria (PKU), there are hundreds of different pathogenic alleles with frequencies that vary among different population groups defined by geography and/ or ethnicity (see Chapter 13). Affected individuals can be homozygoues for the same pathogenic allele, but they are often compound heterozygotes for different pathogenic variants (see Chapter 7). For many conditions, it is convenient to consider all disease-causing alleles together and treat them as a single pathogenic allele, with frequency q, even when there is significant allelic heterogeneity among pathogenic alleles. Similarly, the combined frequency of all benign or nonpathogenic alleles, p, is given by 1 − q. Suppose we would like to know the frequency of all disease-causing PKU alleles in a population for use in genetic counseling, for example, to inform couples of their risk for having a child with PKU. If we were to attempt to determine the frequency of disease-causing PKU alleles directly from genotype frequencies, we would need to know the frequency of heterozygotes in the population, a frequency that cannot be measured directly because of the recessive nature of PKU. This is because heterozygotes are asymptomatic silent carriers (see Chapter 7), and their frequency in the population (i.e., 2pq) cannot be reliably determined directly from phenotype observation. TABLE 10.3 Allele Counts and Frequencies for the ACKR1 Duffy Blood Group Allele rs 36007769 observed and reported by the 1000 Genomes (1KG) Project Allele in ACKR1 rs 36007769 Observed Allele Counts Observed Allele Frequencies G 4990 0.996 A 18 0.004 Total 5008 1.0 From The 1000 Genomes Project Consortium: A global reference for human genetic variation, Nature 526:68–74, 2015. doi:10.1038/nature 15393. Accessed online via Ensembl. TABLE 10.4 Genotype Counts and Frequencies for the ACKR1 Duffy Blood Group Allele rs 36007769 Observed in the 1000 Genomes (1KG) Project and Derived Using HWE rs 36007769 Genotypes Observed Genotype Counts Observed Genotype Frequencies Derived Genotype Frequencies (HWE) G|G 2486 0.993 p 2 = (0.996)2 = 0.992 A|G 18 0.007 2pq = 2(0.996) (0.004) = 0.008 A|A 0 0 q 2=(0)2=0
CHAPTER 10 — Population Genetics for Genomic Medicine 191 makes sense, because the sampling scheme of the project was meant to identify individuals from different parts of the world to enable comparis...
Ch10 · Pt18 192 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE However, the frequency of affected homozygotes/ compound heterozygotes for disease-causing alleles in the population (i.e., q 2) could be determined directly, by counting the number of babies with PKU born over a given time period and identified through newborn screening (see Chapter 19), divided by the total number of babies screened during that same time period. Now, using HWE, we can calculate the pathogenic allele frequency (q) from the observed frequency of homozygotes/ compound heterozygotes alone (q 2), thereby providing an estimate (2pq) of the frequency of heterozygotes for use in genetic counseling. To illustrate this example further, consider a population in which the frequency of PKU is approximately 1 per 4500. If we group all disease-causing alleles together and treat them as a single allele with frequency q, then the frequency of affected individuals q 2 = 1/4500. From this, we calculate q = 0.015, and thus 2pq = 0.029. The carrier frequency for all disease-causing alleles lumped together in this population is therefore approximately 3%. For an individual known to be a carrier of PKU through the birth of an affected child in the family, there would then be an approximately 3% chance that he or she would find a new mate from the same population who would also be a carrier, and this estimate could be used to provide genetic counseling. Note, however, that this estimate applies only to the population in question; if the new mate was from a different genetic ancestral population where the frequency of PKU is much lower (e.g., 1 per 200,000), their chance of being a carrier would be only 0.6%. In this example, all PKU-causing alleles are collapsed for the purpose of estimating q. For other conditions, however, such as hemoglobin disorders that we will consider in Chapter 12, different pathogenic alleles can lead to very different conditions, and therefore it would make no sense to group all pathogenic alleles together, even when the same locus is involved. Instead, the frequencies of alleles leading to different phenotypes (e.g., sickle cell disease and β-thalassemia in the case of different pathogenic alleles at the β-globin locus) are calculated separately. Allele and Genotype Frequencies in X-Linked Conditions Recall from Chapter 7 that, for X-linked genes, there are three female genotypes but only two possible male genotypes. To illustrate the relationship between allele and genotype frequencies when a gene of interest is X linked, we use the trait known as X-linked red-green color blindness, which is caused by structural variants in genes encoding cell receptors that respond to photons of light at wavelengths we perceive as red and green (OPN1LW and OPNMW, respectively) that are adjacent to one another on the X chromosome. We use color blindness as an example because, as far as we know, it is not a deleterious trait (except for possible difficulties with traffic lights), and persons with color blindness are not subject to selection. In this example, we use the symbol cb to represent variants conferring some variation of color blindness and the symbol + for variants without color blindness, with frequencies q and p, respectively (Table 10.5). Because females have two X chromosomes, their genotypes are distributed like autosomal genotypes, but because color blindness variants are recessive, their homozygous and heterozygous genotypes are typically not distinguishable. In contrast, males with only one X chromosome will exhibit the trait with a single copy of a cb variant. As such, the frequency of color blindness in females is much lower than that in males (<1%). Frequencies of the two types of variants (cb and +) can be determined directly from the prevalence of the corresponding phenotypes in males. So, if the prevalence of colorblindness in a population of biologically male individuals (with an X and a Y chromosome) is roughly 8%, the frequency of cb variants in the population is likewise 0.08. Genotypes of unaffected females (homozygous or heterozygous) cannot be determined by looking at phenotypes; frequencies of variants for the trait that were ascertained by looking at phenotype frequencies in male individuals can be used to determine approximate variant frequencies for females using HWE. As shown in Table 10.5, ~15% of females are unaffected carriers. Among these heterozygous unaffected individuals, those who are pregnant with a male fetus have a 50% chance of giving birth to a male child with colorblindness (because they have a 50% chance of passing on the Xcb variant and this is the only X chromosome a male will receive from either parent). In contrast, a female fetus receives two copies of the X chromosome, so the 50% probability of an unaffected carrier transmitting the cb variant is instead the chance that a female fetus will also be a silent carrier. TABLE 10.5 X-Linked Genotype Frequencies and Prevalence of Color Blindness Biological Sex Phenotypes Prevalence Genotypes Genotype Frequencies Variant Frequencies Male (X/Y) Colorblindness Unaffected cb 8%+ 92% [Y] Xcb [Y] X+ 0.08 0.92 q=0.08 p=0.92 Female (X/X) Colorblindness Unaffected cb <1% Xcb Xcb Xcb X+ X+X+ q 2 = (0.08)2 = 0.0064 2pq = 2(0.08)(0.92) = 0.1472 p 2 = (0.92)2 = 0.8464
192 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE However, the frequency of affected homozygotes/ compound heterozygotes for disease-causing alleles in the population (i.e., q 2) could be de...
Ch10 · Pt19 CHAPTER 10 — Population Genetics for Genomic Medicine 193 Violating Assumptions of Hardy-Weinberg Equilibrium Underlying the principle of Hardy-Weinberg equilibrium and its use are several assumptions (see Box 10.2), not all of which can be met (or reasonably inferred to be met) by all populations. In this section, we provide a highlevel overview of the conditions and factors that contribute to violations of HWE assumptions: (1) nonrandom mating, (2) genetic drift, (3) mutation, (4) selection, and (5) gene flow. We have seen examples of these throughout the chapter, as population genetics is concerned with modeling and measuring shifts in allele frequencies. First, we have seen nonrandom mating at work in small genetic isolates, or founder populations, which (by definition) are isolated from reproduction events outside the group. In human populations, assortative mating, underlying population structure, or cryptic relatedness due to shared recent ancestry, and consanguinity can all lead to nonrandom mate choices. Assortative mating is a type of nonrandom mating in which individuals in a population engage in preferential reproductive choices. This may increase the frequency of variants contributing to traits that influence individuals in these reproductive choices and other variants in LD with them. Genetic drift refers to changes in allele frequencies over time due to random chance, which occur more quickly in smaller populations. When a new mutation occurs in a small population, its frequency is represented by only one copy among all the copies of that gene in the population. Random effects of the environment or other chance occurrences that are independent of the genotype (i.e., events that occur for reasons unrelated to whether an individual is carrying a pathogenic variant) can produce significant changes in the frequency of the disease allele when the population is small. Such chance occurrences disrupt Hardy-Weinberg equilibrium and cause the allele frequency to change from one generation to the next. During the next few generations, although the population size of the new group remains small, there may be considerable fluctuation until allele frequencies come to a new equilibrium as the population increases in size. HWE assumes no genetic drift – which requires an absence of migration in and out of the population by groups whose allele frequencies at loci of interest differ drastically from those of the population under HWE assumptions. This is a phenomenon called gene flow, whereby alleles are exchanged into and out of populations via migration and reproduction among individuals from different populations. Gene flow disrupts HWE because allele frequencies can change when new alleles are introduced into a population. Similarly, mutation (see Chapter 4) introduces new allelic variants into a population at random, which can influence stability of allele frequencies. When these newly introduced alleles (either through gene flow or mutation) confer some evolutionary advantage over existing alleles in the population, they will naturally rise in frequency due to pressures of natural selection, thereby disrupting HWE. Changes in allele frequencies due to selection or mutation usually occur slowly, in small increments, and cause much less deviation from HWE, at least for recessive diseases. This is because rates of new mutations are generally well below the frequency of heterozygotes for autosomal recessive diseases. The addition of new pathogenic alleles to the gene pool thus has little effect (in the short term) on allele frequencies for such diseases. In addition, most deleterious recessive alleles are hidden in asymptomatic heterozygotes and thus are not subject to selection. Consequently, selection is not likely to have major short-term effects on allele frequencies of these recessive alleles. Therefore, to a first approximation, HWE may apply even for alleles that cause severe autosomal recessive disease. Importantly, however, for dominant or X-linked conditions, mutation and selection do perturb allele frequencies from what would be expected under HWE, by substantially reducing or increasing certain genotypes in just a few generations. In practice, some violations of HWE we have discussed are more disruptive than others when applying the principle to human populations. For example, violating the assumption of random mating can cause large deviations from the expected frequency of individuals homozygous for an autosomal recessive condition. In contrast, changes in allele frequency due to mutation, selection, or migration usually cause more minor and gradual deviations from HWE. When HWE assumptions do not hold for a particular disease allele at a particular locus, it may be instructive to investigate why the allele and its associated genotypes are not in equilibrium as this may provide clues about the pathogenesis of the condition or point to historical events that have affected the frequency of alleles in different population groups over time. Mutation and Selection Balance in Traits With Different Modes of Inheritance In this section, we examine the concepts of mutation and selection through the lens of fitness, a heuristic device that indicates the likelihood of a mutation at a particular locus being eliminated, becoming stable, or becoming (over time) the predominant allele (or fixed) in a population. The frequency of an allele in a population at any given time represents a balance between the rate at which new alleles appear through mutation and the influence of selection on these alleles. If the mutation rate or the effectiveness of selection is altered, the allele frequency is expected to change. More formally, whether an allele is transmitted to the succeeding generation depends on its fitness, ω, which is a quantitative measure for the expected (average) number of offspring of affected persons who survive to
CHAPTER 10 — Population Genetics for Genomic Medicine 193 Violating Assumptions of Hardy-Weinberg Equilibrium Underlying the principle of Hardy-Weinberg equilibrium and its use are several assumptions...
Ch10 · Pt20 194 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE reproductive age. This measure is called relative fitness when compared with that of an appropriate control group. If a pathogenic allele is just as likely to be represented in the next generation compared to functionally neutral alleles, ω = 1. If an allele causes death or sterility, purifying, or negative, selection acts against it completely, and ω = 0. Values between 0 and 1 indicate transmission of the variant, and values of ω > 1 indicate positive selection increasing the variant’s frequency in a population. A related parameter is the coefficient of selection, s, which is a measure of the loss of fitness and is defined as 1− ω, that is, the proportion of pathogenic alleles that are not passed on and are therefore lost due to negative selection. When a genetic condition limits reproduction such that ω = 0 and s = 1, the variant conferring this trait is referred to as a genetic lethal. In the genetic sense, a variant that prevents reproduction by an adult is just as “lethal” as one that causes a very early miscarriage of an embryo, because in neither case is the variant transmitted to the next generation. Fitness is thus the outcome of the joint effects of survival and fertility. In the biological sense, relative fitness has no connotation of superior endowment but is simply a measure of comparative ability to contribute alleles to the next generation, on average. The frequency of pathogenic alleles in a population represents a balance between loss of pathogenic alleles through the effects of selection and gain of pathogenic alleles through recurrent mutation. A stable allele frequency will then be reached at whatever level balances the two opposing forces: one (selection) that removes pathogenic alleles from the gene pool and one (de novo mutation) that adds new ones back. The mutation rate per generation, µ, at a locus with pathogenic variants must be sufficient to account for the fraction of all pathogenic alleles that are lost by selection from each generation. That is,
µ=sq where µ is the mutation rate, s is the coefficient of selection, and q is the allele frequency. If a pathogenic allele for a condition has a dominant mode of inheritance and is deleterious but no...
Ch10 · Pt21 CHAPTER 10 — Population Genetics for Genomic Medicine 195 homozygotes for the recessive allele. It follows that the mathematical relationship between genotype and allele frequencies described by HWE holds for most practical purposes in the case of an autosomal recessive disease. In contrast to recessive pathogenic alleles, dominant pathogenic alleles are exposed directly to selection. Consequently, the effects of selection and mutation are more obvious and can be more readily measured for dominant traits. A genetic lethal dominant allele, if fully penetrant, will be exposed to selection in heterozygotes, thus removing all alleles responsible for the disorder in a single generation. Several human diseases are thought to be autosomal dominant traits with zero or near-zero fitness and thus always result from de novo, rather than inherited, autosomal dominant variants. This is a point of great significance for genetic counseling, and examples of these conditions are listed in Table 10.6. In some of these conditions, the specific pathogenic alleles are known, and family studies have revealed de novo mutations in affected individuals that were not inherited from the parents. In other conditions, the responsible genes are not known, but paternal age effects (Chapter 4) have been observed, which suggests a possible mechanism of de novo mutations in the paternal germline. The implication for genetic counseling is that parents of a child with an autosomal dominant (but genetically lethal) condition will typically have a very low risk of recurrence in subsequent pregnancies because the condition would require another independent de novo mutation. A caveat to keep in mind is the possibility of germline mosaicism (See Fig. 7.17) and the possibility of abundant de novo mutations in a germline heavily exposed to mutagens. In clinically relevant conditions that have an X-linked recessive mode of inheritance, selection acts on hemizygous males but not in heterozygous females, except for the small proportion of females who are manifesting heterozygotes with reduced fitness (see Chapter 7). In this brief discussion, we assume that heterozygous females do not have reduced fitness. Because males have one X chromosome and females have two, the pool of X-linked alleles in the entire population’s gene pool is partitioned, such that one-third of pathogenic alleles are in males and two-thirds are in females. As we saw in the case of autosomal dominant variants, pathogenic alleles lost through selection must be replaced by recurrent new mutations to maintain the observed disease incidence. If the incidence of an X-linked condition is not changing, and selection is operating (only) against hemizygous males, the mutation rate, µ, must equal the coefficient of selection, s (i.e., the proportion of pathogenic alleles that are not passed on), times q (the pathogenic allele frequency), adjusted by a factor of 3, since selection is operating only on the third of pathogenic alleles in the population that are present in males. The equation is thus: µ=sq/3 For an X-linked genetic lethal condition, s = 1, and one-third of all copies of the pathogenic allele are lost from each generation; so, at equilibrium, these must be replaced by de novo mutations. Thus, roughly one-third of all persons who have X-linked lethal disorders are predicted to carry a de novo mutation, and their unaffected mothers have a low risk of future pregnancies harboring the same disorder (in the absence of germline mosaicism). The remaining two-thirds of mothers of individuals with an X-linked lethal disorder are predicted to be carriers, each with a 50% future risk of conceiving an affected child, given that it is male. However, the prediction that two-thirds of mothers of individuals with an X-linked lethal disorder are carriers of a disease-causing variant assumes that mutation rates in males and in females are equal. Given that the germline mutation rate is higher in males with advanced paternal age than in females, the chance of a de novo mutation occurring in the egg is very low. (Note: the impact on genetic counseling related to these sex-dependent considerations of mutation rates will be discussed in Chapter 17). Most mothers of affected children are carriers, having most likely inherited novel variants from unaffected fathers, which they then have a 50% chance of passing on to their children. Taking advantage of the statistical property that the probability of two independent events occurring is equal to the product of probabilities of each separate event, Pr[A and B] Pr[A] Pr[B] = × where A and B are independent events, we can therefore calculate the total risk of an unaffected carrier of an X-linked condition having a child with the condition as: Pr[male child] Pr[passing on the pathogenic allele] (0.5)( × = 0.5) 0.25 = TABLE 10.6 Genetic Conditions Occurring via De Novo Mutations With Zero Fitness Condition/ Phenotype Description Atelosteogenesis Early lethal form of short-limbed skeletal dysplasia Cornelia de Lange syndrome Intellectual disability, micromelia, synophrys, and other abnormalities; can be caused by pathogenic variants in the NIPBL and other genes Developmental and epileptic encephalopathy Intellectually disability and early-onset seizures; can be caused by de novo variants in >50 different genes Osteogenesis imperfecta, type II Perinatal lethal type, with a defect in type I collagen (COL1A1, COL1A2) (see Chapter 13) Thanatophoric dysplasia Early lethal form of skeletal dysplasia due to specific de novo pathogenic variants in the FGFR3 gene (see Fig. 7.6 C)
CHAPTER 10 — Population Genetics for Genomic Medicine 195 homozygotes for the recessive allele. It follows that the mathematical relationship between genotype and allele frequencies described by HWE h...
Ch10 · Pt22 196 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE such that the probability that an unaffected carrier will give birth to a child with an X-linked genetic lethal condition is 25% or one-fourth, given that the other parent is not affected. In less severe disorders, such as hemophilia A (Case 21), the proportion of affected individuals representing new mutations is less than one-third (~15%). Because the treatment of hemophilia has improved significantly, the total frequency of pathogenic alleles can be expected to rise rapidly and to reach a new equilibrium. Assuming that the mutation rate at this locus stays the same over time, the proportion of those with hemophilia whose pathogenic variant arises de novo will decrease, but the overall incidence of the disease will increase. Such a change would have significant implications for genetic counseling for this condition (see Chapter 17). Although certain pathogenic alleles may be deleterious in homozygotes, there may be environmental conditions in which heterozygotes for some conditions have increased fitness relative to homozygotes for both the pathogenic allele and the reference allele. This is called heterozygote advantage because even a slightly greater relative fitness of heterozygotes can lead to an increase in frequency of an allele that is severely detrimental in homozygotes. This is because heterozygotes greatly outnumber homozygotes in the population. A situation in which selective forces operate to both maintain a deleterious allele and remove it from the gene pool is often referred to as balancing selection. A well-known example of heterozygote advantage is resistance to malaria in individuals who are heterozygous for the pathogenic allele that causes sickle cell disease (Case 42). This pathogenic variant in the β-globin gene HBB has reached its highest frequency in certain regions of Africa and Southeast Asia, where malaria is endemic and heterozygotes have greater relative fitness than either type of homozygote, due to their resistance to malarial infection. In the presence of (mosquito) vectors that carry malaria-inducing parasites, homozygotes without the trait allele are highly susceptible; they may become infected and are severely (even fatally) affected. Homozygotes for the pathogenic allele are even more disadvantaged, with a relative fitness that approaches zero due to severely debilitating hematological disease (see Chapter 12). Heterozygotes, on the other hand, have red blood cells that are inhospitable to the malarial parasite but do not typically undergo the characteristic sickling that leads to pain crises in active sickle cell disease. As such, these heterozygotes have a much greater relative fitness than homozygotes for the typical β-globin allele. Over time, the pathogenic allele for sickle cell disease has reached a frequency as high as 0.15 in some areas of the world that are endemic for malaria, far higher than could be accounted for by recurrent mutation alone. Heterozygote advantage in sickle cell disease offers a clear example of how assumptions of HWE are violated when the mathematical relationship between allele and genotype frequencies diverges from expected values, (in this case) due to the effects of balancing selection. To firmly ground this example, let us consider the sickle cell allele in the β-globin gene HBB, rs 334 (c.20 A>T [p. Glu 7Val]), in which the pathogenic allele β S is under balancing selection due to heterozygote advantage. Let us define the benign (nonpathogenic) allele β+ such that the two alleles β S and β+ give rise to three genotypes: β+|β+ (unaffected; homozygous), β+|β S (unaffected carriers; heterozygous), and β S|β S (affected). In a study of whole-genome sequence data from 2932 individuals aggregated across the 1000 Genomes Project, the African Genome Variation Project, and Qatar, balancing selection on the sickle cell variant β S is estimated to have conferred strong heterozygote advantage, having reached its equilibrium frequency of 12% after 87 generations (while the initial mutation is dated back 259 generations, or ~7300 years ago). Using the β S allele’s reported equilibrium frequency of 0.12 (q), we can calculate the expected ratio of genotypes under HWE (p 2:2pq:q 2) and compare this to observed genotype frequencies in the 1KG dataset for populations that have the β S allele at or near its equilibrium frequency. Given that q = 0.12 and p = 1 – q, the frequency of p = 1 – 0.12 = 0.88. From this, we can use HWE to calculate expected genotype frequencies as follows: Pr[affected homozygotes (β S|β S)] = q 2 = (0.12) * (0.12) = 0.014; Pr[unaffected homozygotes (β+|β+)] = p 2 = (0.88) * (0.88) = 0.774; and Pr[heterozygotes (β S|β+)] = 2pq = 2*(0.88) * (0.12) = 0.211. Thus the expected genotype proportions of p 2:2pq:q 2 are 0.774:0.211:0.014. Investigating rs 334 allele frequencies in the 1KG dataset through Ensembl, two populations sampled from Africa appear to have the β S allele at or near its equilibrium frequency (0.12): the Esan in Nigeria (ESN) with β S at 12.1% and the Mende in Sierra Leone (MSL) with β S at 12.4% in the population. Observed genotype frequencies for these populations are as follows: 0.758:0.242:0 (for ESN) and 0.753:0.247:0 (for MSL). In both these populations, the observed proportions of heterozygous (β S|β+) individuals exceed what was predicted assuming HWE, whereas the observed number of unaffected homozygotes (β+|β+) and affected homozygotes (β S|β S) are below what was predicted. This trend reflects balancing selection at this locus, illustrating how forces of selection, operating both negatively on the relatively rare β S|β S genotype and positively on the more common β S|β+genotype, cause deviation from HWE. The effects of balancing selection on malaria resistance are also apparent in other infectious diseases. For example, many people with the severe renal disease known as focal segmental glomerulosclerosis are homozygotes for certain alleles in the coding region of the APOL1 gene that encodes the apolipoprotein L1. Apolipoprotein L1 is a serum factor that kills the trypanosome parasite
196 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE such that the probability that an unaffected carrier will give birth to a child with an X-linked genetic lethal condition is 25% or one-four...
Ch10 · Pt23 CHAPTER 10 — Population Genetics for Genomic Medicine 197 Trypanosoma brucei, which causes trypanosomiasis (sleeping sickness). The same variants that increase one’s risk of severe kidney disease in homozygotes tenfold over the rest of the population protect heterozygotes carrying these variants against trypanosomes (e.g., T. brucei rhodesiense) that have developed resistance to wild-type apolipoprotein L1. As a result, the frequency of heterozygous carriers for these alleles can be as high as ~45% in parts of the world in which the rhodesiense trypanosomiasis is endemic. GENOMIC VARIATION AND BIASES IN POPULATION DATASETS The type and amount of information available to researchers and clinicians for our work is the foundation for discoveries, diagnostics, treatment regimens, and approaches to measuring outcomes. Missing data is inevitable – but uncertainty arises in the interpretation and portability of findings when the amount and type of data available is missing or of differential quality in a nonrandom ways, that is, if certain groups or patient populations are better represented in databases than others, for example. Ascertainment bias refers to systemic biases in observations that steer our research or interpretations in a direction based on what is observed – without having any knowledge of that which remains unobserved (and could potentially change the result or interpretation if revealed). For example, >80% of genomics research to date has been conducted on people of mostly European ancestries, so population reference data and genomic databases are heavily biased toward a European genomic background. Fig. 10.8 shows broad ancestry groups of GWAS participants from 2003 to 2018, compared with the world’s population, showing Europeans are vastly overrepresented among GWAS participants relative to their global census representation. The impact of this European bias is far-reaching, such that genetic tests are designed primarily to capture variation that commonly contributes to disease on a European genomic background (or is associated with the causal variant via Linkage Disequilibrium; see Chapter 11). Genetic or locus heterogeneity in a trait means that variation in different regions of the genome can be pathogenic for the same trait. The ascertainment bias toward prevalence in European ancestries contributes to higher rates of variants of unknown or uncertain significance (VUS) in non-Europeans, as well as higher rates of false negative diagnoses, due to missing genetic heterogeneity. Higher false positive rates have also been documented, as in the case of a variant associated with hypertrophic cardiomyopathy (and curated as pathogenic based on information from European ancestry individuals), which was later shown to be benign in African American patients – after many had already undergone an invasive prophylactic intervention. Exclusion of individuals with recent African ancestries is a mistake for any study of the genetic underpinnings of health and disease, due to the wealth of genetic variation that exists in African populations. Fig. 10.9 illustrates this point by showcasing not only how much ancestral variation is present across the African continent (relative to reference data) but also how biased representations of global genomic diversity can be when restricted to continental ancestry groupings that seek to “balance” representation across the canonical discrete categories, as in Fig. 10.9B. Information disparity in the genomic knowledgebase between individuals with primarily European ancestries and everyone else is responsible for differences in the utility and accuracy of clinical genetic testing. This further exacerbates health and health care disparities at Figure 10.8 Ancestry of GWAS participants over time, as compared with the global population. Cumulative data, as reported by the GWAS catalog. Individuals whose ancestry is “not reported” are not shown. From Martin et al. Clinical use of current polygenic risk scores may exacerbate health disparities, Nat Genet 51:584–591, 2019. https://doi.org/10.1038/s 41588-019-0379-x
CHAPTER 10 — Population Genetics for Genomic Medicine 197 Trypanosoma brucei, which causes trypanosomiasis (sleeping sickness). The same variants that increase one’s risk of severe kidney disease in h...
Ch10 · Pt24 198 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Population names are sample providers’ collected the samples. self-identifications, or the descriptions used by those who Six geographical labels correspond to the recent origins of the populations that fall mainly into the seven ancestry clusters produced by an algorithm.† A Individual genomes Each horizontal bar corresponds to the genome of a single person, in this case from a population identified or self-identified as African American. This person has 42% Eurasian ancestry, 5% African Great Lakes ancestry and 53% West African ancestry. Individuals Populations from Africa sampled: 85% (unpublished) Populations from Africa Adjusting the sampling and using an African-centric data set creates a more representative view. sampled: 13.5%. East and South Asia, Americas (Indigenous) North Africa, Europe, Americas South Asia**, Horn of Africa African Great Lakes Africa Western Nama Mende Zulu Africa Middle East Europe Central and South Asia East Asia America Oceania Esan Yoruba African Mandinka African American Caribbean Banyarwanda Bakiga Barundi Banyankole Baganda Rwandese Gumuz Somali Luhya Oromo Wolayta Amhara Egyptian Gujarati Mexican British Italian Han Chinese Southern Africa ‡ B A C B Figure 10.9 Compiling genotype data from individuals can show the genetic diversity of populations. (A) An analysis from 2008* suggested significant genetic differences between seven continental populations (B) But only 13.5% of the populations represented were from Africa. Boosting representation to 85% and sampling more broadly across the continent (C) underlines that the level of genetic variation within Africa is equivalent to that seen between continents. Figure adapted from Carlson J, Henn BM, Al-Hindi DR, Ramachandran S. Counter: the weaponization of genetics research by extremists, Nature 610: 2022. **South Asia appears twice because Gujarati people in India have intermediate allele frequencies.
198 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Population names are sample providers’ collected the samples. self-identifications, or the descriptions used by those who Six geographical l...
Ch10 · Pt25 CHAPTER 10 — Population Genetics for Genomic Medicine 199 the population level due to structural racism that disproportionately impacts Black patients of all ancestries (e.g., the well-documented undertreatment of pain in African Americans). We will ground this discussion in the practical context of clinical variant interpretation. Guidelines for clinical variant interpretation protocols published by Association for Molecular Pathology (AMP) and the American College of Medical Genetics and Genomics (ACMG) include population-level information (e.g., allele frequency in population databases), among other data (e.g., functional evidence) to guide decisions about whether to designate a variant as benign, likely benign, likely pathogenic, pathogenic, or uncertain significance. For variants detected in European populations, the genomic knowledgebase is robust and reliable enough such that the absence of a variant from population databases can be considered evidence for pathogenicity. This is because these populations are well represented among reference datasets, so the absence of a variant can be reasonably interpreted as evidence that it is not tolerated in the population. In contrast, most global populations have not been included in genomic datasets at a comparable rate or magnitude, so this interpretation may be subject to limited certainty. Instead, it may be that absence or very low frequency of a variant in population databases reflects insufficient representation, and more rigorous sampling would reveal higher population-allele frequencies than would be consistent with predicted pathogenicity of the variant. Several efforts are underway to improve diversity and inclusion in genomic databases; Box 10.3 details two of these examples. As biomedical researchers and clinicians, we must acknowledge that the ways in which we collect, BOX 10.3 “A TALE OF TWO INITIATIVES” (BY LAURA ARBOUR) The Genome Aggregation Database (gnom AD) is an international collaborative effort of researchers utilizing available data sources of genomic variation compiled through the Broad Institute (https://gnomad.broadinstitute.org). It is used frequently as a clinical tool in the diagnosis of rare, severe, genetic disease. The frequency of variants present in the dataset and reported geographical ancestry of origin are openly available to clinicians and researchers, which aids in the first steps of consideration for pathogenicity of variants (common variants are unlikely to cause severe, early onset disease). The combination of whole genomes and exomes of more than 140,000 unrelated individuals contributes to the genomic reference database. Although there are ongoing efforts to increase diversity in public genomic databases including gnom AD, the problem is that not all populations are represented in available datasets for a multitude of reasons, therefore not all children or families with rare genetic conditions will have the same opportunity for a precise diagnosis in a timely manner. The lack of genomic reference data increases the “genomic divide,” where those with the greatest health disparities benefit least from genomic advances. The impact of lack of Indigenous genomic data is staggering when it is considered that there are more than 370 million Indigenous people spanning 90 countries worldwide (https://www.un.org/esa/socdev/unpfii/ documents/5session_factsheet 1.pdf.). Two initiatives aim to address this issue for Indigenous patients with genetic conditions. In parallel, the Silent Genomes Project (Canada) and the Aotearoa Variome (New Zealand) are developing genomic data reference databases that are prioritized for genomic health care (https://www.frontiersin.org/ articles/10.3389/fpubh.2020.00111/full) and may also be used for health research. These initiates are led or coled by Indigenous scholars, and variant use and release mechanisms are being developed and are informed by long-standing ethical frameworks from within their respective countries (“DNA on Loan” and the Te Mata Ira guidelines for medical genomics with Māori) and are consistent with the recent International Indigenous Data Sovereignty Interest Group “CARE” principles, CARE being the acronym for “Collective Benefit, Authority to Control, Responsibility and Ethics” (https://datascience. codata.org/articles/10.5334/dsj-2020-043/). The CARE principles are Indigenous focused but are meant to complement the “FAIR principles (Findable, Accessible, Interoperable, Reusable) which are Guiding Principles for scientific data management and stewardship” (https:// www.nature.com/articles/sdata 201618). The CARE principles support the notion of benefit for, and selfdetermination of, Indigenous people, consistent with the United Nations Declaration of the Rights of Indigenous Peoples (adopted by the UN General Assembly in 2007 and endorsed by law in Canada [Bill C-15] in 2021). Both the Silent Genomes Project and Aotearoa Variome will start with sequencing the samples of consented individuals within their countries. Storage, use, and release of variants for clinical and possibly research purposes is being informed by local Indigenous perspectives and governance mechanisms. The Silent Genomes Project will also assess the efficacy of the Indigenous Background Variant Library (IBVL) in a cohort of Indigenous children who have gone through diagnosis without the IBVL. The primary goal of both initiatives is to reduce, and not increase, health disparities with genomic advances. References (also integrated above): https://www.un.org/esa/socdev/unpfii/ documents/5session_factsheet 1.pdf https://www.genomics-aotearoa.org.nz/projects/ aotearoa-nz-genomic-variome https://www.bcchr.ca/silent-genomes-project Indigenous genomic databases: Pragmatic considerations and cultural contexts NR Caron, M Chongo, M Hudson, L Arbour, WW Wasserman, S Robertson, S Correard, P Wilcox: Front Public Health 8:111, 2020. doi:10.3389/fpubh.2020.00111; https://www.nature.com/articles/sdata 201618 https://datascience.codata.org/articles/10.5334/ dsj-2020-043/ https://www.un.org/development/desa/indigenouspeoples/declaration-on-the-rights-of-indigenous-peoples.html https://www.mltaikins.com/indigenous/senatepasses-undrip-bill-c-15/
CHAPTER 10 — Population Genetics for Genomic Medicine 199 the population level due to structural racism that disproportionately impacts Black patients of all ancestries (e.g., the well-documented unde...
Ch10 · Pt26 200 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE analyze, visualize, and report or publish our data on human population genetics are crucially important. Our analytic approaches are heavily influenced by the historical, cultural, social, and political contexts in which we conduct our research and clinical practices. How we think about and represent categories of difference and similarity within and among populations matters – both for science and medicine, but also for society. Members of the public with harmful political agendas have weaponized figures illustrating admixture mapping that are published in peer-reviewed journals (e.g., Fig. 10.9B), claiming such clustering methods support their ideologies that are steeped in biological racism. We must counter these efforts and work to prevent further misconceptions from spreading, through responsible and trustworthy research and reporting. Until we have a more robust and complete picture of global genomic variation and how it contributes to complex disease etiology, it is imperative that we exercise caution when implementing existing (and new) methodologies to analyze genome-wide data. For example, polygenic risk scores, or polygenic scores (PRS), have the potential to exacerbate both conceptual and practical issues related to equity in human population genetics and medicine. Figs. 10.10 and 10.11 provide a highlevel overview of how PRS are constructed, and the LD LD LD tag SNP2 causal SNP2 tag SNP3 causal SNP3 predicted phenotype Total number of SNPs m Y j=1 gj j = ∑ tag SNP genotype SNP weight (adjusted effect size) Model Adjustments tag SNP1 causal SNP1 effect sizes '3 GWAS SIGNAL IN DISCOVERY POPULATION SNPs and effect sizes identified by GWAS in discovery population to predict phenotype or trait of interest in a target population. PHENOTYPE PREDICTION IN TARGET POPULATION β β '2 β '1 β Figure 10.10 Construction of a polygenic risk score (PRS) to predict a trait of interest (Y) in a target population using SNPs associated with the trait (gj) and their effect sizes (βj) in a genome-wide association study (GWAS) in a discovery population. Variants identified by GWAS in the discovery population are not necessarily causing or contributing to variation of the trait in the discovery population, but these “tag SNPs” are linked to causal SNP(s) such that they signal a candidate region to be further investigated. ALLELE FREQUENCIES IN DISCOVERY POPULATION ALLELE FREQUENCIES IN TARGET POPULATION Tag SNPs from discovery GWAS signal candidate loci associated with a trait; causal SNPs unknown LD structure, gene x environment, and gene x gene interactions may limit prediction accuracy LD LD LD LD Gx E Gx G LD LD tag SNP2 causal SNP2 tag SNP2 causal SNP2 tag SNP3 tag SNP3 causal SNP3 causal SNP3 causal SNP4 tag SNP1 tag SNP1 = 0.002 causal SNP1 tag SNP1 causal SNP1 No LD causal SNP1 = 0.002 tag SNP2 = 0.0003 causal SNP2 = 0.0003 tag SNP1 = 0.01 causal SNP1 = 0.01 tag SNP1 = 0.005 causal SNP1 = 0.002 tag SNP2 = 0.0008 causal SNP2 = 0.0008 tag SNP3 = 0.01 causal SNP3 = 0.01 Figure 10.11 Allele frequencies of tag SNPs and causal SNPs are equal in the discovery population because these variants are in linkage disequilibrium (LD). When LD structure differs between discovery and target populations, associations between tag SNPs and causal SNPs in the discovery GWAS may not be replicated in the target population, limiting the predictive power of this model. Differences in environmental factors and gene-by-environment interactions (Gx E) between discovery and target populations may impact the accuracy of prediction in the target population, particularly for multifactorial traits. The presence of other variants in the causal pathway with gene-by-gene interactions (Gx G, or epistatic effects) in one population, but not the other, may also limit the portability of PRS between populations.
200 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE analyze, visualize, and report or publish our data on human population genetics are crucially important. Our analytic approaches are heavily...
Ch10 · Pt27 CHAPTER 10 — Population Genetics for Genomic Medicine 201 conceptual as well as technical pitfalls of trying to predict phenotypic variation in one (target) population using GWAS results from another (discovery) population. In this simplified model, PRS construction involves several assumptions about homogeneity in genetic contributions to disease, allele frequencies, LD structure, Gx G, and Gx E between a discovery population and the target (prediction) population. Allele frequencies and LD structure may differ substantively between populations (depending on how they are defined and their underlying characteristics). Similarly, differential Gx E and Gx G between populations can dramatically affect the results of genomic investigations and obscure the role of genetic variants in disease etiology. Clinical professionals must be aware that PRS still have a long way to go before one could argue that they offer enhanced utility beyond the current standards of care and could make matters worse in the meantime. What is true of PRS is the same for all approaches we use in genomic research and medicine. We must critically examine our underlying assumptions about what factors contribute most to health and disease, consider the impact of those assumptions, investigate the history and biases of approaches we seek to use, and question the foundations of what we think we know – to make room for more curiosity and innovation that will lead to novel discoveries and more precise genomic medicine. ACKNOWLEDGMENT Some language and sections included in this chapter were inherited from previous (published) versions of the textbook. Laura Arbour contributed “A Tale of Two Initiatives”. Conversations with clinical genetics professionals and other interdisciplinary collaborations through the NIH-funded Clinical Genome Resource (Clin Gen) Ancestry & Diversity Working Group helped motivate the development of new content for this revision. We thank Sonja Rasmussen for contributing to this chapter. GENERAL REFERENCES Li CC: First course in population genetics, Pacific Grove, 1975, Boxwood Press. Nielsen R, Slatkin M: An introduction to population genetics, Sunderland, 2013, Sinauer Associates, Inc. Dorothy R: Fatal invention: How Science, Politics, and Big Business Re-Create Race in the Twenty-First Century, New York, 2011, New Press. Popejoy AB, Crooks KR, Fullerton SM, et al: Clinical Genome Resource (Clin Gen) Ancestry and Diversity Working Group: Clinical genetics lacks standard definitions and protocols for the collection and use of diversity measures, Am J Hum Genet 107(1):72–82, 2020. https://doi.org/10.1016/j.ajhg.2020.05.005 Royal CD, Novembre J, Fullerton SM, et al: Inferring genetic ancestry: opportunities, challenges and implications, Am J Hum Genet 86:661–673, 2010. REFERENCES FOR SPECIFIC TOPICS American Society of Human Genetics: ASHG denounces attempts to link genetics and racial supremacy, Am J Hum Genet 103:636, 2018. Behar DM, Yunusbayev B, Metspalu M, et al: The genome-wide structure of the Jewish people, Nature 466:238–242, 2010. Borrell LN, Elhawary JR, Fuentes-Afflick E, et al: Race and genetic ancestry in medicine – A time for reckoning with racism, N Engl J Med 384:474–480, 2021. https://doi.org/10.1056/NEJMms 2029562 Corona E, Chen R, Sikora M, et al: Analysis of the genetic basis of disease in the context of worldwide human relationships and migration, PLo S Genet 9:e 1003447, 2013. Gregg AR, Aarabi M, Klugman S, et al: ACMG Professional Practice and Guidelines Committee: Screening for autosomal recessive and X-linked conditions during pregnancy and preconception: a practice resource of the American College of Medical Genetics and Genomics (ACMG, Genet Med 23(10):1793–1806, 2021. https:// doi.org/10.1038/s 41436-021-01203-z Henn BM, Cavalli-Sforza LL, Feldman MW: The great human expansion. Proceedings of the National Academy of Sciences of the United States of America 109:17758–64. PMID 23077256. https://doi. org/10.1007/S12045-019-0830-4 Howes R, Patil A, Piel F, et al: The global distribution of the Duffy blood group. Nat Commun 2:266, 2011. https://doi.org/10.1038/ ncomms 1265 Kaseniit KE, Haque IS, Goldberg JD, Shulman LP, Muzzey D: Genetic ancestry analysis on >93,000 individuals undergoing expanded carrier screening reveals limitations of ethnicity-based medical guidelines, Genet Med 22(10):1694–1702, 2020. https://doi.org/10.1038/ s 41436-020-0869-3 Kumar R, Seibold MA, Aldrich MC, et al: Genetic ancestry in lungfunction predictions, N Engl J Med 363:321–330, 2010. Lewontin RC: The apportionment of human diversity. In: Dobzhansky T, Hecht MK, Steere WC, editors: Evolutionary biology: volume, 6, New York, 1972, Springer. Martin A, Kanai M, Kamatani Y, Okada Y, Neale BM, Daly MJ: Clinical use of current polygenic risk scores may exacerbate health disparities, Nat Genet 51:584–591, 2019. https://doi.org/10.1038/ s 41588-019-0379-x Novembre J, Johnson T, Bryc K, et al: Genes mirror geography within Europe. Nature 456:98–101, 2008. https://doi.org/10.1038/ nature 07331 Novembre J, Peter BM: Recent advances in the study of fine-scale population structure in humans. Curr Opin Genet Dev 41:98–105, 2016. https://doi.org/10.1016/j.gde.2016.08.007 Peterson RE, Kuchenbaecker K, Walters RK, et al: Genome-wide association studies in ancestrally diverse populations: Opportunities, methods, pitfalls, and recommendations, Cell 179(3):589–603, 2019. https://doi.org/10.1016/j.cell.2019.08.051 Richards S, Aziz N, Bale S, et al: Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology, Genet Med 17:405–423, 2015. https://doi.org/10.1038/gim.2015.30 Sankararaman S, Mallick S, Dannemann M, et al: The genomic landscape of Neanderthal ancestry in present-day humans, Nature 507:354–357, 2014. Shriner D, Rotimi CN: Whole-genome-sequence-based haplotypes reveal single origin of the sickle allele during the Holocene Wet Phase, Am J Hum Genet 102(4):547–556, 2018. https://doi.org/10.1016/j. ajhg.2018.02.003. Epub 2018 Mar 8. PMID: 29526279; PMCID: PMC5985360. Wastnedge E, Waters D, Patel S, et al: The global burden of sickle cell disease in children under five years of age: a systematic review and meta-analysis, J Global Health 8(2):021103, 2018. https://www. ncbi.nlm.nih.gov/pmc/articles/PMC6286674/ Wexler (Need REF for Venezuela finding) Wojcik GL, Graff M, Nishimura KK, et al: Genetic analyses of diverse populations improves discovery for complex traits, Nature 570:514– 518, 2019. https://doi.org/10.1038/s 41586-019-1310-4
CHAPTER 10 — Population Genetics for Genomic Medicine 201 conceptual as well as technical pitfalls of trying to predict phenotypic variation in one (target) population using GWAS results from another...
Ch10 · Pt28 202 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE PROBLEMS 1. A short tandem repeat (STR) variant consists of 5 different alleles, each with a frequency of 0.20 in a population. a. What proportion of individuals in this population would you expect to be homozygous at this locus? b. What proportion of the population is likely to be heterozygous at this locus? c. What proportion of individuals would be homozygous, and what proportion would be heterozygous if the 5 alleles had different frequencies of 0.40, 0.30, 0.15, 0.10, and 0.05? 2. In a population with allele frequencies in Hardy-Weinberg equilibrium, three genotypes are present in the following proportions: A/A, 0.81; A/a, 0.18; a/a, 0.01. a. What are the allele frequencies of A and a? b. What will their frequencies be in the next generation, assuming the conditions of Hardy-Weinberg equilibrium hold? 3. In a screening program designed to detect carriers of an autosomal recessive condition, the carrier frequency in a specific founder population was approximately 4%. a. Calculate the frequency of the pathogenic allele in this population (assuming only one). b. If the genetic fitness of affected individuals is zero, what proportion of possible reproductive pairings in this population could produce an affected child? c. What is the prevalence of unaffected carriers among offspring of couples in which both partners are heterozygous for the trait? d. If the pathogenic allele for this condition had the same frequency but were instead inherited through an autosomal dominant mode of inheritance (fully penetrant), what would be the frequency of unaffected adult carriers? What if it were X-linked dominant? X-linked recessive? 4. Which of the following populations is in Hardy-Weinberg equilibrium, based on the information provided? a. A/A, 0.70; A/a, 0.21; a/a, 0.09. b. A/A, 0.32; A/a, 0.64; a/a, 0.04. c. A/A, 0.64; A/a, 0.32; a/a, 0.04. What explanations could you offer to explain the frequencies in those populations that are not in equilibrium? 5. You are consulted by a couple, Meera and Arjun, who tell you that Meera’s sister has Hurler syndrome (a mucopolysaccharidosis) and that they are concerned that they might have a child with the same syndrome. Hurler syndrome is inherited as an autosomal recessive trait with an estimated prevalence of 1 in 90,000 individuals in a large population. a. What is the chance that Arjun is heterozygous for the pathogenic allele? b. What is the chance that Meera is a carrier of the pathogenic allele for Hurler syndrome? c. If Meera and Arjun are not genetically related through their parents (no consanguinity), what is the risk that Meera and Arjun’s first child will have Hurler syndrome? d. If Meera and Arjun share the same population ancestry, or received similar continental-level ancestry results from a direct-to-consumer genetic testing company, what is the risk that their first child will have the syndrome? 6. In a certain population, each of 3 serious neuromuscular conditions—autosomal dominant facioscapulohumeral muscular dystrophy, autosomal recessive Friedreich ataxia, and X-linked recessive Duchenne muscular dystrophy—has an incidence of approximately 1 in 25,000 individuals. a. What are the frequencies of pathogenic alleles for each of these conditions? b. Suppose that each condition could be treated such that affected individuals could have children. What would be the resulting effect on the incidence of each condition? Why? 7. As discussed in this chapter, the autosomal recessive condition tyrosinemia type I has an incidence of 1 in 685 individuals in one population in the province of Quebec, but approximately 1 in 100,000 elsewhere. What is the frequency of the variant associated with tyrosinemia in these two groups? Suggest possible explanations for the difference in allele frequencies between the population in Quebec and populations elsewhere.
202 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE PROBLEMS 1. A short tandem repeat (STR) variant consists of 5 different alleles, each with a frequency of 0.20 in a population. a. What pr...
Select a segment to play