🧬 Part 6: Identifying and Understanding the Genetic Basis of Disease English

← Back to Index
⚠️ For personal study use only. Commercial use is prohibited.
0%
0 / 46 listened Reset
Part 1 Part 2 Part 3 Part 4 Part 5 Part 6 Part 7 Part 8 Part 9 Part 10

Chapter 11: Identifying the Genetic Basis for Human Disease

Ch11 · Pt1 chapter 11 Identifying the Genetic Basis for Human Disease Christian R. Marshall This chapter provides an overview of how geneticists study families and populations to identify genetic contributions to disease. Whether a disease is inherited in a recognizable mendelian pattern, as illustrated in Chapter 7, or occurs at a higher frequency in relatives of affected individuals, as explored in Chapter 9, it is specific genomic variants that either cause disease directly or influence the susceptibility to disease. Genome research has provided geneticists with a catalogue of all known human genes, knowledge of their location and structure, and an ever-growing list of tens of millions of variants in DNA sequence found among individuals in different populations. As we saw in previous chapters, some of these variants are common, others are rare, and still others are ultrarare or even private to families or individuals. Whereas some variants clearly have functional consequences associated with disease risk, others are certainly neutral. For most, their significance for human health and disease is unknown. In Chapter 4, we dealt with the effect of mutation, which alters one or more genes or loci to generate variant alleles and polymorphism. In Chapters 7 and 9, we examined the role of genetic factors in the pathogenesis of various mendelian or complex disorders. In this ­chapter, we discuss how geneticists go about discovering the particular genes implicated in disease and the variants they contain that underlie or contribute to human diseases, focusing on three approaches: The first approach, linkage analysis, is family based. Linkage analysis takes explicit advantage of family pedigrees to follow the inheritance of a disease among family members and to test for consistent, repeated coinheritance of the disease with a particular genomic region or even with a specific variant(s), whenever the disease is passed on in a family. The second approach, genome-wide association analysis, is population based. Association analysis takes advantage of the entire history of a population to look for increased or decreased frequency of a particular allele or set of alleles in a cohort of affected individuals compared with a control set of unaffected individuals from that same population. It is particularly useful for complex and multifactorial diseases that do not show a mendelian inheritance pattern. The third approach involves direct genome-wide sequencing of affected individuals and their parents and/or other individuals in the family or population. Genome-wide sequencing refers to sequencing the entire genome or sequencing the coding portion of the genome, the exome. This approach has been broadly adopted by geneticists, mainly due to the advancement of nextgeneration sequencing technologies that have reduced the cost of DNA sequencing a millionfold from the original reference genome sequenced for the Human Genome Project. Genome-wide sequencing is particularly useful for rare mendelian disorders in which linkage analysis is not possible because there are not enough families to do such analysis, or because the disorder is a genetic lethal that always results from new mutations and is never inherited. Although this approach allows an unbiased analysis of genes, examination of the resulting billions (or in the case of the exome, tens of millions) of bases of DNA requires a robust filtering strategy for identification of disease alleles. Using a strategy to pare down variants coupled with the emergence of matchmaking programs has been extremely successful in gene discovery for hundreds of rare genetic disorders. Use of linkage, association, and genome-wide sequencing to identify disease-associated genes has had an enormous impact on our understanding of the pathogenesis and pathophysiology of many diseases. In time, knowledge of the genetic contributions to disease will also suggest new methods of prevention, management, and treatment. GENETIC BASIS FOR LINKAGE ANALYSIS AND ASSOCIATION A fundamental feature of human biology is that each generation reproduces by combining haploid gametes containing 23 chromosomes that resulted from independent assortment and recombination of homologous chromosomes (see Chapter 2). To understand fully the concepts underlying genetic linkage analysis and tests for association, it is necessary to review briefly the
chapter 11 Identifying the Genetic Basis for Human Disease Christian R. Marshall This chapter provides an overview of how geneticists study families and populations to identify genetic contributions t...
Ch11 · Pt2 204 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE behavior of chromosomes and genes during meiosis, as they are passed from one generation to the next. Some of this information repeats the classic material on gametogenesis presented in Chapter 2, illustrating it with new information that has become available as a result of the Human Genome Project and its applications to the study of human variation. Independent Assortment and Homologous Recombination in Meiosis During meiosis I, homologous chromosomes line up in pairs along the meiotic spindle. The paternal and maternal homologues exchange homologous segments by crossing over and creating new chromosomes that are a patchwork-like effect consisting of alternating portions of the grandparental chromosomes (see Fig. 2.15). In the family illustrated in Fig. 11.1, examples of recombined chromosomes are shown in the offspring of each generation, with the individual in generation III shown to inherit a maternal chromosome that contains segments derived from all four of his maternal grandparents’ chromosomes (generation I). The creation of such patchwork chromosomes emphasizes the notion of human genetic individuality: each chromosome inherited by a child from a parent is never exactly the same as either of the two copies of that chromosome in the parent. Homologous chromosomes differ substantially at the DNA sequence level. As discussed in Chapter 4, these differences at the same position (locus) on a pair of homologous chromosomes are alleles. Alleles that are common (generally considered to be those carried by 1% or more of the population) constitute a polymorphic locus. Allelic variants on homologous chromosomes allow geneticists to trace each segment of a chromosome inherited by a particular child to determine if and where recombination events have occurred along the homologous chromosomes. There are now hundreds of millions of genetic markers across diverse populations available to serve as genetic markers for this purpose. Alleles at Loci on Different Chromosomes Assort Independently Assume there are two polymorphic loci, 1 and 2, on different chromosomes, with alleles A and a at locus 1 and alleles B and b at locus 2 (Fig. 11.2). Suppose an I II III Figure 11.1 The effect of recombination on the origin of various portions of a chromosome. Because of crossing over in meiosis, the copy of the chromosome the boy (generation III) inherited from his mother is a mosaic of segments of all four of his grandparents’ copies of that chromosome. The blank chromosome represents the chromosome inherited from the boy’s father. A Locus 1→ b←Locus 2 B b B B a A A A a a b B A a Meiosis: homologous chromosomes line up randomly in one of two orientations or Gametes Parental combinations (AB and ab) Nonparental combinations (Ab and a B) b b a B Figure 11.2 Independent assortment of alleles at two loci, 1 and 2, when they are located on different chromosomes. Assume that alleles A and B were inherited from one parent, a and b from the other. The two chromosomes can line up on the metaphase plate in meiosis I in one of two equally likely combinations, resulting in independent assortment of the alleles on these two chromosomes.
204 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE behavior of chromosomes and genes during meiosis, as they are passed from one generation to the next. Some of this information repeats the c...
Ch11 · Pt3 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 205 individual’s genotype at these loci is Aa and Bb; that is, she is heterozygous at both loci, with alleles A and B inherited from her father and alleles a and b inherited from her mother. The two different chromosomes will line up on the metaphase plate at meiosis I in one of two combinations with equal likelihood. After recombination and chromosomal segregation are complete, there will be four possible combinations of alleles in a gamete: AB, ab, Ab, and a B. Each combination is as likely to occur as any other, a phenomenon known as independent assortment. Because AB gametes ­contain only her paternally derived alleles, and ab gametes only her maternally derived alleles, these gametes are designated parental. In contrast, Ab or a B gametes, each containing one paternally derived allele and one maternally derived allele, are termed nonparental gametes. On average, half (50%) of gametes will be parental (AB or ab) and 50% nonparental (Ab or a B). Alleles at Loci on the Same Chromosome Assort Independently If At Least One Crossover Between Them Always Occurs Now suppose that an individual is heterozygous at two loci, 1 and 2, with alleles A and B paternally derived and a and b maternally derived, but the loci are on the same chromosome (Fig. 11.3). Genes that reside on the same chromosome are said to be syntenic (literally, “on the same thread”), regardless of how close together or how far apart they lie on that chromosome. How will these alleles behave during meiosis? We know that between one and four crossovers occur between homologous chromosomes during meiosis I when there are two chromatids per homologous chromosome. If no crossing over occurs within the segment of the chromatids between the loci 1 and 2 (and ignoring whatever happens in segments outside the interval between these loci), then the chromosomes we see in the gametes will be AB and ab, which are the same as the original parental chromosomes; a parental chromosome is therefore a nonrecombinant chromosome. If crossing over occurs at least once in the segment between the loci, the resulting chromatids may be either nonrecombinant or Ab and a B, which are not the same as the parental chromosomes; such a nonparental chromosome is therefore a recombinant chromosome (shown in Fig. 11.3). One, two, or more recombinations occurring between two loci at the four-chromatid stage result in gametes that are 50% nonrecombinant (parental) and 50% recombinant (nonparental), which is precisely the same proportions one sees with independent assortment of alleles at loci on different chromosomes. Thus if two syntenic loci are sufficiently far apart on the same chromosome to ensure at least one crossover between them in every meiosis, then the ratio of recombinant to A A A a a A a A a a b b b b A A A a a A a A a a b b A A A a a a b b A A A a a a b b A A A a a a b b b b A a A a b b A a A a A a A a b b b b b 8 NR 8 R NR:R=1:1 2 AB 2 ab 0 Ab 0 a B 0 AB 0 ab 2 Ab 2 a B 1 AB 1 ab 1 Ab 1 a B 1 AB 1 ab 1 Ab 1 a B NR NR R R NR NR R R NR NR R R NR NR R R 2 NR 2 R NR:R=1:1 B B B B B b B B B b B B B b B B B b B B B B B B B B B B B Figure 11.3 Crossing over between homologous chromosomes (black horizontal lines) in meiosis is shown between chromatids of two homologous chromosomes on the left. Crossovers result in new combinations of maternally and paternally derived alleles on the recombinant chromosomes present in gametes, shown on the right. If no crossing over occurs in the interval between loci 1 and 2, only parental (nonrecombinant) allele combinations, AB and ab, occur in the offspring. If one or two crossovers occur in the interval between the loci, half the gametes will contain a nonrecombinant combination of alleles and half the recombinant combination. The same is true if more than two crossovers occur between the loci (not illustrated here). NR, Nonrecombinant; R, recombinant. nonrecombinant genotypes will be, on average, 1: 1— just as if the loci were on separate chromosomes and assorting independently.
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 205 individual’s genotype at these loci is Aa and Bb; that is, she is heterozygous at both loci, with alleles A and B inherited from her fa...
Ch11 · Pt4 206 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Recombination Frequency and Map Distance Frequency of Recombination as a Measure of Distance Between Loci Suppose now that two loci are on the same chromosome but are either far apart, very close together, or somewhere in between (Fig. 11.4). As we just saw, when the loci are far apart (see Fig. 11.4A), at least one crossover will occur in the segment of the chromosome between loci 1 and 2, and there will be gametes of both the nonrecombinant genotypes (AB and ab) and recombinant genotypes (Ab and a B), in equal proportions (on average) in the offspring. On the other hand, if two loci are so close together on the same chromosome that crossovers never occur between them, there will be no recombination; the nonrecombinant genotypes (parental chromosomes AB and ab in Fig. 11.4B) are transmitted together all of the time, and the frequency of the recombinant genotypes Ab and a B will be 0. In between these two extremes is the situation in which two loci are far enough apart that one recombination between the loci occurs in some meioses but not in others (see Fig. 11.4C). In this situation we observe nonrecombinant combinations of alleles in the offspring when no crossover occurred and recombinant combinations when a recombination has occurred. The frequency of recombinant chromosomes at the two loci will fall between 0 and 50%. The crucial point is that the closer together two loci are, the smaller the recombination frequency and the fewer recombinant genotypes in the offspring. Detecting Recombination Events Requires Heterozygosity and Knowledge of Phase Detecting the recombination events between loci requires that (1) a parent be heterozygous (informative) at both loci and (2) we know which allele at locus 1 is on the same chromosome as which allele at locus 2. In an individual who is heterozygous at two syntenic loci, one with alleles A and a, the other B and b, phase refers to which allele at the first locus is on the same chromosome with which allele at the second locus (Fig. 11.5). The set of alleles on the same homologue (A and B, or a and b) are said to be in cis (or in coupling) and form what is referred to as a haplotype. In contrast, alleles on the different homologues (A and b, or a and B) are in trans (or in repulsion) (see Fig. 11.5). A Locus 1→ Locus 2→ a B b A B a b A b a B Nonrecombinant = Recombinant (AB + ab) = (Ab + a B) A Locus 1→ Locus 2→ a B b A B a b Nonrecombinant only (AB + ab) only A Locus 1→ Locus 2→ a B b A B a b A B a B Nonrecombinant > Recombinant (AB + ab) > (Ab + a B) A B C Figure 11.4 Assortment of alleles at two loci, 1 and 2, when they are located on the same chromosome. (A) The loci are far apart and at least one crossover between them is likely to occur in every meiosis. (B) The loci are so close together that crossing over between them is not observed, regardless of the presence of crossovers elsewhere on the chromosome. (C) The loci are close together on the same chromosome but far enough apart that crossing over occurs in the interval between the two loci only in some meioses but not in most others. In coupling (cis): Aand Baandb In repulsion (trans): aand BAandb B A b a Figure 11.5 Possible phases of alleles A and a and alleles B and b.
206 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Recombination Frequency and Map Distance Frequency of Recombination as a Measure of Distance Between Loci Suppose now that two loci are on t...
Ch11 · Pt5 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 207 Fig. 11.6 shows a pedigree of a family with multiple individuals affected by autosomal dominant retinitis pigmentosa (RP), a degenerative disease of the retina that causes progressive blindness in association with abnormal retinal pigmentation. As shown, individual I-1 is heterozygous at both marker locus 1 (with alleles A and a) and marker locus 2 (with alleles B and b), as well as heterozygous for the disorder (D is the disease allele, d is the normal allele). The alleles A-D-B form one haplotype, and a-d-b the other. Because we know her spouse is homozygous at all three loci and can only pass on the a, b, and d alleles, we can easily determine which alleles the children received from their mother and thus trace the inheritance of her RP-causing allele or her normal allele at that locus, as well as the alleles at both marker loci in her children. Close inspection of Fig. 11.6 allows one to determine whether each child has inherited a recombinant or a nonrecombinant haplotype from the mother. However, if the mother (I-1) had been homozygous bb at locus 2, then all children would have inherited a maternal b allele, regardless of whether they received a mutant D or normal d allele at the RP9 locus. Because she is not informative at locus 2 in this scenario, it would be impossible to determine whether recombination had occurred. Similarly, if the information provided for the family in Fig. 11.6 was simply that individual I-1 was heterozygous, Bb, at locus 2 and heterozygous for an autosomal dominant form of RP, but the phase was not known, one could not determine which of her children were nonrecombinant between the RP9 locus and locus 2 and which were recombinant. Thus, determination of who is or is not a recombinant requires that we know whether the B or b allele at locus 2 was on the same chromosome as the mutant D allele for RP in individual I-1 (see Fig. 11.6). Linkage and Recombination Frequency Linkage is the term used to describe a departure from the independent assortment of two loci, or, in other words, the tendency for alleles at loci that are close together on the same chromosome to be transmitted together, as an intact unit, through meiosis. Analysis of linkage depends on determining the frequency of recombination as a measure of how close two loci are to each other on a chromosome. A common notation for recombination frequency (as a proportion, not a percentage) is the Greek letter theta, θ, where θ varies from 0 (no recombination at all) to 0.5 (independent assortment). If two loci are so close together that θ = 0 between them (as in Fig. 11.4B), they are said to be completely linked; if they are so far apart that θ = 0.5 (as in Fig. 11.4A), they are assorting independently and are unlinked. In between these two extremes are various degrees of linkage. Genetic Maps and Physical Maps The map distance between two loci is a theoretical concept that is based on actual data—the extent of observed recombination, θ, between the loci. Map distance is measured in units called centimorgans (c M), defined as the genetic length over which, on average, one crossover occurs in 1% of meioses. (The centimorgan is 1100 of a “morgan,” named after Thomas Hunt Morgan, who first observed genetic recombination in the fruit fly, Drosophila.) Therefore a recombination fraction of 1% (i.e., θ = 0.01) translates approximately into a map distance of 1 c M. As we discussed before in this chapter, the recombination frequency between two loci increases proportionately with the distance between two loci only up to a point, because once markers are far enough apart that at least one recombination will always occur, the observed recombination frequency will equal 50% (θ = 0.5), no matter how physically far apart the two loci are. To accurately measure true genetic map distance between two widely spaced loci, therefore, one has to use markers spaced at short genetic distances (≤1 c M) in the interval between these two loci, and then add up the values of θ between the intervening markers; the values of θ between pairs of closely neighboring markers will be good approximations of the genetic distances between them. Using this approach, the genetic length of an entire human genome has been measured and, 1 I II 2 1 2 3 4 5 6 7 8 Locus 2 RP9 Locus 1 B D A b d a b d A b d A b D A b d A B D a B D A b d a b d a b d A b d A Figure 11.6 Coinheritance of the gene for an autosomal dominant form of retinitis pigmentosa (RP), with marker locus 2 and not with marker locus 1. Only the mother’s contribution to the children’s genotypes is shown. The mother (I-1) is affected with this dominant disease and is heterozygous at the RP9 locus (Dd) as well as at loci 1 and 2. She carries the A and B alleles on the same chromosome as the mutant RP9 allele (D). The unaffected father is homozygous normal (dd) at the RP9 locus as well as at the two marker loci (AA and bb); his contributions to his offspring are not considered further. Two of the three affected offspring have inherited the B allele at locus 2 from their mother, whereas individual II-3 inherited the b allele. The five unaffected offspring have also inherited the b allele. Thus, seven of eight offspring are nonrecombinant between the RP9 locus and locus 2. However, individuals II-2, II-4, II-6, and II-8 are recombinant for RP9 and locus 1, indicating that meiotic crossover has occurred between these two loci.
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 207 Fig. 11.6 shows a pedigree of a family with multiple individuals affected by autosomal dominant retinitis pigmentosa (RP), a degenerati...
Ch11 · Pt6 208 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE interestingly, found to differ between the sexes. When measured in female meiosis, genetic length of the human genome is ~60% greater (≈4596 c M) than when it is measured in male meiosis (2868 c M), and this sex difference is consistent and uniform across each autosome. The sex-averaged genetic length of the entire haploid human genome, which is estimated to contain ~3.3 billion base pairs of DNA, or ≈3300 Mb (see Chapter 2), is 3790 c M, for an average of ~1.15 c M/Mb. Pairwise measurements of recombination between genetic markers separated by 1 Mb or more gives a fairly constant ratio of genetic distance to physical distance of ~1 c M/Mb. However, when recombination is measured at much higher resolution, such as between markers spaced less than 100 kb apart, recombination per unit length becomes nonuniform and can range over four orders of magnitude (0.01–100 c M/ Mb). When viewed on the scale of a few tens of kilobase pairs of DNA, the apparent linear relationship between physical distance in base pairs and recombination between polymorphic markers located millions of base pairs of DNA apart is, in fact, the result of an averaging of so-called hot spots of recombination interspersed among regions of little or no recombination. Hot spots occupy only ~6% of sequence in the genome and yet account for ~60% of all the meiotic recombination in the human genome. The impact of this nonuniformity of recombination at high resolution is discussed next, as we address the phenomenon of linkage disequilibrium. Linkage Disequilibrium It is generally the case that the alleles at two loci will not show any preferred phase in the population if the loci are linked, but at a distance of 0.1 to 1 c M or more. For example, suppose loci 1 and 2 are 1 c M apart. Suppose further that allele A is present on 50% of the chromosomes in a population and allele a on the other 50%, whereas at locus 2, a disease susceptibility allele S is present on 10% of chromosomes and the protective allele s is on 90% (Fig. 11.7). Because the frequency of the A-S haplotype (freq(A-S)) is simply the product of the frequencies of the two alleles—freq(A) × freq(S) = 0.5 × 0.1 = 0.05—the alleles are said to be in linkage equilibrium (see Fig. 11.7A). That is, the frequencies of the four possible haplotypes, A-S, A-s, a-S, and a-s follow directly from the allele frequencies of A, a, S, and s. However, as we examine haplotypes involving loci that are very close together, we find that knowing the allele frequencies for these loci individually does not allow us to predict the four haplotype frequencies. The frequency of any one of the haplotypes, freq(A-S) for example, may not be equal to the product of the frequencies of the individual alleles that make up that haplotype; in this situation, freq(A-S) ≠ freq(A) × freq(S), and the alleles are thus said to be in linkage disequilibrium (LD). The deviation (“delta”) between the expected and actual haplotype frequencies is called D and is given by:
D freq(-) freq( - freq - freq( - A S a s A s a s) ()) D ≠ 0 is equivalent to saying the alleles are in LD, whereas D = 0 means the alleles are in linkage equilibrium. Examples of LD are illustrated i...
Ch11 · Pt7 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 209 haplotype is present on only 1% of chromosomes in the population (see Fig. 11.7C). The A-S haplotype has a frequency much below what one would expect on the basis of the frequencies of alleles A and S in the population as a whole, and D < 0, whereas the haplotype a-S has a frequency much greater than expected and D > 0. In other words, chromosomes carrying the susceptibility allele S are enriched for allele a at the expense of allele A, compared with chromosomes that carry the protective allele s. Note, however, that the individual allele frequencies are unchanged; it is only how they are distributed into haplotypes that differ, and this is what determines if there is LD. Linkage Disequilibrium Has Both Biologic and Historical Causes What causes LD? When a pathogenic allele first enters the population (by mutation or by immigration of a founder who carries the altered allele), the particular set of alleles at polymorphic loci linked to the disease locus constitutes a disease-associated haplotype (Fig. 11.8). The degree to which this original disease-associated haplotype will persist over time depends in part on the probability that recombination removes the diseaseassociated allele from the original haplotype and onto chromosomes with different sets of alleles at these linked loci. The speed with which recombination will move the pathogenic allele onto a new haplotype depends on a number of factors: The number of generations (and therefore the number of opportunities for recombination) since the mutation first appeared. The frequency of recombination per generation between the loci. The smaller the value of θ, the greater is the chance that the disease-associated haplotype will persist intact. Processes of natural selection for or against particular haplotypes. If a haplotype combination undergoes either positive selection (and is preferentially passed on) or experiences negative selection (and is less readily passed on), it will be either over- or under-represented in that population. Measuring Linkage Disequilibrium Although conceptually valuable, the discrepancy, D, between the expected and observed frequencies of haplotypes is not a good way to quantify LD because it varies not only with degree of LD but also with the allele frequencies themselves. To quantify varying degrees of LD, therefore, geneticists often use a measure derived from D, referred to as D′ (see Box 11.1). D′ is designed to vary from 0, indicating linkage equilibrium, to a maximum of ±1, indicating very strong LD. LD is a result, not only of genetic distance, but of the amount of time A B Fragmentation of original chromosome by recombination as population expands through multiple generations Pathogenic variant located within region of linkage disequilibrium Pathogenic variant on founder chromosome Figure 11.8 (A) With each generation, meiotic recombination exchanges the alleles that were initially present at polymorphic loci on a chromosome on which a disease-associated (pathogenic) variant arose () for other alleles present on the homologous chromosome. Over many generations, the only alleles that remain in coupling phase with the variant are those at loci so close to the disease associated locus that recombination between the loci is very rare. These alleles are in linkage disequilibrium with the pathogenic variant and constitute a disease-associated haplotype. (B) Affected individuals in the current generation (arrows) carry the pathogenic variant (X) in linkage disequilibrium with the disease-associated haplotype (individuals in blue). Depending on the age of the pathogenic variant and other population genetic factors, a disease-associated haplotype ordinarily spans a region of DNA of a few kb to a few hundred kb. (Modified from original figures of Thomas Hudson, Mc Gill University, Canada.)
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 209 haplotype is present on only 1% of chromosomes in the population (see Fig. 11.7C). The A-S haplotype has a frequency much below what on...
Ch11 · Pt8 210 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE during which recombination had a chance to occur and the possible effects of selection for or against particular haplotypes. Different populations, therefore, living in different environments and with different histories can have different values of D′ between the same two alleles at the same locus in the genome. shorter haplotypes more rapidly than average, resulting in linkage equilibrium between SNPs on one side and the other side of the hot spot. The correlation is by no means exact, and many apparent boundaries between LD blocks are not located over evident recombination hot spots. This lack of perfect correlation should not be surprising, given what we have already surmised about LD: it is affected not only by how likely a recombination event is (i.e., where the hot spots are) but also by the age of the population, the frequency of the haplotypes originally present in the founding members of that population, and whether there has been either positive or negative selection for particular haplotypes. STRATEGIES FOR DISCOVERY OF DISEASEASSOCIATED GENES In clinical medicine, a disease state is defined by a collection of phenotypic findings seen in a patient or group of patients. Designating such a disease as “genetic”—inferring the existence of a gene whose alteration is responsible for or contributes to the disease—comes from detailed genetic analysis, applying the principles outlined in Chapters 7 and 9. However, surmising the existence of a gene or genes in such a way does not tell us which of the ~20,000 coding and ~18,000 noncoding genes in the genome is involved, what its function might be, or how it causes or contributes to the disease. Strategies for discovery of genes associated with human disease have evolved over the years, from gene mapping to genome-wide sequencing. Combining these two approaches has provided an effective strategy. Mapping of such genes has historically been a critical and necessary first step in identifying the gene(s) in which certain variants are responsible for causing or increasing susceptibility to disease. Mapping the gene focuses attention on a region of the genome in which to carry out a systematic analysis of all the genes in that region, to identify variation that contributes to the disease. The marked fall in cost of DNA sequencing over the last decade has made it feasible to take a genomewide sequencing approach to gene discovery. Sequencing the genomes (or just the coding portion of the genome, the exome) of cohorts with similar phenotypes, followed by systematic filtering, has proven a powerful approach for gene discovery. This is especially true for disorders with a new (de novo) dominant mechanism that may be intractable to gene mapping. Incorporation of gene mapping information into the filtering strategy of genomewide sequencing has also proven an effective approach to narrow in on loci that are associated with disease. Regardless of strategy, identification of the gene that harbors the DNA variants responsible for either causing a mendelian disorder or increasing susceptibility to a genetically complex disease, allows the full spectrum of variation in that gene to be studied. We can determine the degree of allelic heterogeneity, the penetrance of BOX 11.1 MEASURING LINKAGE DISEQUILIBRIUM D′ = D/F where D freq(-)freq(-) freq(-)freq(-)
A S a s A s a S and F is a correction factor that helps account for the allele frequencies. The value of F depends on whether D itself is a positive or negative number. F = the smaller of freq(A) × fr...
Ch11 · Pt9 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 211 G A A G G T C C T T C T C C T C T C C C T G T T T C T C C C C T T T T G T G G G A T A A A 40% 30% 11% 9% 8% 98% Haplotype frequency 40% 60% 100% Allele frequency 1 2 3 4 5 6 7 8 9 C T C G G T T T C A A T 42% 31% 26% 99% Haplotype frequency G A 10 CLUSTER 1 CLUSTER 2 1 SNP# 2 3 4 5 6 7 8 9 10 11 14 13 12 50 kb 100 kb 150 kb 11 12 13 14 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 0.9 1.0 1.0 1.0 1.0 1.0 1.0 1.0 0.8 1.0 1.0 1 2 3 4 5 6 7 8 9 10 11 12 13 1.0 14 1.0 0 50 100 100 50 1.0 c M/Mb Mb 89 kb 14 kb 42 kb 1.0 0.8 1.0 1.0 0.6 1.0 1.0 0.9 0.9 1.0 A B C Figure 11.9 (A) A 145-kb region of chromosome 4 containing 14 single nucleotide polymorphic loci (SNPs). In cluster 1, containing SNPs 1 through 9, five of the 29 = 512 theoretically possible haplotypes are responsible for 98% of all the haplotypes in the population, reflecting substantial linkage disequilibrium (LD) among these SNP loci. Similarly, in cluster 2, only three of the 24 = 16 theoretically possible haplotypes involving SNPs 11 to 14 represent 99% of all the haplotypes found. In contrast, alleles at SNP 10 are found in linkage equilibrium with the SNPs in cluster 1 and cluster 2. (B) A schematic diagram in which each red box contains the pairwise measurement of the degree of LD between two SNPs (e.g., the arrow points to the box, outlined in black, containing the value of D′ for SNPs 2 and 7). The higher the degree of LD, the darker the color in the box, with maximum D′ values of 1.0 occurring when there is complete LD. Two LD blocks are detectable, the first containing SNPs 1 through 9, and the second SNPs 11 through 14. Between blocks, the 14-kb region containing SNP 10 shows no LD with neighboring SNPs 9 or 11 or with any of the other SNP loci. (C) A graph of the ratio of map distance to physical distance (c M/Mb), showing that a recombination hot spot is present in the region between SNP 10 and cluster 2, with values of recombination that are 50- to 60-fold above the average of ~1.15 c M/Mb for the genome. (Based on data and diagrams provided by Thomas Hudson, Quebec Genome Center, Montreal, Canada.)
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 211 G A A G G T C C T T C T C C T C T C C C T G T T T C T C C C C T T T T G T G G G A T A A A 40% 30% 11% 9% 8% 98% Haplotype frequency 40%...
Ch11 · Pt10 212 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE different alleles, whether there is a correlation between certain alleles and various aspects of the phenotype (genotype-phenotype correlation), and the frequency of disease-causing or predisposing variants in various populations. Those with the same or similar disorders can also be examined to determine locus heterogeneity. Once the gene and variants in that gene are identified in affected individuals, highly specific methods of diagnosis— including prenatal diagnosis and carrier screening (see Chapter 18)—can be offered to patients and their families. The variants associated with disease can then be modeled in other organisms, which allows us to use powerful genetic, biochemical, and physiologic tools to better understand the disease pathogenesis. Finally, armed with an understanding of gene function and how disease-alleles affect that function, we can begin to develop specific therapies to prevent or ameliorate the disorder (see Chapter 14). Indeed, much of the material in the next few chapters about the etiology, pathogenesis, mechanism, and treatment of various diseases begins with identification of the genes involved. Here, we examine the major approaches used to discover these genes, as outlined at the beginning of this chapter. MAPPING HUMAN DISEASE GENES BY LINKAGE ANALYSIS Determining Whether Two Loci Are Linked Linkage analysis is a method of mapping genes that uses studies of recombination in families to determine whether two genes show linkage when passed from one generation to the next. We use information from the known or suspected mendelian inheritance pattern (dominant, recessive, X-linked) to determine which family members have inherited a recombinant or a nonrecombinant chromosome. To decide whether two loci are linked and, if so, how close or far apart they are, we rely on two pieces of information. First, using the family data in hand, we estimate θ, the recombination frequency between the two loci. Next, we ascertain whether θ is statistically significantly different from 0.5, which is the fraction expected for unlinked loci. Estimating θ and, at the same time, determining the statistical significance of any deviation of θ from 0.5, relies on a statistical tool called the likelihood ratio (as discussed later in the chapter). Linkage analysis begins with a set of actual family data with N individuals. Based on a mendelian inheritance model, count the number of chromosomes, r, that show recombination between the allele causing the disease and informative alleles at various polymorphic loci around the genome (so-called markers). The number of chromosomes that do not show recombination is therefore N − r. With each meiosis, the recombination fraction, θ, is the unknown probability that a recombination will occur between the two loci; the probability that no recombination occurs is, therefore, 1 − θ. Because each meiosis is an independent event, one multiplies the probability of a recombination, θ, or of no recombination, (1 − θ), for each chromosome. The formula for the likelihood (probability) of observing this number of recombinant and nonrecombinant chromosomes when θ is unknown is {N!/r!(N − r)!}θr (1 − θ)(N−r). (The factorial term, N!/r!(N − r)!, accounts for all the possible birth orders in which the recombinant and nonrecombinant children can appear in the pedigree.) Calculate a second likelihood based on the null hypothesis that the two loci are unlinked, i.e., that θ = 0.50. The ratio of the likelihood of the family data supporting linkage with unknown θ to the likelihood that the loci are unlinked is the odds in favor of linkage and is given by:
Likelihood of the data if loci are linked at distance Likelihood of t he data if loci are unlinked( 0.5) {N!/r!(N r)!} (1) { r N r () N!/r!(N r)!} r N r  12 12 () Fortunately, the factorial ter...
Ch11 · Pt11 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 213 alleles at locus 1 or alleles at any of the other hundreds of marker loci tested on the other autosomes. Thus, although the RP locus involved in this family could, in principle, have mapped anywhere in the human genome the linkage data suggest that the responsible RP locus lies in the region of chromosome 7 near marker locus 2. To provide a quantitative assessment of this suspicion, suppose we let θ be the “true” recombination fraction between RP and locus 2—the fraction we would see if we had unlimited offspring to test. The likelihood ratio for this family is
()(1) ( ( 1 7 1 7 12 12)) and reaches a maximum LOD score of Zmax = 1.1 at θmax = 0.125. The value of θ that maximizes the likelihood ratio, θmax, may be the best estimate for θ given the data, but h...
Ch11 · Pt12 214 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE D-A and d-a the other half (which assumes the alleles in these haplotypes are in linkage equilibrium). To calculate the overall likelihood of this pedigree, we then add the likelihood calculated assuming one phase in the mother to that assuming the other phase. The overall likelihood = 12 0 12
() ()() 1 1 3 3 0 ; the likelihood ratio for this pedigree, then, is: 12 12 18 () ()() () 1 1 3 0 3 0 giving a maximum LOD score of Zmax = 0.602 at θmax = 0. If, however, additional genotype informa...
Ch11 · Pt13 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 215 Alternatively, if the association study was designed as a cross-sectional or cohort study, the strength of an association can be measured by the relative risk (RR). The RR is the ratio of the proportion of those with the disease who carry a particular marker allele ([a/(a + b)]) to the proportion of those without the disease who carry that marker ([c/(c + d)]). RR a (a b c (c d)
) Again, an RR that differs from 1 means there is an association of disease with the genetic marker, whereas RR = 1 means there is no association. (The RR introduced here should not be confused with R...
Ch11 · Pt14 216 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE To illustrate these approaches we first consider a case-control study of cerebral vein thrombosis (CVT), which we introduced in Chapter 9. In this study, suppose a group of 120 individuals with CVT and 120 matched controls were genotyped for the 20210 G>A allele in the prothrombin gene (see Chapter 9). Cases With CVT Controls Without CVT Totals 20210 G>A allele present 23 4 27 20210 G>A allele absent 97 116 213 Total 120 120 240 CVT, Cerebral vein thrombosis. Because this is a case-control study, we will calculate an odds ratio: OR = (23/4)/(97/116) = ≈6.9 with 95% confidence limits of 2.3 to 20.6. The effect size of 6.9 is substantial, and 95% confidence limits exclude 1.0, thereby demonstrating a strong and statistically significant association between the 20210 G>A allele and CVT. Stated simply, individuals carrying the prothrombin 20210 G>A allele have nearly seven times greater odds of having the disease than do those who do not carry this allele. To illustrate a longitudinal cohort study, calculating RR instead of OR—consider statin-induced myopathy, a rare but well-recognized adverse drug reaction that can develop in some individuals during statin therapy to lower cholesterol. In one study, subjects enrolled in a cardiac protection study were randomized to receive 40 mg of the statin drug, simvastatin, or placebo. Over 16,600 participants exposed to the statin were genotyped for a variant (Val 174Ala) in the SLCO1B1 gene— which encodes a hepatic drug transporter, and were watched for development of the adverse drug response. Out of the entire genotyped group exposed to the statin, 21 developed myopathy. Examination of their genotypes showed that the RR for developing myopathy associated with the presence of the Val 174Ala allele was ~2.6, with 95% confidence limits of 1.3 to 5.1. Thus, there is a statistically significant association between the Val 174Ala allele and statin-induced myopathy. Those carrying this allele are at moderately increased risk for developing this adverse drug reaction, relative to those who do not carry this allele. One common misconception concerning an association study is that the more significant the p-value, the stronger is the association. In fact, a significant p-value for an association does not provide information concerning the magnitude of the effect of an associated allele on disease susceptibility. Significance is a statistical measure that describes how likely it is that the population sample used for the association study could have yielded an observed OR or RR that differs from 1.0, simply by chance. In contrast, the actual magnitude of the OR or RR—how far it diverges from 1.0—is a measure of the impact a particular variant (or genotype or haplotype) on increasing or decreasing disease likelihood. Genome-Wide Association Studies The Haplotype Map (Hap Map) Association studies for human disease genes were once limited to particular sets of variants in restricted sets of genes. These were chosen, either for convenience or because they were thought to be involved in a pathophysiologic pathway relevant to a disease, making them logical candidate genes for the disease under investigation. Many such association studies were undertaken before the Human Genome Project era, using HLA or blood group loci, for example, because these were highly polymorphic and easily genotyped in case-control studies. Ideally, however, one would like to test systematically for an association between any disease of interest and every one of the tens of millions of rare and common alleles in the genome, in an unbiased fashion without preconception of what genes and genetic variants might be contributing to the disease. Association analyses on a genome scale are referred to as genome-wide association studies (GWAS). Such an undertaking for all known variants is impractical for many reasons. It can, however, be approximated by genotyping cases and controls for a mere 300,000 to 1 million individual variants located throughout the genome, to search for association with the disease or trait in question. The success of this approach depends on exploiting LD: as long as a variant responsible for altering disease susceptibility is in LD with one or more of the genotyped variants within an LD block, a positive association should be detectable between that disease and the alleles in the LD block. Developing such a set of markers led to the launch of the Haplotype Mapping (Hap Map) Project, one of the biggest human genomics efforts to follow completion of the Human Genome Project. The Hap Map Project began in four geographically distinct groups—a primarily European population, a West African population, a Han Chinese population, and a population from Japan—and included collecting and characterizing millions of SNP loci and developing methods to genotype them rapidly and inexpensively. Hap Map version 3 expanded coverage and diversity to include genotyping of 1.6 million common SNP loci as well as common CNVs in more than 1000 reference individuals from 11 global populations. Subsequently, whole genome sequencing has been applied to many populations in what is referred to as the 1000 Genomes Project, resulting in a massive expansion in the database of DNA variants available for GWAS among different populations around the globe. Gene Mapping by Genome-Wide Association Studies The purpose of the Hap Map was not just to gather basic information about the distribution of LD across
216 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE To illustrate these approaches we first consider a case-control study of cerebral vein thrombosis (CVT), which we introduced in Chapter 9. I...
Ch11 · Pt15 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 217 the human genome. Its primary purpose was to provide a powerful new tool for finding the genetic variants that contribute to human disease and other traits, by making possible an approximation to an idealized, fullscale, genome-wide association. The driving principle behind this approach is straightforward: detecting an association with alleles within an LD block pinpoints the genomic region within the block as likely to contain the disease-associated allele. Consequently, although the approach does not typically pinpoint the actual variant responsible functionally for the disease association, this region will be the place to focus additional studies to find the allelic variant(s) directly involved in the disease process. Historically, detailed analysis of conditions associated with high-density variants in the class I and class II HLA regions has exemplified this approach (see Box 11.3). However, with the tens of ­millions of variants now available in different populations, this approach can be broadened to examine the genetic basis of virtually any complex disease or trait. Indeed, to date, thousands of GWAS have uncovered an enormous number of naturally occurring variants associated with a variety of genetically common and complex multifactorial diseases. These range from diabetes and inflammatory bowel disease to rheumatoid arthritis and neuropsychiatric disease, and include traits such as stature and pigmentation. Research to uncover the underlying biologic basis for these associations will be ongoing for years to come. Finding the Genes Contributing to a Complex Disease by Genome-Wide Association: Age-Related Macular Degeneration Genome-wide association has proven effective in identifying hundreds of genes and alleles associated with genetically complex disorders. The power of these approaches has increased enormously with the introduction of highly efficient and less expensive technologies for genome analysis. Here we describe an example of using GWAS to find multiple allelic variants in genes that increase susceptibility to age-related macular degeneration (AMD) (Case 3), a devastating disorder that robs older adults of their vision. AMD is a progressive degenerative disease of the portion of the retina responsible for central vision. It causes blindness in 1.75 million Americans older than 50 years. The disease is characterized by the presence of drusen, which are clinically visible, discrete extracellular deposits of protein and lipids behind the retina in the region of the macula (Case 3). Although there is ample evidence for a genetic contribution to the disease, most individuals with AMD are not in families with a likely mendelian pattern of inheritance. Environmental contributions are also important, as shown by the increased risk for AMD in cigarette smokers compared with nonsmokers. Initial case-control GWAS of AMD revealed association of two common SNP loci near the complement factor H (CFH) gene. The most frequent at-risk haplotype containing these alleles was seen in 50% of cases versus only 29% of controls (OR = 2.46; 95% confidence interval [CI], 1.9–53.11). Homozygosity for this haplotype was found in 24.2% of cases, compared to only 8.3% of the controls (OR = 3.51; 95% CI, 2.13 − 5.78). A search through the SNPs within the LD block containing the AMD-associated haplotype revealed a nonsynonymous SNP in the CFH gene that substituted a histidine for tyrosine at position 402 of the CFH BOX 11.3 HUMAN LEUKOCYTE ANTIGEN AND DISEASE ASSOCIATION Among more than 1000 genome-trait or genome-disease associations from around the genome, the region with the highest concentration of associations to different phenotypes is the human leukocyte antigen (HLA) region. In addition to the association of specific alleles and haplotypes to type 1 diabetes discussed in Chapter 9, association of various HLA polymorphisms has been demonstrated for a wide range of conditions. Most, but not all of these are autoimmune; that is, associated with an abnormal immune response apparently directed against one or more self-antigens. These associations are thought to be related to variation in the immune response resulting from polymorphism in immune response genes. The functional basis of most HLA-disease associations is unknown. HLA molecules are integral to T-cell recognition of antigens. Different HLA alleles are thought to result in structural variation in these cell surface molecules, leading to differences in capacity of the proteins to interact with antigen and the T-cell receptor in the initiation of an immune response. This affects such critical processes as immunity against infections and self-tolerance to prevent autoimmunity. Ankylosing spondylitis, a chronic inflammatory disease of the spine and sacroiliac joints, is one example. More than 95% of those with ankylosing spondylitis are HLA-B27 positive; the risk for developing ankylosing spondylitis is at least 150 times higher for people who have certain HLA-B27 alleles than for those who do not. These alleles lead to HLA-B27 heavy chain misfolding and inefficient antigen presentation. In other disorders, the association between a particular HLA allele or haplotype and a disease is not due to functional differences in immune response genes themselves. Instead, the association is due to a particular allele being present at a very high frequency on chromosomes that also happen to contain disease-causing variants in another gene within the major histocompatibility complex region. One example is hemochromatosis (Case 20), a common disorder of iron overload. More than 80% of individuals with hemochromatosis are homozygous for a common variant, Cys 282Tyr, in the hemochromatosis gene (HFE) and have HLA-A*0301 alleles at their HLA-A locus. The association is not the result of HLA-A*0301, however. HFE is involved with iron transport or metabolism in the intestine; HLA-A, as a class I immune response gene, has no effect on iron transport. The association is due to proximity of the two loci and LD between the Cys 282Tyr HFE mutation and the A*0301 allele at HLA-A.
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 217 the human genome. Its primary purpose was to provide a powerful new tool for finding the genetic variants that contribute to human dise...
Ch11 · Pt16 218 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE protein (Tyr 402His). The Tyr 402His alteration, which has an allele frequency of 26 to 29% in European and African populations, showed an even stronger association with AMD than did the two SNPs that showed an association in the original GWAS. Given that drusen contain complement factors and that CFH is found in retinal tissues around drusen, it is believed that the Tyr 402His variant is less protective against the inflammation that is thought to be responsible for drusen formation and retinal damage. Thus, Tyr 402His is likely to be the variant at the CFH locus responsible for increasing the risk for AMD. More recent GWAS of AMD, using more than 7600 cases and more than 50,000 controls and millions of variants genome wide, have revealed that alleles at a minimum of 19 loci are associated with AMD, with genome-wide significance of P < 5 × 10−8. A popular way to summarize GWAS in graphic form is to plot the −log 10 significance levels for each associated variant in a Manhattan plot (so named because it is thought to bear a somewhat fanciful similarity to the skyline of New York City) (Fig. 11.11). The ORs for AMD of these variants range from a high of 2.76 for a gene of unknown function, ARMS2, and 2.48 for CFH to 1.1 for many other genes involved in multiple pathways, including the complement system, atherosclerosis, blood vessel formation, and others. In this example of AMD, a complex disease, GWAS led to the identification of strongly associated common SNPs that in turn were in LD with a common coding SNP in the gene that appears to be the functional variant involved in the disease. This discovery in turn led to the identification of other SNPs in the complement cascade and elsewhere that can also predispose to or protect against the disease. Taken together, these results give important clues to the pathogenesis of AMD and suggest that the complement pathway might be a fruitful target for novel therapies. Equally interesting is that GWAS revealed that a novel gene of unknown function, ARMS2, is also involved, thereby opening up an entirely new line of research into the pathogenesis of AMD. Gene Mapping by Analysis of Copy Number Variation Association studies have been successful in uncovering both common and rare risk alleles for common neuropsychiatric disorders. The Psychiatric Genomics Consortium (PGC) is a large international consortium that promotes global collaboration for the study of 11 psychiatric disorders: ADHD, Alzheimer disease, autism, bipolar disorder, eating disorders, major depressive disorder, obsessive-compulsive disorder/Tourette syndrome, posttraumatic stress disorder, schizophrenia, substance use disorders, and all other anxiety disorders. For complex brain disorders such as these, it has become clear that extremely large numbers of cases and controls (i.e., >10,000 samples) are necessary for markers to reach statistical significance; these numbers are well beyond what single study analysis can achieve. The PGC GWAS group aims to conduct rigorous large-scale-analysis for 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 0 5 10 15 100 200 300 400 CFH CF1 C2-CFB ARMS2-HTRA1 B3GALTL RAD51B UPC CETP C3 APOE TIMP3 SLC16A8 VEGFA TNFRSF10A COL15A1-TGFBR1 FRK-COL10A1 IER3-DDR1 ADAMS9 COL8A1 –log 10P Chromosome Figure 11.11 Manhattan plot of genome-wide association studies (GWAS) of age-related macular degeneration using ~1 million genomewide single nucleotide polymorphism (SNP) alleles located along all 22 autosomes on the x-axis. Each blue dot represents the statistical significance [expressed as −log 10(P) plotted on the y-axis], confirming a previously known association; green dots are the statistical significance for novel associations. The discontinuity in the y-axis is needed because some of the associations have extremely small P values < 1 × 10−16. (From Fritsche LG, Chen W, Schu M, et al: Seven new loci associated with age-related macular degeneration. Nature Genet 17:1783–1786, 2013.)
218 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE protein (Tyr 402His). The Tyr 402His alteration, which has an allele frequency of 26 to 29% in European and African populations, showed an e...
Ch11 · Pt17 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 219 psychiatric disorders (i.e., to gather data from multiple studies and platforms into one large dataset and perform GWAS). This approach has proven effective in increasing the number of associated loci recognized for disorders (e.g., for schizophrenia, the number of associated loci increased from 22 using 20,000 subjects to 108 using 150,000 subjects). One of the beneficial by-products of genotyping so many individuals, typically performed with SNP microarray, is the ability to interrogate the data for copy number variation (CNV). Rare CNVs are known to cause several neuropsychiatric conditions, some with high penetrance. Since highly penetrant CNVs are individually rare and tend to have nonrecurrent breakpoints, large cohorts are needed to show statistical association and precisely map regions and genes. For example, in Fig. 11.12 we can see a CNV map overlapping the NRXN1 gene on chromosome 2 in 20,000 cases and controls. This method has been used to find dozens of genomic loci or specific genes associated with neuropsychiatric disorders. Pitfalls in Design and Analysis of GWAS Association methods are powerful tools for pinpointing precisely the genes that contribute to genetic disease, by demonstrating not only the genes, but also the particular alleles responsible. These methods are relatively easy to perform because one needs samples only from a set of unrelated affected individuals and controls, not laborious collection of samples from many members of a pedigree. Association studies must be interpreted with caution, however. One serious limitation of association studies is the problem of totally artifactual association caused by population stratification (see Chapter 10). If a population is stratified into separate subpopulations (e.g., by ethnicity or religion) and members of one subpopulation rarely mate with members of other subpopulations, then a disease that happens to be more common in one subpopulation can appear (incorrectly) to be associated with any alleles that also happen to be more common in that subpopulation than in the population as a whole. Factitious association due to population stratification can be minimized, however, by careful selection of matched controls. In particular, one form of quality control is to make sure the cases and controls have similar frequencies of alleles whose frequencies differ markedly between populations (ancestry informative markers as we discussed in Chapter 10). If the frequencies seen in cases and controls are similar, then unsuspected or cryptic stratification is unlikely. In addition to the problem of stratification producing false-positive associations, false-positive results in GWAS can arise if an inappropriately lax test for statistical significance is applied. This is because, as the number of alleles being tested for a disease association Position (hg 18) Genes Breakpoint association: –log(z P) 4 2 0 Case CNVs Control CNVs 50M NRXN1INM_004801 NRXN1INM_001135659 NRXN1INM_138735 51M Figure 11.12 Copy number breakpoint mapping at the NRXN1 gene locus in cases diagnosed with schizophrenia. Manhattan plot of copy number variation breakpoint associations at the NRXN1 locus in cases with schizophrenia versus controls. The three isoforms of NRXN1 are shown in pink. Rare copy number deletions (red bars) and duplications (blue) are mapped in schizophrenia cases (n = 21,094) and population controls (n = 20,227).
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 219 psychiatric disorders (i.e., to gather data from multiple studies and platforms into one large dataset and perform GWAS). This approach...
Ch11 · Pt18 220 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE increases, the risk of finding associations by chance alone also increases—a concept in statistics known as the problem of multiple hypothesis testing. To understand why the cutoff for statistical significance must be much more stringent when multiple hypotheses are being tested, imagine flipping a coin 50 times and having it come up heads 40 times. Such a result has a probability of occurring only once in ~100,000 times. However, if the same experiment were repeated a million times, chances are greater than 99.999% that at least one coin flip experiment out of the million performed will result in 40 or more heads! Thus, even rare events that occur by chance alone in an experiment become frequent when the experiment is repeated over and over again. This is why, when testing for an association with hundreds of thousands to millions of variants across the genome, tens of thousands of variants could appear associated with P < 0.05 by chance alone. This makes a typical cutoff for statistical significance of P < 0.05 far too low to point to a true association. Instead, a significance level of P < 5 × 10−8 is considered to be more appropriate for GWAS that tests hundreds of thousands to millions of variants. Even with appropriately stringent cutoffs for genome-wide significance, however, false-positive results due to chance alone will still occur. To take this into account, a properly performed GWAS usually includes a replication study in a different, completely independent group of individuals to show that alleles near the same locus are associated. A caveat, however, is that alleles that show association may be different in different ancestral groups. Finally, it is important to emphasize that if an association is found between a disease and a marker allele that is part of a dense haplotype map, one cannot infer a functional role for that marker allele in increasing disease susceptibility. Because of the nature of LD, all alleles in LD with an allele at a locus involved in the disease will show apparently positive association, whether or not they have any functional relevance in disease predisposition. An association based on LD is still quite useful, however; for the marker alleles to appear associated, they likely sit within an LD block that also harbors the actual disease locus. Importance of Associations Discovered with GWAS There is vigorous debate regarding the interpretation of GWAS results and their value as a tool for human genetic studies. The debate arises primarily from a misunderstanding of what an OR or RR means. Many properly executed GWAS yield significant associations, but of very modest effect size (similar to the OR of 1.1 just mentioned for AMD). In fact, significant associations of smaller and smaller effect size have become more common as larger and larger sample sizes are used. This has led to the suggestion that GWAS are of little value because the effect size of the association, as measured by OR or RR, is too small to implicate the gene and pathway identified by that variant in the pathogenesis of the disease. This is faulty reasoning on two accounts. First, ORs are a measure of the impact of a specific allele (e.g., the CFH Tyr 402His allele for AMD) on complex pathogenetic pathways, such as the alternative complement pathway, of which CFH is a component. The subtlety of that impact is determined by how that allele perturbs the biologic function of the gene in which it is located, not by whether the gene harboring that allele might be important in disease pathogenesis. Studies, for example, of individuals with different autoimmune disorders, such as rheumatoid arthritis, systemic lupus erythematosus, and Crohn disease, reveal modest associations, but with some of the same variants. This suggests common pathways leading to these distinct but related diseases, an observation that may illuminate their pathogenesis. Second, even if the effect size of any one variant is small, GWAS demonstrate that many of these disorders are indeed extremely polygenic, even more so than previously suspected. Thousands of variants, most of which individually contribute little to disease likelihood (ORs between 1.01 and 1.1) in aggregate, account for a substantial fraction of clustering of these diseases within certain families (see Chapter 9). Indeed genotypes from all of the variants can be combined into a number called a polygenic risk score (PRS) (see Chapter 9). The PRS reflects a person’s inherited susceptibility to a disease, and the effect size can be as high as with some rare penetrant variants. Most alleles found by GWAS may indeed have modest effect size, but there is a critical and perhaps most fundamental finding of GWAS: the genetic architecture of some of the most common complex diseases may involve hundreds to thousands of loci harboring variants of small effect in many genes and pathways. These genes and pathways are important to our understanding of how complex diseases occur, even if each allele exerts but subtle effects on gene regulation or protein function and on disease susceptibility. It is also important to note the potential interplay between common and rare variants in disease susceptibility. This has been studied in depth in schizophrenia, with the identification of both rare variants with high OR and common variants of low OR contributing to risk (Fig. 11.13). An overall risk profile that combines a PRS with rare variants will perhaps be come available for many disorders. Thus, GWAS remains an important human genetics research tool for dissecting the many contributions to complex disease, regardless of whether the individual variants associated substantially raise the likelihood of the disease in individuals who carry them (see Chapter 17). There are more than 300,000 mapped associations in the genome (https://www.ebi.ac.uk/gwas/home). We
220 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE increases, the risk of finding associations by chance alone also increases—a concept in statistics known as the problem of multiple hypothes...
Ch11 · Pt19 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 221 anticipate many more genetic variants responsible for complex diseases to be identified by genome-wide association and that deep sequencing of such regions will uncover the variants or collections of variants functionally responsible for the disease. Such findings should provide powerful insights and potential therapeutic targets for many of the common diseases that cause so much morbidity and mortality. FINDING GENES RESPONSIBLE FOR DISEASE BY GENOME-WIDE SEQUENCING Thus far in this chapter we have focused on two approaches to map and then identify genes involved in disease: linkage analysis and GWAS. Now we turn to a third approach, involving direct genome-wide sequencing of affected individuals, along with their parents and/ or other family members, or a population cohort with the same clinical diagnosis. Characteristics, strengths, and weaknesses of linkage, association, and genomewide sequencing methods for disease gene identification are summarized in Box 11.4. The development of vastly improved and highthroughput methods of DNA sequencing has cut the cost of sequencing by six orders of magnitude from that spent for the Human Genome Project’s reference sequence. This has opened new possibilities for discovering the genes and variants responsible for diseases, particularly for rare mendelian disorders. As introduced in Chapter 4, these new technologies make it possible to generate a whole genome sequence (GS) or, in what is a cost-effective compromise, sequence for the less than 2% of the genome containing the exons of genes, referred to as a whole exome sequence or exome sequencing (ES). Comparison of Exome and Whole Genome Sequencing Both exome and genome sequencing fall into the category of genome-wide sequencing (i.e., taking an unbiased approach to interrogating a genome). ES involves sequencing of the coding portion of the genome, where, during the library preparation step, gene exons are targeted or captured. There are several commercially available kits available for ES and all involve either a hybridization or PCR step to enrich for the exonic portion of the genome. Further customization is often possible to boost and thus provide better coverage for certain variant types (e.g., mitochondrial variants) or clinically relevant regions (e.g., known pathogenic variants). The application of ES has been instrumental in the discovery of genes responsible for rare mendelian disorders, given its cost effectiveness, allowing the sequencing of more samples compared to more costly GS. However, there are both design and technical limitations in using ES, including the inability to analyze noncoding changes (e.g., deep intronic or regulatory regions) unless specifically targeted, the dropout of some coding sequence due to capture inefficiencies (high GC content exons), and limited ability to resolve more complex genetic mechanisms (e.g., structural rearrangements, repeat expansions). With the cost of sequencing continuing to drop it is becoming more feasible to use GS for gene discovery. The application of GS can address many of the technical limitations of ES (see Box 11.5). First, since there is no capture process, GS is not prone to dropout of more complex or high GC regions and, thus, provides better coverage of the coding regions of the genomes. Second, noncoding regions of the genome are 22q11.2 del 20 1.0 0.0001 0.01 Allele frequency in population Penetrance/OR 0.1 0.5 1q21 del 3q29 del CNVs with Low frequency with high OR SNPs with High frequency with low OR 16q11.2 dup Figure 11.13 Common and rare variant allele spectrum of schizophrenia. Penetrance in the form of an odds ratio (OR) is shown on the y-axis with allele frequency on the y-axis. Both rare and common variation contribute to schizophrenia risk with rare CNVs acting with OR of 10 to 20. Common SNPs associated with schizophrenia have OR of ~1 to 1.5. (From Sullivan PF, Daly MJ, O’Donovan M. Genetic architectures of psychiatric disorders: the emerging picture and its implications. Nat Rev Genet 13(8):537-551, 2012. https://doi. org/10.1038/nrg 3240.)
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 221 anticipate many more genetic variants responsible for complex diseases to be identified by genome-wide association and that deep sequen...
Ch11 · Pt20 222 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE BOX 11.4 METHODS OF DISCOVERY: COMPARISON OF LINKAGE, ASSOCIATION METHODS, AND GENOME-WIDE SEQUENCING Linkage Association Genome-Wide Sequencing Follows inheritance of a disease trait and regions of the genome from individual to individual in family pedigrees Looks for regions of the genome harboring disease alleles; uses polymorphic loci to mark which region an individual has inherited from which parent Uses hundreds to thousands of informative markers across the genome Not designed to find the specific variant responsible for or predisposing to the disease; can only demarcate where the variant can be found within (usually) one or a few megabases Relies on recombination events occurring in families during only a few generations to allow measurement of the genetic distance between a disease gene and markers on chromosomes Requires sampling of families, not just people affected by the disease Loses power when disease has complex inheritance with substantial lack of penetrance Most often used to map disease-causing variants with strong enough effects to cause a mendelian inheritance pattern Tests for altered frequency of particular alleles or haplotypes in individuals with a disease, compared with population controls Examines particular alleles or haplotypes for their contribution to the disease Uses anywhere from a few markers in targeted genes to hundreds of thousands of markers for genome-wide analyses Can occasionally pinpoint the variant functionally responsible for the disease; more often, defines a disease-containing haplotype over a 1- to 10-kb interval Relies on finding a set of alleles, including the disease gene, that remained together for many generations due to lack of recombination among the markers Can be carried out on case-control or cohort samples from populations Is sensitive to population stratification artifact, although this can be controlled by proper case-control designs or the use of family-based approaches Is the best approach for finding variants with small effect that contribute to complex traits Determines variation in the whole genome or coding sequence (exome) in unbiased approach in families or cohorts Requires robust filtering strategy to narrow down rare variants based on segregation and expected disease mode Filtering can be within a family or across individuals with the same clinical diagnosis Can be used in conjunction with linkage data to aid in narrowing down region of the genome with causative variant Designed to precisely identify causative variant that is functionally responsible for the disease Does not rely on linked variants and is particularly useful for finding de novo dominant variants that are intractable to linkage or association studies Can use either family- or cohort-based analysis Designed to identify rare disease-causing variants but can perform genome-wide association studies Confirmation of gene disease association greatly enhanced with submission to Gene Matcher
sequenced, including deep intronic, regulatory regions and noncoding RNA, which may harbor pathogenic variation. Third, GS has the potential to detect nearly all classes of genetic variation, includin...
Ch11 · Pt21 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 223 gene or in regulatory sequence located some distance from a gene, as introduced in Chapter 3. However, these are currently more difficult to assess; as a simplifying assumption, it is reasonable to focus initially on protein coding genes. If the filtering scheme does not yield interesting candidates, one can always go back and assess noncoding variants. Splicing assessment algorithms are becoming better at predicting splice variants in deep intronic regions. 2. Population frequency. Keep rare variants from step 1 by discarding common variants with allele frequencies greater than expected for a rare disorder. For filtering, one would typically pick an allele frequency cutoff of between 0.03 and 0.05. Common variants are highly unlikely to be responsible for a rare disease whose population prevalence is much less than the q 2 predicted by Hardy-Weinberg equilibrium (see Chapter 10). 3. Deleterious nature of the variant. Keep variants from step 2 that cause loss of function changes, including nonsense, frameshift, or those that alter highly conserved (canonical) splice sites. Keep nonsynomous variants that are predicted to be damaging. Discard synonymous or intronic changes that have no predicted effect on gene function. 4. Consistency with likely inheritance pattern. If the disorder is considered most likely to be autosomal recessive, keep any variants from step 3 that are found in both copies of a gene in an affected child. The child need not be homozygous for the same deleterious variant but could be a compound heterozygote for two different deleterious changes in the same gene (see Chapter 7). If this hypothesized mode of inheritance is correct, then the parents should both be heterozygous for the variant(s). If the parents are consanguineous, the candidate genes and variants may be further filtered by requiring that the child be a true homozygote for the same variant derived from a single common ancestor (see Chapter 10). If the disorder is severe and seems more likely to be due to a new mutation for a dominant trait, keep variants from step 3 that are de novo changes in the child and are not present in either parent. Finally, if there is suspicion of an X-linked disorder and the child is male, one can focus on hemizygous variants that are inherited from a heterozygous mother. In the end, millions of variants can be filtered down to a handful occurring in a small number of genes. Once the filtering reduces the number of genes and alleles to a manageable number, they can be further assessed for other characteristics. Do any of the genes have a known function or tissue expression pattern that would be expected of a potential disease gene? Is the gene involved in other disease phenotypes, or does it have a role in pathways with other genes in which variation can cause similar or different phenotypes? Is the gene under constraint such that loss of function variants are 4–5,000,000 variants 3,000,000,000 bp in genome Not located within or near an exon Too frequent in variant databases to cause disease Variants with no predicted functional consequence ~40,000 variants ~1500 variants ~200 variants Inconsistent with Inheritance model Clinical correlation and/or makes biological sense 5–10 variants 1–2 variants Figure 11.14 Representative filtering scheme for whole-genome sequencing of a family consisting of two unaffected parents and an affected child, reducing the millions of variants detected down to a small number that can be assessed for biologic and disease relevance. The initial enormous collection of variants is reduced into smaller and smaller bins by applying filters that remove variants unlikely to be causative, based on assuming that variants of interest are likely to be located near a gene, will disrupt its function, and are rare. Each remaining candidate gene is then assessed for whether the variants are inherited in a manner that fits the most likely inheritance pattern of the disease, whether a variant occurs in a candidate gene that makes biologic sense given the phenotype in the affected child, and whether other affected individuals also have causative variants in that gene. product, population frequency, and inheritance pattern. Once these basic filters have been applied to narrow the number of variants, they can be further interrogated for clinical correlation, plausible biologic function, known expression, prior observation in other cases, and previous classifications. One example of a filtering scheme that can be used to sort through these variants is shown in Fig. 11.14; the order of filtering can be altered to produce similar results. 1. Location with respect to protein-coding genes. Keep variants that are within or near exons of protein-coding genes, and discard variants deep within introns or intergenic regions. It is possible, of course, that the variant responsible might lie in a noncoding RNA
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 223 gene or in regulatory sequence located some distance from a gene, as introduced in Chapter 3. However, these are currently more difficu...
Ch11 · Pt22 224 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE rarely observed? Finally, has the variant been observed and classified in others, or has pathogenic variation in the gene been observed in others with the disease? Finding causative variation in one of these genes in other affected individuals would lend evidence that this was the responsible gene and variant in the original trio. In some cases, one gene from the list in step 4 may rise to the top as a candidate because its involvement makes biologic or genetic sense, or it is known to be causative in other affected individuals. In other cases, however, the gene responsible may turn out to be entirely unanticipated on biologic grounds, or may not be causative in other affected individuals because of locus heterogeneity (i.e., pathogenic variants in other as yet undiscovered genes can cause a similar disease). Such variant assessments require extensive use of public genomic databases and software tools. These include the human genome reference sequence, databases of allele frequencies (e.g., 1000 genomes and gnom AD), software that assesses how deleterious an amino acid substitution might be to gene function, collections of known diseasecausing variants (e.g., Clin Var or disease specific), and databases of functional networks and biologic pathways. The enormous expansion of this information over the past few years, fueled by genome-wide sequencing of millions of cases and controls, has played a crucial role in facilitating gene discovery and molecular diagnosis of rare mendelian disorders, as we discuss in the next section. FILTERING STRATEGIES FOR IDENTIFICATION OF DISEASE-CAUSING GENES In the previous section we discussed filtering schemes for a single family to identify variants causing disease. But what if nothing is identified or there are interesting candidates that are lacking enough evidence to definitively link to disease? For rare mendelian disorders, analysis of a single affected individual is often inadequate for gene discovery, making other study designs and strategies necessary. This involves traditional mapping approaches as well as other common genetic strategies that have been adapted for genome-wide sequencing. Generally, sequencing of multiple affected individuals within a family or sequencing of unrelated individuals with the same clinical diagnosis is necessary. Deciding which cases are the most informative to sequence is highly influenced by the suspected mode of inheritance and whether the variants are expected to be inherited or due to new (de novo) mutations. Fig. 11.15 shows some examples of strategies employed for gene identification using genome-wide sequencing. This is not an exhaustive list but includes strategies for finding diseasecausing genes based on variant inheritance and sharing across affected individuals. For families demonstrating autosomal recessive inheritance, genome-wide sequencing and filtering strategies are relatively straightforward. In those families with no known consanguinity or not belonging to a founder population, compound heterozygous is the most likely disease model. Sequencing of affected sibling pairs is highly effective since there are very few shared compound heterozygous variants. In smaller families, sequencing of parents is useful to filter out variants that are both coming from one parent (in cis). In those families with expected or known consanguinity, causative variants are expected to be homozygous. Sequencing of sibling pairs or other affected relatives and prioritizing homozygous variants is an extremely effective strategy that has led to many new gene discoveries. If possible, sequencing of more distantly related individuals will aid in further reducing the number of shared homozygous variants. Homozygosity mapping with SNP microarrays to predetermine chromosomal loci with overlapping stretches of homozygosity in affected individuals can be used to narrow regions. However, with the cost of sequencing decreasing, current strategies typically only use the genome-wide data for homozygosity mapping, or to sequence more individuals and look for shared homozygous variants. Finally, for those families showing X-linked recessive inheritance, one can look for shared variants on chromosome X in males and filter out autosome variants. However, one must A B Autosomal Recessive Autosomal Dominant X-linked C D E Figure 11.15 Disease gene identification strategies using genome-wide sequencing for different suspected modes of inheritance. Strategies (A–E) are detailed in the text. Filled symbols indicate affected individuals, and empty symbols are unaffected. Obligate carriers are denoted with a dot. Cross lines depict individuals undergoing genome-wide sequencing, with circles below indicating variant sets and overlapping strategies for detection of causative variants.
224 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE rarely observed? Finally, has the variant been observed and classified in others, or has pathogenic variation in the gene been observed in o...
Ch11 · Pt23 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 225 be cautious in this approach that there is unequivocal evidence that the transmission is X-linked and not potentially autosomal recessive. The use of genome-wide sequencing for identifying genes causing inherited dominant disorders can be more challenging. This strategy involves looking for shared variants in affected individuals and excluding variants present in unaffected individuals. In principle, finding the causative variant can be difficult owing to the number of rare variants that are private to a family. To effectively filter variants down to a manageable number to interpret, sequencing of many affected and unaffected individuals is necessary, which can be costly especially if using GS. This can be partially mitigated in large families by using the principles of linkage that we discussed earlier in the chapter. Performing linkage to define a disease locus, then strategically sequencing a few distantly related individuals, can be a cost-effective approach to gene identification. If there are no candidates in the linked region, it raises the possibility of variant classes that may not be detectable due to technical limitations of genome-wide sequencing or that the causative variant is noncoding and difficult to interpret without functional assays. Genome-wide sequencing has been particularly successful for discovery of disorders caused by de novo dominant variants—a disease mode that is intractable to linkage studies. These disorders tend to be severe genetic lethals, with variants not passed on to subsequent generations. There are relatively few de novo events per exome (1–2) and genome (50–70); overlapping de novo events in two unrelated individuals with a similar phenotype can be sufficient to show disease causation, although overlap by chance is possible. The de novo overlap strategy has been particularly useful in finding causes for disorders with high genetic and phenotypic heterogeneity, such as syndromic forms of autism spectrum disorder and intellectual disability. Pooling of large case cohorts is often necessary to find de novo variants in the same gene, given their individually rare participation in a genotypically heterogeneous pool. Genes identified using this strategy are more likely to overlap by chance as the number of family trios sequenced increases. Thorough investigation of genotype and phenotype correlation is imperative to establish disease gene associations. Each of the strategies above can be hindered by genetic and phenotypic heterogeneity and the rarity of the disorders that remain unsolved. Indeed, much of the low-hanging fruit for rare disease associations has been discovered, and the majority of cases that undergo genome-wide sequencing are unsolved. It is clear that for study of exceedingly rare disorders or new disease mechanisms, international collaboration is necessary, as no one cohort will likely have the numbers to make robust disease associations. One such effort is the Matchmaker Exchange (MME), which uses a federated network of data sharing to solve undiagnosed cases and facilitates publication of case series describing new disease genes (Fig. 11.16). The MME provides a genetic Gene 1 Gene 5 Gene 4 Gene 7 Gene 6 Gene 2 Gene 3 Gene 3 Gene 3 Matchmaker exchange Matchmaker Exchange Statistics (last updated February 2022) MME Node DECIPHER (UK) Gene Matcher (USA) IRUD (Japan) My Gene 2 (USA) Patient Matcher (Sweden) Phenome Central (Canada) RD-Connect GPAP (Europe) 80,000 42,534 3,578 2,521 8,945 12,118 12,114 7,929 40,544 64,852 62 1,599 14 8,904 4,901 1,174 8,967 13,436 55 1,302 25 3,014 821 1,224 seqr (USA) Patients/Cases Total Patients/Cases in MME Unique Genes Figure 11.16 International collaboration through matchmaker exchange (MME). For rare mendelian disorders, federated networks of data sharing can facilitate solutions for undiagnosed cases and establish genotype-phenotype correlations. Candidate genes without established disease associations, sequenced as either part of research or clinical service, can be submitted to MME via several projects. Matching genes can be linked between submitters for more detailed phenotype-genotype correlation. In the example here, researchers and clinicians from three centers have submitted candidate genes (solid lines) for affected individuals to MME, with gene 3 in common. Direct connections between submitters (dotted bidirectional arrows) can then be established to further study potential causal relationships. The insert includes a snapshot of federated databases that link into MME with the number of cases and genes submitted.
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 225 be cautious in this approach that there is unequivocal evidence that the transmission is X-linked and not potentially autosomal recessi...
Ch11 · Pt24 226 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE matchmaking service that connects multiple researchers or clinicians who have patients with similar phenotypes and variants in the same candidate gene. Multiple nodes feed into MME, with a growing collective dataset of more than 150,000 cases submitted from 99 countries. As sequencing becomes less expensive, these datasets will continue to grow. The increased application of genome-wide sequencing coupled with sharing of data has played a pivotal role in our understanding of gene-disease associations, leading to a rapid increase in gene discovery over the last 10+ years. Since the application of GS or ES to rare mendelian disorders was first described in 2009, many hundreds of such disorders have been studied, and the causative variants found among hundreds of previously unrecognized disease genes (Fig. 11.17). These discoveries feed back into diagnostic testing where they not only provide information useful for genetic counseling in the families involved, but may inform clinical management and the potential development of effective treatments. The application of genome-wide sequencing has grown in diagnostic testing, notably in individuals with genetically heterogenous disorders, leading to a diagnostic yield of ~30%. The success rate of this approach will only increase as the costs of sequencing continue to fall and with improved ability to interpret the likely functional consequences of sequence changes in the genome. Example: Identification of the Gene Causing Postaxial Acrofacial Dysostosis The genome-wide sequencing approach just outlined was used in the study of a family in which two siblings affected with a rare congenital malformation, known as postaxial acrofacial dysostosis (POAD), were born to two unaffected, unrelated parents. Individuals with this disorder have small jaws, missing or poorly developed digits on the ulnar sides of their hands, underdevelopment of the ulna, cleft lip, and clefts (colobomas) of the eyelids. The disorder was thought to be autosomal recessive, because some parents of an affected child are consanguineous, and because a few families are like the one here, with multiple affected siblings born to unaffected parents—both findings that are hallmarks of recessive inheritance (see Chapter 7). This small family alone was clearly inadequate for linkage analysis. Instead, all four members of the family had their entire genomes sequenced and analyzed. From an initial list of more than 4 million variants and assuming autosomal recessive inheritance of the disorder in both affected children, a filtering scheme similar to that described earlier (see Figs. 11.14 and 11.15) yielded only four possible candidate genes. One of these, DHODH, had rare damaging variants in two other unrelated individuals with POAD, thereby identifying this gene as responsible for the disorder in these families. DHODH encodes dihydroorotate dehydrogenase, a mitochondrial enzyme involved in pyrimidine biosynthesis, and was not suspected on biologic grounds to be the gene responsible for this malformation syndrome. Limitations of Genome-Wide Sequencing and Future Outlook Although the genome-wide sequencing approach has proved powerful for both gene discovery and diagnosis of rare mendelian disease, it still has limitations. Most groups Growth of Gene-phenotype Relationships 31 Dec 2021 7500 6750 6000 5250 4500 3750 Year Phenotypes with known molecular basis Genes with phenotype-causing variant 3000 2250 1500 750 0 1987 1990 1993 1996 1999 2002 2005 2008 2011 2014 2017 2020 Figure 11.17 Growth of phenotype and disease gene associations. Growth of gene association with disease phenotype has steadily increased due to the advent of genome-wide sequencing combined with data sharing.
226 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE matchmaking service that connects multiple researchers or clinicians who have patients with similar phenotypes and variants in the same cand...
Ch11 · Pt25 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 227 report diagnostic yields in the 20 to 40% range, depending on clinical indication, which leaves the majority of cases without an identified causative variant and answer. There are several reasons for this observation. First, some disorders are intractable to standard genome-wide sequencing, including methylation disorders (e.g., Prader-Willi and Angelman syndromes) or those involving certain types of Uniparental disomy (UPD) if parents are not sequenced (i.e., heterodisomy). Second (and related), genome-wide sequencing may miss certain classes of variation that are difficult to detect by routine short-read sequencing alone (e.g., balanced changes, repeat expansions). As discussed earlier, whole genome sequencing has technological advantages over exome sequencing, with the ability to detect a broader range of variation; however, there are still limitations in accurately detecting or resolving complex regions of the genome. Third, variation is detected that is difficult to interpret with our current understanding of the genome. This is particularly true of genome sequencing and the rare variation detected in noncoding and regulatory genomic regions. Finally, at present, whole genome sequencing cannot actually sequence the entire genome in any one individual. Cost limits most of the sequencing to short-reads and aligning back to a reference genome (termed resequencing) to detect variation. Some sequence is too complex and/or highly homologous or repetitive to be technically sequenced or mapped back to the reference genome. Additionally, sequence novel to one’s genome, i.e., not in the reference assembly, will be filtered out even though it may be pathogenic. Such limitations are starting to be addressed in several ways. Improvements in both sequencing technology and informatics algorithms for variant detection allow a more accurate catalogue of genomic variation. The increasing number of genomes sequenced and the subsequent aggregation of data allow more accurate interpretation of variation. In addition, the advancement of other -omic technologies, such as RNA sequencing or methylation experiments, are proving to be valuable tests ancillary to genome-wide sequencing for the interpretation of noncoding variation. The emergence of long-read sequencing technology (reads up to 10–100 kb in length) has provided insight into regions of the genome that have been intractable to short-read sequencing, including highly homologous or repetitive regions, with detection of variation previously unseen. Long-read technology can also be used to advance de novo assembly of genomes, rather than reference-based assembly, to get a more complete picture of the genome. These advances will lead to a better understanding of our genome, its variation, and its relation to disease. ACKNOWLEDGMENT We wish to thank Ada Hamosh and Gregory Costain for contributing to this chapter. GENERAL REFERENCES Altshuler D, Daly MJ, Lander ES: Genetic mapping in human disease, Science 322:881–888, 2008. Boycott KM, Azzariti DR, Hamosh A, et al: Seven years since the launch of the Matchmaker Exchange: The evolution of genomic matchmaking, Hum Mutat, 2022. https://doi.org/10.1002/humu.24373. Online ahead of print. PMID: 35537081. Boycott KM, Vanstone MR, Bulman DE, et al: Rare-disease genetics in the era of next-generation sequencing: Discovery to translation, Nat Rev Genet 14:681–691, 2013. Gilissen C, Hoischen A, Brunner HG, et al: Disease gene identification strategies for exome sequencing, Eur J Hum Genet 20:490–497, 2012. Manolio TA: Genomewide association studies and assessment of the risk of disease, NEJM 363:166–176, 2010. Risch N, Merikangas K: The future of genetic studies of complex human diseases, Science 273:1516–1517, 1996. Sullivan PF, Daly MJ, O’Donovan M: Genetic architectures of psychiatric disorders: The emerging picture and its implications, Nat Rev Genet 13:537–551, 2012. Terwilliger JD, Ott J: Handbook of human genetic linkage, Baltimore, 1994, Johns Hopkins University Press. REFERENCES FOR SPECIFIC TOPICS Abecasis GR, Auton A, Brooks LD, et al: An integrated map of genetic variation from 1,092 human genomes, Nature 491:56–65, 2012. Bainbridge MN, Wiszniewski W, Murdock DR, et al: Whole-genome sequencing for optimized patient management, Science Transl Med 3:87re 3, 2011. Bush WS, Moore JH: Genome-wide association studies, PLo S Computational Biol 8:e 1002822, 2012. Denny JC, Bastarache L, Ritchie MD, et al: Systematic comparison of phenome-wide association study of electronic medical record data and genome-wide association data, Nat Biotechnol 31:1102–1110, 2013. Fritsche LG, Chen W, Schu M, et al: Seven new loci associated with age-related macular degeneration, Nat Genet 17:1783–1786, 2013. Gonzaga-Jauregui C, Lupski JR, Gibbs RA: Human genome sequencing in health and disease, Ann Rev Med 63:35–61, 2012. Hindorff LA, Mac Arthur J, Morales J, et al: A catalog of published genome-wide association studies, 2015. www.genome.gov/ gwastudies International Hap Map Consortium: A second generation human haplotype map of over 3.1 million SNPs, Nature 449:851–861, 2007. Kircher M, Witten DM, Jain P, et al: A general framework for estimating the relative pathogenicity of human genetic variants, Nat Genet 46:310–315, 2014. Koboldt DC, Steinberg KM, Larson DE, et al: The next-generation sequencing revolution and its impact on genomics, Cell 155:27–38, 2013. Lionel AC, Costain G, Monfared N, et al: Improved diagnostic yield compared with targeted gene sequencing panels suggests a role for whole-genome sequencing as a first-tier genetic test, Genet Med 20:435–443, 2018. Manolio TA: Bringing genome-wide association findings into clinical use, Nat Rev Genet 14:549–558, 2014. Marshall CR, Howrigan DP, Merico D, et al: Contribution of copy number variants to schizophrenia from a genome-wide study of 41,321 subjects, Nat Genet 49:27–35, 2016. Matise TC, Chen F, Chen W, et al: A second-generation combined linkage-physical map of the human genome, Genome Res 17:1783– 1786, 2007. Roach JC, Glusman G, Smit AF, et al: Analysis of genetic inheritance in a family quartet by whole-genome sequencing, Science 328:636– 639, 2010. Robinson PC, Brown MA: Genetics of ankylosing spondylitis, Mol Immunol 57:2–11, 2014. SEARCH Collaborative Group: SLCO1B1 variants and statin-induced myopathy—A genomewide study, NEJM 359:789–799, 2008. Stahl EA, Wegmann D, Trynka G, et al: Bayesian inference analyses of the polygenic architecture of rheumatoid arthritis, Nat Genet 44:4383–4391, 2012. Yang Y, Muzny DM, Reid JG, et al: Clinical whole-exome sequencing for the diagnosis of mendelian disorders, NEJM 369:1502–1511, 2013. Yuen RK, Merico D, Bookman M, et al: Whole genome sequencing resource identifies 18 new candidate genes for autism spectrum disorder, Nat Neurosci 20:602–611, 2017.
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 227 report diagnostic yields in the 20 to 40% range, depending on clinical indication, which leaves the majority of cases without an identi...
Ch11 · Pt26 228 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE PROBLEMS 1. In the early days of gene mapping the Huntington disease (HD) locus was found to be tightly linked to a DNA common variant on chromosome 4. In the same study, however, linkage was ruled out between HD and the locus for the polymorphic MNSs blood group, which also maps to chromosome 4. What is the explanation? 2. LOD scores (Z) between a common variant in the α-globin locus on the short arm of chromosome 16 and an autosomal dominant disease (e.g., polycystic kidney disease) was analyzed in a series of British and Dutch families, with the following data: θ 0.00 0.01 0.10 0.20 0.30 0.40 Z −∞ 23.4 24.6 19.5 12.85 5.5 Z 25.85 at 0.05 max max
How would you interpret these data? In a subsequent study, a large family from Sicily with what looks like the same disease was also investigated for linkage to α-globin, with the following results: θ...
Ch11 · Pt27 CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 229 PROBLEMS—CONT’D Pedigree of X-linked hemophilia. The affected grandfather in the first generation has the disease (allele h) and allele M at a polymorphic locus on the X chromosome. 7. Relative risk calculations are used for cohort studies and not case-control studies. To demonstrate why, imagine a case-control study for the effect of a genetic variant on disease susceptibility. The investigator has ascertained as many affected individuals (a + c) as possible and then arbitrarily chooses a set of (b + d) controls. They are genotyped as to whether a variant is present: a/(a + c) of the affected have the variant, whereas b/(b + d) of the controls have the variant. Disease Present Disease Absent Variant present a b Variant absent c d a+c b+d Calculate the odds ratio (OR) and relative risk (RR) for the association between the variant being present and the disease being present. Now, imagine the investigator arbitrarily decided to use three times as many unaffected individuals, 3 × (b + d), as controls. The investigator has every right to do so because it is a case-control study and the numbers of affected and unaffected are not determined by the prevalence of the disease in the population being studied, as they would be in a cohort study. Assume the distribution of the variant remains the same in this control group as with the smaller control group—that is, 3b/[3 × (b + d)] = b/(b + d) carrying the allele. Disease Present Disease Absent Variant present a 3b Variant absent c 3d a+c 3 × (b + d) Recalculate the OR and RR with this new control group. Do the same when an arbitrary control group is an n-tuple of the original control group—that is, the size of the control group is n × (b + d).
CHAPTER 11 — IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 229 PROBLEMS—CONT’D Pedigree of X-linked hemophilia. The affected grandfather in the first generation has the disease (allele h) and allele...

Chapter 12: The Molecular Basis of Genetic Disease

Ch12 · Pt1 The Molecular Basis of Genetic Disease Gregory Costain GENERAL PRINCIPLES AND LESSONS FROM THE HEMOGLOBINOPATHIES The term molecular disease, introduced in 1949, refers to disorders in which the primary disease-­causing event is an alteration, either inherited or acquired, affecting a gene(s), its structure, and/­or its expression. In this chapter we first outline the basic DNA variant types and associated mechanisms underlying monogenic (single-­gene) disorders. We then illustrate their molecular and clinical consequences using inherited diseases of hemoglobin—­the ­hemoglobinopathies—­as examples. This overview of mechanisms is expanded on in Chapter 13 to include other genetic diseases that illustrate key principles of genetics in medicine. A genetic disease occurs when an alteration in the DNA of an essential gene changes the amount or function, or both, of the gene products—­typically messenger RNA (mRNA) and protein, but occasionally specific noncoding RNAs (ncRNAs) with structural or regulatory functions. Although almost all known single-­gene disorders result from variants that affect the function of a protein, a few exceptions to this generalization are now known. These exceptions are diseases due to variants in ncRNA genes, including micro RNA (miRNA) genes that regulate specific target genes, and mitochondrial genes that encode transfer RNAs (tRNAs; see Chapter 13). In this chapter we restrict our attention to diseases caused by defects in protein-­coding genes. It is essential to understand genetic disease at the molecular level because this knowledge is the foundation of rational therapy. By 2022, the online version of Mendelian Inheritance in Man listed over 7000 phenotypes for which the molecular basis is known. Although it is impressive that the basic molecular defect has been found in so many disorders, it is sobering to realize that the pathophysiology is not entirely understood for any genetic disease. Sickle cell disease (Case 42), discussed later in this chapter, was the first disease to be characterized at the molecular level and remains among the best characterized of all inherited disorders; even here, knowledge is incomplete. Genetically informed therapies for hemoglobinopathies are now emerging as a realistic prospect in the clinic, thanks in part to an increasingly sophisticated understanding of the genetic pathomechanisms. EFFECT OF PATHOGENIC VARIANTS ON PROTEIN FUNCTION DNA variants within protein-­coding genes have been primarily found to cause disease through one of four different effects on protein function (Fig. 12.1). The most common effect by far is a loss of function of the protein. Many important conditions arise, however, from other mechanisms: a gain of function, the acquisition of a novel property by the affected protein, or the expression of a gene at the wrong time (heterochronic expression) and/­or in the wrong place (ectopic expression). Loss-­of-­Function Variants The loss of function of a gene may result from alteration of its coding, regulatory, or other critical sequences due to nucleotide substitutions, deletions, insertions, or rearrangements. A loss of function due to deletion, leading to a reduction in gene dosage, is exemplified by the α-­thalassemias (Case 44), which are most commonly due to deletion of α-­globin genes (see later discussion); by chromosome-­loss diseases (Case 27), such as monosomies like Turner syndrome (see Chapter 6) (Case 47); and by acquired somatic variants that occur in tumor-­suppressor genes in many cancers, such as retinoblastoma (Case 39) (see Chapter 16). Many other types of variants can also lead to a complete loss of function, and all are illustrated by the β-­thalassemias (see later discussion), a group of hemoglobinopathies that result from a reduction in the abundance of β-­globin, one of the major adult hemoglobin proteins in red blood cells. The severity of a disease due to loss-­of-­function variants generally correlates with the amount of function lost. In many instances, the retention of even a small percent of residual function by the abnormal protein greatly reduces the severity of the disease. Gain-­of-­Function Variants Variants may also enhance one or more of the normal functions of a protein; in a biologic system, however, more is not necessarily better, and disease may result. It chapter 12
The Molecular Basis of Genetic Disease Gregory Costain GENERAL PRINCIPLES AND LESSONS FROM THE HEMOGLOBINOPATHIES The term molecular disease, introduced in 1949, refers to disorders in which the prima...
Ch12 · Pt2 232 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE is critical to recognize when a disease is due to a gain-­of-­ function variant because the treatment must necessarily differ from disorders due to other mechanisms, such as loss-­of-­function variants. Gain-­of-­function variants fall into two broad classes: Variants that increase the production of a normal protein. Some variants cause disease by increasing the synthesis of a normal protein in cells in which the protein is normally present. The most common variants of this type are due to increased gene dosage, which generally results from duplication of part or all of a chromosome. As discussed in Chapter 6, the classic example is trisomy 21 (Down syndrome), which is due to the presence of three copies of chromosome 21. Other important diseases arise from the increased dosage of single genes, including one form of familial Alzheimer disease due to a duplication of the amyloid precursor protein (β APP) gene (see Chapter 13), and the peripheral nerve degeneration Charcot-­Marie-­Tooth disease type 1 A (Case 8), which generally results from duplication of the gene for peripheral myelin protein 22 (PMP22). Variants that enhance one normal function of a protein. Rarely, a variant in the coding region may increase the ability of each protein molecule to perform one or more of its normal functions, even though this increase is detrimental to the overall physiologic role of the protein. For example, the missense variant that creates hemoglobin Kempsey locks hemoglobin into its high oxygen affinity state, thereby reducing oxygen delivery to tissues. Another example of this mechanism is the missense variation in the FGFR3 gene that causes achondroplasia (Case 2), the most common skeletal dysplasia. Novel Property Variants In a few diseases, a change in the amino acid sequence confers a novel property on the protein, without necessarily altering its normal functions. The classic example of this mechanism is sickle cell disease (Case 42), which, as we will see later in this chapter, is due to an amino acid substitution that has no effect on the ability of sickle hemoglobin to transport oxygen. Rather, unlike normal Variants disrupting RNA stability or RNA splicing Variants affecting gene regulation or dosage Variants in coding region Decreased amount CAUSE OF DISEASE Loss of protein function (the great majority) Gain of function Novel property (infrequent) Ectopic or heterochronic expression (uncommon, except in cancer) Inappropriate expression (wrong time, place) HPFH Many oncogenes Increased amount Trisomies Charcot-Marie-Tooth disease type 1A -Thalassemias Monosomies Tumor-suppressor variants Hb Hammersmith -Thalassemias Hb Kempsey Achondroplasia Hb S MUTATION Protein abnormal (if unstable decreased amount) Protein structure normal Figure 12.1 A general outline of the mechanisms by which disease-­causing variants produce disease. Variants in the coding region result in structurally abnormal proteins that have a loss or gain of function or a novel property that causes disease. Variants in noncoding sequences are of two general types: those that alter the stability or splicing of the messenger RNA (mRNA) and those that disrupt regulatory elements or change gene dosage. Variants in regulatory elements alter the abundance of the mRNA or the time or cell type in which the gene is expressed. Variants in either the coding region or regulatory domains can decrease the amount of the protein produced. HPFH, Hereditary persistence of fetal hemoglobin.
232 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE is critical to recognize when a disease is due to a gain-­of-­ function variant because the treatment must necessarily differ from disorders...
Ch12 · Pt3 CHAPTER 12 — The Molecular Basis of Genetic Disease 233 hemoglobin, sickle hemoglobin chains aggregate when they are deoxygenated and form abnormal polymeric fibers that deform red blood cells. That novel property variants are infrequent is not surprising because most amino acid substitutions are either neutral or detrimental to the function or stability of a protein that has been finely tuned by evolution. Variants Associated With Heterochronic or Ectopic Gene Expression An important class of variants includes those that lead to inappropriate expression of the gene at an abnormal time or place. These variants occur in the regulatory regions of the gene. Cancer can be driven by expression of a gene that normally promotes cell ­proliferation—­ a proto-oncogene—­in cells in which the gene is not normally expressed (see Chapter 16). Some variants in hemoglobin regulatory elements lead to the continued expression in adults of the γ-­globin gene, which is normally expressed at high levels only in fetal life. Such γ-­globin gene variants cause a benign phenotype called hereditary persistence of fetal hemoglobin (Hb F), as we explore later in this chapter. HOW VARIANTS DISRUPT THE FORMATION OF BIOLOGICALLY NORMAL PROTEINS Disruptions of the normal functions of a protein that result from the different types of variants outlined earlier can be well exemplified by the broad range of diseases due to variants in the globin genes, as we will discuss in the second part of this chapter. To form a biologically active protein (such as the hemoglobin molecule), information must be transcribed from the nucleotide sequence of the gene to the mRNA and then translated into the polypeptide, which then undergoes progressive stages of maturation (see Chapter 3). Variants can disrupt any of these steps (Table 12.1). As we shall see next, abnormalities in five of these stages are illustrated by various hemoglobinopathies; the others are exemplified by diseases to be presented in Chapter 13. THE RELATIONSHIP BETWEEN GENOTYPE AND PHENOTYPE IN GENETIC DISEASE Key molecular concepts that can account for differences in the observed clinical phenotype associated with a genetic disease are: Allelic heterogeneity Locus heterogeneity Effect of modifier genes Each of these concepts is illustrated by variants in the α-­globin and/­or β-­globin genes (Table 12.2). Allelic Heterogeneity Genetic heterogeneity is most commonly due to the presence of multiple alleles at a single locus, a situation referred to as allelic heterogeneity (see Chapter 7 and Table 12.1). In some instances there may be a clear genotype-­phenotype correlation between a specific allele and a specific phenotype. The most common explanation for the effect of allelic heterogeneity on the clinical TABLE 12.1 Eight Steps at Which DNA Variants Can Disrupt the Production of a Normal Protein Step Phenotype Example Transcription
Thalassemias due to reduced or absent production of a globin mRNA because of deletions or variants in regulatory or splice sites of a globin gene Hereditary persistence of fetal hemoglobin, which resu...
Ch12 · Pt4 234 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE phenotype is that alleles that confer more residual function on the altered protein are often associated with a milder form of the principal phenotype associated with the disease. In some instances, however, alleles that confer some residual protein functions are associated with only one or a subset of the phenotypes seen with a missing or completely nonfunctional allele (frequently termed a null allele). As we will explore more fully in Chapter 13, this situation prevails with certain variants of the cystic fibrosis gene, CFTR, that lead to a phenotypically different condition—congenital absence of the vas deferens, but not to the other manifestations of cystic fibrosis. An important exception to this rule relates to variants that act in a dominant-­negative fashion, as exemplified by select missense variants in the genes encoding components of type I collagen that result in a more severe form of osteogenesis imperfecta than with a null allele (see Chapter 13). A second explanation for allele-­based differences in phenotype is that a specific property of the protein may be more perturbed by a particular variant. This situation is well illustrated by Hb Kempsey, a β-­globin allele that maintains the hemoglobin in a high oxygen affinity structure. This causes polycythemia because the reduced peripheral delivery of oxygen is misinterpreted by the hematopoietic system as being due to an inadequate production of red blood cells. The consequences of a specific variant on the function of a protein can be unpredictable. No one would have foreseen that the β-­globin allele associated with sickle cell disease would lead to the formation of globin polymers that deform erythrocytes to a sickle cell shape (see later in this chapter). However, sickle cell disease is also unusual in that it results only from a single specific variant—­the p. Glu 6Val substitution in the β-­globin chain—­whereas most genetic diseases can arise from any of a number of different DNA-­level variants in the corresponding gene. Locus Heterogeneity Genetic heterogeneity also arises when variants at more than one locus can result in a specific clinical condition—a situation termed locus heterogeneity (see Chapter 7). This phenomenon is illustrated by the finding that thalassemia can result from variants in either the α-­globin or β-­globin chain genes (see Table 12.2). Once locus heterogeneity has been documented, careful comparison of the phenotype associated with each gene sometimes reveals that the phenotype is not as homogeneous as initially believed. Modifier Genes Sometimes even the most robust genotype-­phenotype relationships are found not to hold for a specific individual. Such phenotypic variation can, in principle, be ascribed to nongenetic (e.g., environmental, stochastic) factors or to the action of other genes, termed modifier genes (see Chapter 9). Identified modifier genes for specific human monogenic disorders are growing in number; however, there remain relatively few examples with clinically relevant effect sizes or therapeutic significance. As described later in this chapter, individuals with β-­thalassemia who also have a deletion at the α-­globin locus can have a less severe phenotype. HUMAN HEMOGLOBIN AND ASSOCIATED DISEASES To illustrate in greater detail the concepts introduced in the first section of this chapter, we now turn to disorders of hemoglobin. These hemoglobinopathies are collectively the most common monogenic diseases in humans, and major contributors to global morbidity. The World Health Organization estimates that more than 5% of the world’s population are heterozygous carriers of genetic variants associated with clinically important disorders of hemoglobin. Hemoglobinopathies are also important because their molecular and biochemical pathology is better understood than perhaps that of any other group of genetic diseases. Indeed, our understanding of the basic anatomy of a gene (Chapter 3) arose in large part from studying these prototypical monogenic disorders of hemoglobin. Before the hemoglobinopathies are discussed in depth, it is important to briefly introduce the normal aspects of the globin genes and hemoglobin biology. Structure and Function of Hemoglobin Hemoglobin is the oxygen carrier in vertebrate red blood cells. Each hemoglobin molecule consists of four subunits: two α-­ (or α-­like) globin chains and two β-­ (or β-­like) globin chains. Each subunit is composed of a polypeptide chain, globin, and a prosthetic group, heme. The latter is an iron-­containing pigment that combines with TABLE 12.2 Types of Heterogeneity Associated With Genetic Disease Type of Heterogeneity Definition Example From the Hemoglobinopathies Genetic Allelic heterogeneity The occurrence of more than one allele at a locus β-­Thalassemia Locus heterogeneity The association of more than one locus with a clinical phenotype Thalassemia can result from variants in either the α-­globin or β-­globin genes Clinical or phenotypic The association of more than one phenotype with variants at a single locus Sickle cell disease and β-­thalassemia each result from distinct β-­globin gene variants
234 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE phenotype is that alleles that confer more residual function on the altered protein are often associated with a milder form of the principal...
Ch12 · Pt5 CHAPTER 12 — The Molecular Basis of Genetic Disease 235 oxygen to give the molecule its oxygen-­transporting ability (Fig. 12.2). The predominant adult human hemoglobin, Hb A, has an α2β2 structure in which the four chains are folded and fit together to form a globular tetramer. As with all proteins that have been strongly conserved throughout evolution, the tertiary structure of globins is constant; virtually all globins have seven or eight helical regions (depending on the chain) (see Fig. 12.2). Variants that disrupt this tertiary structure invariably have pathologic consequences. In addition, variants that substitute a highly conserved amino acid or that replace one of the nonpolar residues—which form the hydrophobic shell that excludes water from the interior of the molecule, are likely to cause a hemoglobinopathy (see Fig. 12.2). Like all proteins, globin has sensitive areas, in which variants cannot occur without affecting function, and insensitive areas, in which variations are more freely tolerated. The Globin Genes In addition to Hb A, with its α2β2 structure, there are five other normal human hemoglobins, each of which has a tetrameric structure like that of Hb A, consisting of two α or α-­like chains and two non-­α chains (Fig. 12.3A). The genes for the α and α-­like chains are clustered in a tandem arrangement on chromosome 16. Note that there are two identical α-­globin genes, designated α1 and α2, on each homologue. The β-­ and β-­like globin genes, located on chromosome 11, are close family members that, as described in Chapter 3, undoubtedly arose from a common ancestral gene (see Fig. 12.3A). Illustrating this close evolutionary relationship, the β-­ and δ-­globins differ in only 10 of their 146 amino acids. Developmental Expression of Globin Genes and Globin Switching The expression of the various globin genes changes during development, a process referred to as globin switching (see Fig. 12.3B). Note that the genes in the α-­ and β-­globin clusters are arranged in the same transcriptional orientation and, remarkably, the genes in each cluster are situated in the same order in which they are expressed during development. The temporal switches of globin synthesis are accompanied by changes in the principal site of erythropoiesis (see Fig. 12.3B). The three embryonic globins are made in the yolk sac from the third to eighth weeks of gestation, but at approximately the fifth week, hematopoiesis begins to move from the yolk sac to the fetal liver. Hb F (α2γ2), the predominant hemoglobin throughout fetal life, constitutes ~70% of total hemoglobin at birth. In adults, however, Hb F represents only a few percent of the total hemoglobin, although this can vary from less than 1% to ~5% in different individuals. β-­chain synthesis becomes significant near the time of birth, and by 3 months of age almost all hemoglobin is of the adult form: Hb A (α2β2) (see Fig. 12.3B). In diseases due to variants that decrease the abundance of β-­globin, such as β-­thalassemia (see later section), strategies to increase the normally small amount of γ-­globin (and therefore of Hb F [α2γ2]) produced in adults are proving to be successful in ameliorating the disorder (see Chapter 14). The Developmental Regulation of β-­Globin Gene Expression: The Locus Control Region Elucidation of the mechanisms that control expression of the globin genes has provided generalizable insights into both normal and pathologic biologic processes. The expression of the β-­globin gene is only partly controlled by the promoter and two enhancers in the immediate flanking DNA (see Chapter 3). A requirement for additional regulatory elements was first suggested by the identification of individuals who had no gene expression from any of the genes in the β-­globin cluster, even though the genes themselves (including their individual regulatory elements) were intact. These informative patients were found to have large deletions upstream of the β-­globin complex that removed an ~20 ­kb domain—now called the locus control region (LCR), located ~6 kb upstream of the ε-­globin gene (Fig. 12.4). The resulting disease, εγδβ-­ thalassemia, is described later in this chapter. These cases show us that the LCR is required for the expression of all genes in the β-­globin cluster. The LCR is defined by five DNase I hypersensitive sites (see Fig. 12.4): genomic regions that are unusually open to certain proteins (including the enzyme DNase I) used experimentally to reveal potential regulatory sites. Within the context of the epigenetic packaging of chromatin (see Chapter 3), these sites maintain an open chromatin Heme E B G b His 92 b Phe 42 H F Helix A D C Figure 12.2 The structure of a hemoglobin subunit. Each subunit has eight helical regions, designated A to H. The two most conserved amino acids are shown: p. His 92, the histidine to which the iron of heme is covalently linked; and p. Phe 42, the phenylalanine that wedges the porphyrin ring of heme into the heme “pocket” of the folded protein. See discussion of Hb Hammersmith and Hb Hyde Park, which have substitutions for p. Phe 42 and p. His 92, respectively, in the β-­globin molecule.
.
Ch12 · Pt6 236 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Developmental period Embryonic Fetal Adult Hemoglobins Hb Gower 2 α2ε2 Hb F α2γ2 β-like genes α-like genes Hb Gower 1 ζ2ε2 Hb A2 α2δ2 Hb A α2β2 Birth 5' 5' 3' 3' β Gγ α2 α1 ζ ε Aγ δ Hb Portland ζ2γ2 48 42 36 30 24 18 12 6 Postnatal age (weeks) Birth 36 30 24 18 12 6 Gestational age (weeks) 10 20 30 40 50 Site of erythropoiesis Percentage of total globin synthesis Liver Spleen Bone marrow α β α γ β ε ζ γ δ Yolk sac A B Figure 12.3 Organization of the human globin genes and hemoglobins produced in each stage of human development. (A) The α-­like genes are on chromosome 16, the β-­like genes on chromosome 11. The curved arrows refer to the switches in gene expression during development. (B) Development of erythropoiesis in the human fetus and infant. Types of cells responsible for hemoglobin synthesis, organs involved, and types of globin chain synthesized at successive stages are shown. (A, Redrawn from Stamatoyannopoulos G, Nienhuis AW: Hemoglobin switching. In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders; B, redrawn from Wood WG: Haemoglobin synthesis during fetal development, Br Med Bull 32:282–­287, 1976.) Normal 10 kb 4321 5 LCR Gγ Aγ ψβ δ β ε Gγ Aγ ψβ δ β ε Hispanic εγδβthalassemia Deletion 10 kb Figure 12.4 The β-­globin locus control region (LCR). Each of the five regions of open chromatin (arrows) contains several consensus binding sites for both erythroid-­specific and ubiquitous transcription factors. The precise mechanism by which the LCR regulates gene expression is unknown. Also shown is a deletion of the LCR that has led to εγδβ-­thalassemia, which is discussed in the text. (Redrawn from Kazazian Jr HH, Antonarakis S: Molecular genetics of the globin genes. In Singer M, Berg P, editors: Exploring genetic mechanisms, Sausalito, 1997, University Science Books.)
236 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Developmental period Embryonic Fetal Adult Hemoglobins Hb Gower 2 α2ε2 Hb F α2γ2 β-like genes α-like genes Hb Gower 1 ζ2ε2 Hb A2 α2δ2 Hb A α...
Ch12 · Pt7 CHAPTER 12 — The Molecular Basis of Genetic Disease 237 configuration that gives transcription factors access to the regulatory elements that mediate the expression of each of the β-­globin genes in erythroid cells (see Chapter 3). The LCR, along with its associated DNA-­binding proteins, interacts with the genes of the β-­globin locus to form a nuclear domain called the active chromatin hub, where β-­globin gene expression takes place. The sequential switching of gene expression that occurs among the five members of the β-­globin gene complex during development results from the sequential association of the active chromatin hub with the different genes in the cluster, as the hub moves from the most proximal gene in the complex (the ε-­globin gene in embryos) to the most distal (the δ-­ and β-­globin genes in adults). The clinical significance of the LCR could extend beyond those individuals with deletions of the LCR who fail to express the genes of the β-­globin cluster. Components of the LCR may prove relevant to gene therapy (see Chapter 14) for disorders of the β-­globin cluster, wherein a goal is for the therapeutic normal copy of the gene in question to be expressed at the correct time in life and in the appropriate tissue. Knowledge of the molecular mechanisms that underlie globin switching may also make it feasible to up-­regulate the expression of the γ-­globin gene in those with β-­thalassemia (who have variants only in the β-­globin gene) because Hb F (α2γ2) is an effective oxygen carrier in adults who lack Hb A (α2β2) (see Chapter 14). Gene Dosage, Developmental Expression of the Globins, and Clinical Disease The differences, both in the gene dosage of the α-­ and β-­globins (four α-­globin and two β-­globin genes per diploid genome) and in their patterns of expression during development, are important to an understanding of the pathogenesis of many hemoglobinopathies. A variant in a β-­globin gene affects 50% of the β chains, whereas a single α-­chain variant affects only 25% of the α chains. β-­globin variants have no prenatal consequences because γ-­globin is the major β-­like globin before birth, with Hb F constituting 75% of the total hemoglobin at term (see Fig. 12.3B). In contrast, because α chains are the only α-­like components of hemoglobin 6 weeks after conception, α-­globin variants cause severe disease in both fetal and postnatal life. THE HEMOGLOBINOPATHIES Hereditary disorders of hemoglobin can be divided into the following three broad groups, which, in some rare instances, overlap: Structural alterations of the amino acid sequence of the globin polypeptide, altering properties such as its ability to transport oxygen, or reducing its stability. An example is sickle cell disease (Case 42) due to a missense variant that makes deoxygenated β-­globin relatively insoluble, changing the shape of the red cell (Fig. 12.5). Thalassemias, which are diseases that result from the decreased abundance of one or more of the globin chains (Case 44). The decrease can result from reduced production of a globin chain or, less commonly, from a variant that destabilizes the chain. The resulting imbalance in the ratio of the α:β chains underlies the pathophysiology of these conditions. Examples include promoter variants that decrease expression of the β-­globin mRNA to cause β-­thalassemia. Hereditary persistence of fetal hemoglobin, a group of clinically benign conditions that impair the perinatal switch from γ-­globin to β-­globin synthesis. An example of a causal variant is a deletion that removes both the δ-­ and β-­globin genes but leads to continued postnatal expression of the γ-­globin genes, to produce Hb F, which is an effective oxygen transporter (see Fig. 12.3). Hemoglobin Structural Alterations Most variant hemoglobins result from single nucleotide variants in one of the globin genes. More than 500 abnormal hemoglobins have been described, and approximately half of these are clinically significant. The hemoglobin structural alterations can be separated B A Figure 12.5 Scanning electron micrographs of red cells from a patient with sickle cell disease. (A) Oxygenated cells are round and full. (B) The classic sickle cell shape is produced only when the cells are in the deoxygenated state. (From Kaul DK, Fabry ME, Windisch P, et al: Erythrocytes in sickle cell anemia are heterogeneous in their rheological and hemodynamic characteristics, J Clin Invest 72:22, 1983.)
CHAPTER 12 — The Molecular Basis of Genetic Disease 237 configuration that gives transcription factors access to the regulatory elements that mediate the expression of each of the β-­globin genes in e...
Ch12 · Pt8 238 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE into the following three classes, depending on the clinical phenotype (Table 12.3): Alterations that cause hemolytic anemia, most commonly because they make the hemoglobin tetramer unstable. Alterations with modified oxygen transport, due to increased or decreased oxygen affinity or to the formation of methemoglobin—a form of globin incapable of reversible oxygenation. Alterations due to variants in the coding region that cause thalassemia because they reduce the abundance of a globin polypeptide. Most of these variants impair the rate of synthesis of the mRNA or otherwise affect the level of the encoded protein. Hemolytic Anemias Hemoglobins With Novel Physical Properties: Sickle Cell Disease. Sickle cell hemoglobin is of great clinical importance in many parts of the world, affecting millions. The causal variant is a single nucleotide substitution that changes the codon of the sixth amino acid of β-­globin from glutamic acid to valine (GAG → GTG: p. Glu 6Val) (see Table 12.3). Homozygosity for this variant is the cause of sickle cell disease (Case 42). The disease has a characteristic geographic distribution: occurring most frequently in equatorial Africa and less commonly in the Mediterranean area, India, Spanish-­ speaking regions in the Western Hemisphere, or in countries to which people from these regions have migrated. Approximately 1 in 400 Black persons in the United States is born with sickle cell disease. Clinical Features. Sickle cell disease is a severe autosomal recessive hemolytic condition characterized by a tendency of the red blood cells to become grossly abnormal in shape (i.e., take on a sickle shape) under conditions of low oxygen tension (see Fig. 12.5). Heterozygotes—who are said to have sickle cell trait, are, generally, clinically unaffected, but their red cells can sickle when subjected to very low oxygen pressure. Occasions when this occurs are uncommon, although heterozygotes appear to be at risk for splenic infarction, especially at high altitude (e.g., in airplanes with reduced cabin pressure) or when exerting themselves to extreme levels in athletic competition. The heterozygous state is present in ~8% of Black individuals in the United States, but in areas where the sickle cell allele (β S) frequency is high (e.g., West Central Africa), up to 25% of the newborn population is heterozygous for the allele. The Molecular Pathology of Hb S. In the 1950s, Vernon Ingram discovered that the abnormality in sickle cell hemoglobin was a replacement of one of the 146 amino acids in the β chain of the hemoglobin molecule. All the clinical manifestations of sickle cell hemoglobin are consequences of this single change in the β-­globin gene. Ingram’s discovery was the first demonstration in any organism that a variant in a structural gene could cause an amino acid substitution in the corresponding protein. Because the substitution is in the β-­globin chain, the formula for sickle cell hemoglobin is written as α2β2 S or, more precisely, α2Aβ2 S. A heterozygote has a mixture of the two types of hemoglobin, A and S, summarized as α2Aβ2A/­α2Aβ2 S, as well as a hybrid hemoglobin tetramer, written as α2Aβ Aβ S. Strong evidence indicates that the sickle cell variant arose in West Africa, but that it occurred independently elsewhere. The β S allele has attained high frequency in malaria endemic areas of the world because it confers protection against malaria in heterozygotes (see Chapter 10). Sickling and Its Consequences. The molecular and cellular pathology of sickle cell disease is summarized in Fig. 12.6. Hemoglobin molecules containing the altered β-­globin subunits are normal in their ability to perform their principal function of binding oxygen (provided they have not polymerized, as described next), but in deoxygenated blood they are only one-­fifth as soluble as normal hemoglobin. Under conditions of low oxygen tension, this relative insolubility of deoxyhemoglobin S causes the sickle hemoglobin molecules to aggregate in the form of rod-­shaped polymers or fibers (see Fig. 12.6). These molecular rods distort the α2β2 S erythrocytes to a sickle shape that prevents them from squeezing single file through capillaries—as do normal red cells, thereby blocking blood flow and causing local ischemia. They may also cause disruption of the red cell membrane TABLE 12.3 Major Classes of Hemoglobin Structural Alterations Class Examplea Amino Acid Substitution Pathophysiological Effect of Variant Inheritance Hb S β chain: p. Glu 6Val Deoxygenated Hb S polymerizes → sickle cells → vascular occlusion and hemolysis AR Hb Hammersmith β chain: p. Phe 42Ser An unstable Hb → Hb precipitation → hemolysis; also low oxygen affinity AD Hb M-­Hyde Park β chain: p. His 92Tyr The substitution makes oxidized heme iron resistant to methemoglobin reductase → Hb M, which cannot carry oxygen → cyanosis (asymptomatic) AD Hb Kempsey β chain: p. Asp 99Asn The substitution keeps the Hb in its high oxygen affinity structure → less oxygen to tissues → polycythemia AD Hb E β chain: p. Glu 26Lys The variant → an abnormal Hb and decreased synthesis (abnormal RNA splicing) → mild thalassemiab (see Fig. 12.11) AR a Hemoglobin variants are often named after a location related to the first described patient(s). b Additional β-­chain gene variants that cause β-­thalassemia are depicted in Table 12.5. AD, Autosomal dominant; AR, autosomal recessive; Hb M, methemoglobin (See text.)
238 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE into the following three classes, depending on the clinical phenotype (Table 12.3): Alterations that cause hemolytic anemia, most commonly b...
Ch12 · Pt9 CHAPTER 12 — The Molecular Basis of Genetic Disease 239 (hemolysis) and release of free hemoglobin, which can have deleterious effects on the availability of vasodilators, such as nitric oxide, thereby exacerbating the ischemia. Modifier Genes Determine the Clinical Severity of Sickle Cell Disease. It has long been known that a strong modifier of the clinical severity of sickle cell disease is the patient’s level of Hb F (α2γ2), higher levels being associated with less morbidity and lower mortality. The physiologic basis of the ameliorating effect of Hb F is clear: Hb F is a perfectly adequate oxygen carrier in postnatal life and inhibits the polymerization of deoxyhemoglobin S. Until recently, however, it was not certain whether the variation in Hb F expression was heritable. Genome-­ wide association studies (GWAS) (see Chapter 11) have demonstrated that single nucleotide variants (SNPs) at three polymorphic loci (SNPs)—­the γ-­globin gene and two genes that encode transcription factors, BCL11A and MYB—­account for 40 to 50% of the variation in the levels of Hb F in individuals with sickle cell disease. Moreover, the Hb F–­associated SNPs are associated with the painful clinical episodes thought to be due to capillary occlusion caused by sickled red cells (see Fig. 12.6). Individuals with heterozygous loss-offunction variants in BCL11A (gene) have a rare neurogenetic disorder but also hereditary persistence of fetal hemoglobin. The genetically driven variations in the level of Hb F are also associated with variation in the clinical severity of β-­thalassemia (discussed later) because the reduced abundance of β-­globin (and thus of Hb A [α2β2]) in that disease is partly alleviated by higher levels of γ-­globin and, thus, of Hb F (α2γ2). The discovery of these genetic modifiers of Hb F abundance not only explains much of the variation in the clinical severity of sickle cell disease and β-­thalassemia, but highlights a general principle introduced in Chapter 9: modifier genes can play a major role in determining the clinical and physiologic severity of a single-­gene disorder. BCL11A, a Silencer of γ-­Globin Gene Expression in Adult Erythroid Cells. The identification of genetic modifiers of Hb F levels, particularly BCL11A, has opened great therapeutic potential. The product of the BCL11A gene is a transcription factor that normally silences γ-­globin expression, thus shutting down Hb F production postnatally. Accordingly, drugs that suppress BCL11A activity postnatally, thereby increasing the expression of Hb F, might be of great benefit to those with sickle cell disease and β-­thalassemia (see Chapter 14). In addition, preliminary clinical trial data suggest that post-transcriptional genetic silencing of BCL11A may be an effective treatment for sickle cell disease. Trisomy 13, Micro RNAs, and MYB—Another Silencer of γ-­Globin Gene Expression. The indication from GWAS that MYB is an important regulator of γ-­globin expression has received further support from an unexpected direction: studies investigating the basis for the persistent increased postnatal expression of Hb F that is observed in individuals with trisomy 13 (see Chapter 6). Two miRNAs, mi R-­15a and mi R-­ 16-­1, directly target the 3′ untranslated region (UTR) of the MYB mRNA, thereby reducing MYB expression. The genes for these two miRNAs are located on chromosome 13; their extra dosage in trisomy 13 is predicted to reduce MYB expression to below normal levels, thereby partly relaxing the postnatal suppression of γ-­globin gene expression normally mediated by the MYB protein. This leads to increased expression of Hb F (Fig. 12.7). Unstable Hemoglobins. The unstable hemoglobins are due largely to single nucleotide variants that cause denaturation of the hemoglobin tetramer in mature red blood cells. The denatured globin tetramers are insoluble and precipitate to form inclusions (Heinz bodies) that damage the red cell membrane and cause hemolysis of mature red blood cells in the vascular tree (Fig. 12.8, showing a Heinz body due to β-­thalassemia). Normal codon Sickle cell codon GAG GTG Amino acid substitution β6 Glu Val Hb S Cell heterogeneity Oxy Deoxy Hb S solution Hb S fiber Vaso-occlusion Figure 12.6 The pathogenesis of sickle cell disease. (Redrawn from Ingram V: Sickle cell disease: molecular and cellular pathogenesis. In Bunn HF, Forget BG, editors: Hemoglobin: Molecular, genetic, and clinical aspects, Philadelphia, 1986, WB Saunders.)
CHAPTER 12 — The Molecular Basis of Genetic Disease 239 (hemolysis) and release of free hemoglobin, which can have deleterious effects on the availability of vasodilators, such as nitric oxide, thereb...
Ch12 · Pt10 240 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE The amino acid substitution in the unstable hemoglobin, Hb Hammersmith (β-­chain p. Phe 42Ser; see Table 12.3), leads to denaturation of the tetramer and consequent hemolysis. This variant is notable because the substituted phenylalanine residue is one of the two amino acids that are conserved in all globins in nature (see Fig. 12.2). It is, therefore, not surprising that substitutions of this phenylalanine produce serious alterations in hemoglobin function. In normal β-­globin, the bulky phenylalanine wedges the heme into a “pocket” in the folded β-­globin monomer. Its replacement by serine, a smaller residue, creates a gap that allows the heme to slip out of its pocket. In addition to its instability, Hb Hammersmith has a low oxygen affinity, which can cause cyanosis in heterozygotes carriers. In contrast to variants that destabilize the tetramer, other variants destabilize the globin monomer and never form the tetramer, causing chain imbalance and thalassemia (see following section). Variants With Altered Oxygen Transport Variants that alter the ability of hemoglobin to transport oxygen, although rare, are of general interest because they illustrate how a variant can impair one function of a protein (in this case, oxygen binding and release) and yet leave the other properties of the protein relatively intact. For example, the variants that affect oxygen transport generally have little or no effect on hemoglobin stability. Methemoglobins. Oxyhemoglobin is the form of hemoglobin that is capable of reversible oxygenation; its heme iron is in the reduced (or ferrous) state. The heme iron tends to oxidize spontaneously to the ferric Euploid erythroid progenitor Micro RNAs 15a and 16-1 MYB MYB Fetal hemoglobin Midgestation Birth Trisomy 13 progenitor Figure 12.7 A model demonstrating how elevations of micro RNAs 15a and 16-­1 in trisomy 13 can result in elevated fetal hemoglobin expression. The basal level of these micro RNAs moderates expression of targets such as the MYB gene during erythropoiesis. In the case of trisomy 13, elevated levels of these micro RNAs result in additional down-­regulation of MYB expression, which in turn results in a delayed switch from fetal to adult hemoglobin and persistent expression of fetal hemoglobin. (Redrawn from Orkin SH: Disorders of hemoglobin synthesis: The thalassemias. In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders, pp. 106–­126.) A B C Figure 12.8 Visualization of one pathologic effect of the deficiency of β chains in β-­thalassemia: the precipitation of the excess normal α chains to form a Heinz body in the red blood cell. Peripheral blood smear and Heinz body preparation. The peripheral smear (A) shows “bite” cells with pitted-­out semicircular areas of the red blood cell membrane as a result of removal of Heinz bodies by macrophages in the spleen, causing premature destruction of the red cell. The Heinz body preparation (B) shows increased Heinz bodies in the same specimen when compared to a control (C). (From Hoffman R, Furie B, Mc Glave P, et al: Hematology: Basic principles and practice, ed 5, 2008, Elsevier.)
240 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE The amino acid substitution in the unstable hemoglobin, Hb Hammersmith (β-­chain p. Phe 42Ser; see Table 12.3), leads to denaturation of the...
Ch12 · Pt11 CHAPTER 12 — The Molecular Basis of Genetic Disease 241 form and the resulting molecule—referred to as methemoglobin, is incapable of reversible oxygenation. If significant amounts of methemoglobin accumulate in the blood, cyanosis results. Maintenance of the heme iron in the reduced state is the role of the enzyme, methemoglobin reductase. In several altered globins (either α or β), substitutions in the region of the heme pocket affect the heme-­globin bond in a way that makes the iron resistant to the reductase. Although heterozygotes for these abnormal hemoglobins are cyanotic (a sign), they are asymptomatic. The homozygous state is presumably lethal. One example of a β-­chain methemoglobin is Hb Hyde Park (see Table 12.3), in which the conserved histidine (p. His 92 in Fig. 12.2) to which heme is covalently bound has been replaced by tyrosine (p. His 92Tyr). Hemoglobins With Altered Oxygen Affinity. Variants that alter oxygen affinity demonstrate the importance of subunit interaction for the normal function of a multimeric protein such as hemoglobin. In the Hb A tetramer, the α:β interface has been highly conserved throughout evolution. It is subject to significant movement between the chains when the hemoglobin shifts from the oxygenated (relaxed) to the deoxygenated (tense) form of the molecule. Substitutions in residues at this interface, exemplified by the β-­globin mutant Hb Kempsey (see Table 12.3), prevent the normal oxygen-­related movement between the chains; the variant “locks” the hemoglobin into the high oxygen affinity state, reducing oxygen delivery to tissues and causing polycythemia. Thalassemia: An Imbalance of Globin-­Chain Synthesis The thalassemias are collectively the most common human single-­gene disorders in the world (Case 44). They are a heterogeneous group of diseases of hemoglobin synthesis in which variants reduce the synthesis or stability of either the α-­globin or β-­globin chain to cause α-­thalassemia or β-­thalassemia, respectively. The resulting imbalance in the ratio of the α:β chains underlies the pathophysiology. The chain that is produced at the normal rate is in relative excess; in the absence of a complementary chain with which to form a tetramer, the excess normal chains eventually precipitate in the cell, damaging the membrane and leading to premature red blood cell destruction. The excess β or β-­like chains are insoluble and precipitate in both red cell precursors (causing ineffective erythropoiesis) and in mature red cells (causing hemolysis) because they damage the cell membrane. The result is anemia (lack of red blood cells) in which the red cells are both hypochromic (i.e., pale red cells) and microcytic (i.e., small red cells). The name thalassemia (from the Greek thalassa, sea) was first used to signify that the disease was discovered in persons of Mediterranean origin. Both α-­thalassemia and β-­thalassemia, however, have a high frequency in many populations; α-­thalassemia is more prevalent and more widely distributed. The high frequency of thalassemia is due to the protective advantage against malaria that it confers on carriers, analogous to the heterozygote advantage of sickle cell hemoglobin carriers (see Chapter 10). There is a characteristic distribution of the thalassemias in a band around the Old World—­in the Mediterranean, the Middle East, and parts of Africa, India, and Asia. An important clinical consideration is that alleles for both types of thalassemia, as well as for structural alterations in hemoglobin, do often coexist in an individual. As a result, interactions may occur among different alleles of the same globin gene, or among variant alleles of different globin genes. The α-­Thalassemias Genetic disorders of α-­globin production disrupt the formation of both fetal and adult hemoglobins (see Fig. 12.3), causing intrauterine as well as postnatal disease. In the absence of α-­globin chains with which to associate, the chains from the β-­globin cluster are free to form a homotetrameric hemoglobin. Hemoglobin with a γ4 composition is known as Hb Barts, and the β4 tetramer is called Hb H. Because neither of these hemoglobins is capable of releasing oxygen to tissues under normal conditions, they are completely ineffective oxygen carriers. Consequently, fetuses with severe α-­thalassemia and high levels of Hb Barts suffer severe intrauterine hypoxia and develop massive generalized fluid accumulation: a condition called hydrops fetalis. In milder α-­thalassemias, an anemia develops because of the gradual precipitation of the Hb H in the erythrocyte. The formation of Hb H inclusions in mature red cells and the removal of these inclusions by the spleen damages the cells, leading to their premature destruction. Deletions of the α-­Globin Genes. The most common molecular causes of α-­thalassemia are gene deletions. The high frequency of deletions in the α-­chain genes, compared with, for example, the β-­chain genes, is a consequence of having two identical α-­globin genes on each chromosome 16 (see Fig. 12.3A). This arrangement of tandem homologous α-­globin genes facilitates misalignment, due to homologous pairing and subsequent recombination between the α1 gene domain on one chromosome and the corresponding α2 gene region on the other (Fig. 12.9). Evidence supporting this pathogenic mechanism of nonallelic homologous recombination is provided by reports of individuals with a triplicated α-­globin gene complex. Deletions or other alterations of one, two, three, or all four copies of the α-­globin gene cause a proportionately severe hematologic abnormality (Table 12.4). Individuals with two normal and two abnormal α-­globin genes are said to have α-­thalassemia trait.
.
Ch12 · Pt12 242 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE This can result from either of two genotypes (−−/­αα or −α/­−α), differing in whether the deletions are in cis or in trans. The α-­thalassemia trait is distributed throughout the world. However, heterozygosity for deletion of both copies of the α-­globin gene in cis (−−/­αα genotype) is largely restricted to Southeast Asians. Offspring of two carriers of this deletion allele may receive two −−/­ −− chromosomes, leading to Hb Barts (γ4) and hydrops fetalis. In other populations, however, α-­thalassemia trait is usually the result of the trans −α/­−α genotype, which cannot give rise to −−/­−− offspring. In addition to α-­thalassemia variants that result in deletion of the α-­globin genes, variants that delete only the LCR of the α-­globin complex can cause α-­thalassemia. In fact, similar to the observations with respect to the β-­globin LCR, such deletions were critical for demonstrating the existence of this regulatory element at the α-­globin locus. Other Forms of α-­Thalassemia. In all the classes of α-­thalassemia described earlier, deletions in the α-­globin genes or variants in their cis-­acting sequences account for the reduction of α-­globin synthesis. Other types of α-­thalassemia occur much less commonly. One important rare form of α-­thalassemia is ATR-­X syndrome, which is associated with both α-­thalassemia and intellectual disability. It illustrates the importance of epigenetic packaging of the genome in the regulation of gene expression (see Chapters 3 and 8). The X chromosome ATRX gene encodes a chromatin remodeling protein that functions, in trans, to activate the expression of the α-­globin genes. The ATRX protein belongs to a family of proteins that function within large multiprotein complexes to change DNA topology. ATR-­X syndrome is one of several monogenic diseases that result from variants in chromatin remodeling proteins (chromatinopathies; see Chapter 8). ATR-­X syndrome was initially recognized as unusual because the first families in which it was identified were northern Europeans, a population in which the deletion forms of α-­thalassemia are uncommon. All affected individuals were males with severe intellectual disability, together with a wide range of other abnormalities, including characteristic facial features, skeletal defects, and urogenital malformations. This diversity of phenotypes suggests that ATRX regulates the expression of numerous other genes besides the α-­globins. In those with ATR-­X syndrome, the reduction in α-­globin synthesis is due to accumulation at the α-­globin gene cluster of a histone variant (see Chapter 3) called macro H2A. This accumulation reduces α-­globin gene expression and causes α-­thalassemia. To date, all variants in the ATRX gene associated with ATR-­X syndrome involve partial loss of function, leading to hematologic defects that are mild, compared with those seen in the classic forms of α-­thalassemia. Individuals with ATR-­X syndrome have abnormalities in DNA methylation patterns to indicate that the Single-gene complex Triple-gene complex Homologous pairing and unequal crossover ψα1 α2 α1 ψα1 α2 α1 ψα1 α2 α1 α ψα1 α Figure 12.9 The probable mechanism underlying the most common form of α-­thalassemia, which is due to deletions of one of the two α-­globin genes on a chromosome 16. Misalignment, homologous pairing, and recombination between the α1 gene on one chromosome and the α2 gene on the homologous chromosome result in the deletion of one α-­globin gene. (Redrawn from Kazazian HH: The thalassemia syndromes: Molecular basis and prenatal diagnosis in 1990, Semin Hematol 27:209–­228, 1990.) TABLE 12.4 Clinical States Associated With α-­Thalassemia Genotypes Clinical Condition Number of Functional α Genes α-­Globin Gene Genotype α-­Chain Production “Normal” 4 αα/­αα 100% Silent carrier 3 αα/­α− 75% α-­Thalassemia trait (mild anemia, microcytosis) 2 α−/­α− or αα/­−− 50% Hb H (β4) disease (moderately severe hemolytic anemia) 1 α−/­−− 25% Hydrops fetalis or homozygous α-­thalassemia (Hb Barts: γ4) 0 −−/­−− 0%
242 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE This can result from either of two genotypes (−−/­αα or −α/­−α), differing in whether the deletions are in cis or in trans. The α-­thalassem...
Ch12 · Pt13 CHAPTER 12 — The Molecular Basis of Genetic Disease 243 ATRX protein is also required to establish or maintain the methylation pattern in certain domains of the genome. This may be by modulating access of the DNA methyltransferase enzyme to its binding sites. This finding is noteworthy because variants in MECP2, which encodes a protein that binds to methylated DNA, cause Rett syndrome (Case 40) by disrupting the epigenetic regulation of genes in regions of methylated DNA, leading to neurodevelopmental regression. Normally, ATRX and the Me CP2 protein interact; impairment of this interaction due to ATRX variants may contribute to the intellectual disability seen in ATR-­X syndrome. The β-­Thalassemias The β-­thalassemias share many features with α-­thalassemia. In β-­thalassemia, the decrease in β-­globin production causes a hypochromic, microcytic anemia. An imbalance in globin synthesis is due to the excess of α chains. The latter are insoluble and precipitate (see Fig. 12.8) in both red cell precursors (causing ineffective erythropoiesis) and mature red cells (causing hemolysis) because they damage the cell membrane. In contrast to α-­globin, however, the β chain is important only in the postnatal period. Consequently, the onset of β-­thalassemia is not apparent until a few months after birth, when β-­globin normally replaces γ-­globin as the major non–­α chain (see Fig. 12.3B). Only the synthesis of the major adult hemoglobin, Hb A, is reduced. The level of Hb F is increased in β-­thalassemia, not because of a reactivation of the γ-­globin gene expression that was switched off at birth, but because of selective survival and perhaps increased production of the minor population of adult red blood cells that contain Hb F. In contrast to α-­thalassemia, the β-­thalassemias are usually due to single nucleotide variants rather than to deletions (Table 12.5). In many regions of the world where β-­thalassemia is common, there are so many different β-­thalassemia variants that individuals with this condition are more likely to be compound heterozygotes (i.e., carrying two different β-­thalassemia alleles) than to be homozygotes for one allele. Most individuals with two β-­thalassemia alleles have thalassemia major: a condition characterized by severe anemia and the need for lifelong medical management. When the β-­thalassemia alleles allow so little production of β-­globin that no Hb A is present, the condition is designated β0-­thalassemia. Clinically, affected individuals are dependent on red blood cell transfusions. If some Hb A is detectable, the affected individual has β+-­thalassemia. Although the severity of the clinical disease depends on the combined effect of the two alleles present, until recently, survival into adult life was unusual. Infants with homozygous β-­thalassemia present with anemia once the postnatal production of Hb F decreases—generally before 2 years of age. At present, in most countries, treatment of the thalassemias is based on correction of the anemia and the increased marrow expansion by blood transfusion; the consequent excess iron accumulation is controlled by administration of chelating agents. Bone marrow transplantation is effective, but this is an option only if an HLA-­matched family member can be found. Gene therapy options are now emerging in clinical practice (see Chapter 14). Carriers of one β-­thalassemia allele are clinically well and are said to have thalassemia minor. Such individuals have hypochromic, microcytic red blood cells TABLE 12.5 The Molecular Basis of Some Causes of Simple β-­Thalassemia Type Example Phenotype RNA splicing defects (see Fig. 12.11C) Abnormal acceptor site of intron 1: AG → GG β0 Promoter variants Variant in the ATA box β+ −31 −30 −29 −28 −31 −30 −29 −28 A T A A → G T A A Abnormal RNA cap site A → C transversion at the mRNA cap site β+ Polyadenylation signal defects AATAAA → AACAAA β+ Nonsense variants Codon 39 gln → stop β0 CAG → UAG Codon 16 (1 bp deletion) Frameshift variants Normal trp gly lys val asn β0 15 16 17 18 19 UGG GGC AAG GUG AAC UGG GCA AGG UGA Variant trp ala arg stop Synonymous variants Codon 24 β+ gly → gly GGU → GGA Derived in part from Weatherall DJ, Clegg JB, Higgs DR, et al: The hemoglobinopathies. In Scriver CR, Beaudet AL, Sly WS, et al, editors: The metabolic and molecular bases of inherited disease, ed 7, New York, 1995, Mc Graw-­Hill, pp. 3417–­3484; Orkin SH: Disorders of hemoglobin synthesis: the thalassemias. In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders, pp. 106–­126. a One other hemoglobin structural variant that causes β-­thalassemia is shown in Table 12.3. mRNA, Messenger RNA.
CHAPTER 12 — The Molecular Basis of Genetic Disease 243 ATRX protein is also required to establish or maintain the methylation pattern in certain domains of the genome. This may be by modulating acces...
Ch12 · Pt14 244 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE and may have a slight anemia that can be misdiagnosed initially as iron deficiency. The diagnosis of thalassemia minor can be supported by hemoglobin electrophoresis, which generally reveals an increase in the level of Hb A2 (α2δ2). In many countries, thalassemia minor is sufficiently common to require diagnostic distinction from iron deficiency anemia and to be a frequent source of referral for prenatal diagnosis of affected homozygous fetuses (see Chapter 18). α-­Thalassemia Alleles as Modifier Genes of β-­Thalassemia. In human genetics, one of the best examples of a modifier gene comes from co-existence of β-­thalassemia and α-­thalassemia alleles in a population. In such populations, β-­thalassemia homozygotes may also inherit an α-­thalassemia allele. The clinical severity of the β-­thalassemia is sometimes ameliorated by the presence of the α-­thalassemia allele, which acts as a modifier. The imbalance of globin chain synthesis that occurs in β-­thalassemia (due to the relative excess of α chains) is reduced by the decrease in α-­chain production that results from the α-­thalassemia gene deletion. β-­Thalassemia, Complex Thalassemias, and Hereditary Persistence of Fetal Hemoglobin. Almost every type of DNA variant known to reduce the synthesis of an mRNA or protein has been identified as a cause of β-­thalassemia. The following overview of these genetic defects is, therefore, instructive about variant mechanisms in general, by describing the molecular basis of one of the most common and severe genetic diseases in the world. Variants of the β-­globin gene complex are separated into two broad groups with different clinical phenotypes. One group, which accounts for the great majority of patients, impairs the production of β-­globin alone and causes simple β-­thalassemia. The second group consists of large deletions that cause the complex thalassemias, in which the β-­globin gene is removed, as well as one or more of the other genes—­or the LCR—­in the β-­globin cluster. Finally, we are informed about the regulation of globin gene expression through some deletions within the β-­globin cluster that do not cause thalassemia, but rather, a benign phenotype termed the hereditary persistence of fetal hemoglobin (i.e., the persistence of γ-­globin gene expression throughout adult life). Molecular Basis of Simple β-­Thalassemia. Simple β-­thalassemia results from a remarkable diversity of molecular defects, predominantly single nucleotide variants, in the β-­globin gene (Fig. 12.10; see Table 12.5). Most variants causing simple β-­thalassemia lead to a decrease in the abundance of the β-­globin mRNA. These include promoter variants, RNA splicing variants (the most common), mRNA capping or tailing variants, and frameshift or nonsense variants that introduce premature termination codons within the coding region of the gene. A few hemoglobin structural alterations also impair processing of the β-­globin mRNA, as exemplified by Hb E (described later). RNA Splicing Variants. Most β-­thalassemia cases with a decreased abundance of β-­globin mRNA have abnormalities in RNA splicing. Dozens of defects of * * ** ** Transcription RNA splicing Cap site RNA cleavage Initiator codon Small deletion Unstable globin Nonsense codon Frameshift 100 bp 5' 3' 1 2 3 Figure 12.10 Representative point variants and small deletions that cause β-­thalassemia. Note the distribution of variants throughout the gene and that the variants affect virtually every process required for the production of normal β-­globin. More than 100 different β-­globin point variants are associated with simple β-­thalassemia.
.
Ch12 · Pt15 CHAPTER 12 — The Molecular Basis of Genetic Disease 245 this type have been described, and their combined clinical burden is substantial. These variants have acquired high visibility because their effects on splicing are often unexpectedly complex, and analysis of the altered mRNAs has contributed extensively to knowledge of the sequences critical to normal RNA processing (introduced in Chapter 3). The splice defects are separated into three groups (Fig. 12.11) depending on the region of the unprocessed RNA in which the variant is located. Splice junction variants include those at the canonical 5′ donor or 3′ acceptor splice junctions of the introns or in the consensus sequences surrounding the junctions. The critical nature of the conserved GT dinucleotide at the 5′ intron donor site and of the AG at the 3′ intron acceptor site (see Chapter 3) is demonstrated by the complete loss of normal splicing that results from variants in these dinucleotides (see Fig. 12.11B). Inactivation of the normal acceptor site elicits the use of other acceptor-­like sequences elsewhere in the RNA precursor molecule. These alternative sites are termed cryptic splice sites because they are not used by the splicing apparatus if the correct site is available. Cryptic donor or acceptor splice sites can be found in either exons or introns. Intronic variants enhance the use of a cryptic splice site by making it more similar or identical to the normal splice site. The activated cryptic site then competes with the normal site, with variable effectiveness. This reduces the abundance of the normal mRNA by decreasing splicing from the correct site, which remains perfectly intact (see Fig. 12.11C). Cryptic splice site variants are often leaky, which means that some use of the normal site occurs, producing a β+-­ thalassemia phenotype. Coding sequence changes that affect splicing result from variants in the open reading frame that activate a cryptic splice site in an exon, whether or not they also change the amino acid sequence (see Fig. 12.11D). For example, a mild form of β+-­thalassemia results from a variant in codon 24 (see Table 12.5) that activates a cryptic splice site but does not change the encoded amino acid (both GGT and GGA code for glycine [see Table 3.1]); this is an example of a synonymous variant that is not neutral in its effect. Nonfunctional mRNAs. Some mRNAs are nonfunctional and cannot direct the synthesis of a complete polypeptide, because the variant generates a premature stop codon, which prematurely terminates translation. Two β-­thalassemia variants near the amino terminus exemplify this effect (see Table 12.5). In one (p. Gln 39Ter), the failure in translation is due to a single nucleotide substitution that creates a nonsense variant. In the other, a frameshift variant results from a single base pair deletion early in the open reading frame, removing the first nucleotide from codon 16, which normally encodes glycine. In the reading frame that results, a premature stop codon is quickly encountered downstream, well before the normal termination signal. Because no β-­globin is made from these alleles, both types of nonfunctional mRNA variants cause β0-­thalassemia in the homozygous state. In some instances, frameshifts near the carboxyl terminus of the protein allow most of the mRNA to be translated normally or to produce elongated globin chains, resulting in a variant hemoglobin rather than null alleles. In addition to ablating the production of the β-­globin polypeptide, premature stop variants, including the two described earlier, often lead to reduced abundance of the abnormal mRNA; indeed, the mRNA may be undetectable. The mechanism underlying this phenomenon—called nonsense-­mediated mRNA decay, appears to be restricted to nonsense codons located more than 50 bp upstream of the final exon-­exon junction. Defects in Capping and Tailing of β-­Globin mRNA. Several β+-­thalassemia variants highlight the critical nature of post-transcriptional modifications of mRNAs. For example, the 3′ UTR of almost all mRNAs ends with a poly A sequence, and if this sequence is not added, the mRNA is unstable. As introduced in Chapter 3, polyadenylation of mRNA first requires enzymatic cleavage of the mRNA, which occurs in response to a signal for the cleavage site, AAUAAA, that is found near the 3′ end of most eukaryotic mRNAs. Individuals with a substitution that changes the signal sequence to AACAAA produce only a minor fraction of correctly polyadenylated β-­globin mRNA. Hemoglobin E: A Structurally Altered Hemoglobin With Thalassemia Phenotypes Hb E is probably the most common structurally abnormal hemoglobin in the world, occurring at high frequency in Southeast Asia, where there are at least 1 million homozygotes and 30 million heterozygotes. Hb E is a β-­globin variant (p. Glu 26Lys) that reduces the rate of synthesis of the abnormal β chain. It is another example of a coding sequence variant that impairs normal splicing by activating a cryptic splice site (see Fig. 12.11D). Although Hb E homozygotes are asymptomatic and only mildly anemic, individuals who are genetic compounds of Hb E and another β-­thalassemia allele have clinically relevant phenotypes that are largely determined by the severity of the other allele. Complex Thalassemias and the Hereditary Persistence of Fetal Hemoglobin As mentioned earlier, large deletions that cause the complex thalassemias remove the β-­globin gene plus one or more other genes—­or the LCR—­from the β-­globin cluster. Thus, affected individuals have reduced expression
CHAPTER 12 — The Molecular Basis of Genetic Disease 245 this type have been described, and their combined clinical burden is substantial. These variants have acquired high visibility because their eff...
Ch12 · Pt16 246 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Exon 1 Intron 1 Exon 2 Exon 3 Intron 2 Intron 2 donor site: GT Intron 2 acceptor site: AG Normal splicing pattern Intron 2 Intron 1 bp 110 β+ mutation in a cryptic acceptor site reduced use of unaffected normal site preferred use of mutant site Exon 1 Exon 2 Exon 3 β+ Mutation Consensus acceptor site Normal sequence CCTATTAG T YYYYNYAG G CCTATTGG T 90% 10% Normal splice site unaffected New splice site in intron Mutation creating a new splice acceptor site in an intron Exon 1 Exon 2 Intron 2 Exon 3 40% Hb E: Exon 1 mutation in a cryptic donor site reduced use of normal site moderate use of cryptic site New splice site, in a codon 60% Codon β+ Mutation Donor consensus Normal exon 1 sequence GGTGGTAAGGCC AAGGTAAGT GGTGGTGAGGCC 24 25 26 27 Hb E codon 26 GAG->AAG glu->lys Mutation enhancing a cryptic splice donor site in an exon Intron 2 Exon 1 Exon 2 Exon 3 Intron 2 cryptic acceptor site Consensus acceptor site 3' part of intron 2 Intron 2 acceptor site β0 mutation no splicing from the mutant site use of an intron 2 cryptic site TTTCTTTCAG G YYYYYYNYAG G Intron 2 Exon 3 Intron 2 Exon 3 β0 Mutation..... CGG CTC..... Normal:..... CAG CTC..... Mutation destroying a normal splice acceptor site and activating a cryptic site A B C D Figure 12.11 Examples of variants that disrupt normal splicing of the β-­globin gene to cause β-­thalassemia. (A) Normal splicing pattern. (B) An intron 2 variant (IVS2-­2A>G) in the normal splice acceptor site aborts normal splicing. This variant results in the use of a cryptic acceptor site in intron 2. The cryptic site conforms perfectly to the consensus acceptor splice sequence (where Y is either pyrimidine, T or C). Because exon 3 has been enlarged at its 5′ end by inclusion of intron 2 sequences, the abnormal alternatively spliced messenger RNA (mRNA) made from this mutant gene has lost the correct open reading frame and cannot encode β-­globin. (C) An intron 1 variant (G > A in nucleotide 110 of intron 1) activates a cryptic acceptor site by creating an AG dinucleotide and increasing the resemblance of the site to the consensus acceptor sequence. The globin mRNA thus formed is elongated (19 extra nucleotides) at the 5′ side of exon 2; a premature stop codon is introduced into the transcript. A β+ thalassemia phenotype results because the correct acceptor site is still used, although at only 10% of the wild-­type level. (D) In the Hb E defect, the missense variant (p. Glu 26Lys) in codon 26 in exon 1 activates a cryptic donor splice site in codon 25 that competes effectively with the normal donor site. Moderate use is made of this alternative splicing pathway, but the majority of RNA is still processed from the correct site, and mild β+ thalassemia results. (Modified from Stamatoyannopoulos G, Grosveld F: Hemoglobin switching. In Stamatoyannopoulos G, Majerus PW, Perlmutter RM, et al, editors: The molecular basis of blood diseases, ed 3, Philadelphia, 2001, WB Saunders.)
246 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Exon 1 Intron 1 Exon 2 Exon 3 Intron 2 Intron 2 donor site: GT Intron 2 acceptor site: AG Normal splicing pattern Intron 2 Intron 1 bp 110 β...
Ch12 · Pt17 CHAPTER 12 — The Molecular Basis of Genetic Disease 247 of β-­globin and one or more of the other β-­like chains. These disorders are named according to the genes deleted (e.g., [δβ]0-­thalassemia or [Aγδβ]0-­thalassemia) (Fig. 12.12). Deletions that remove the β-­globin LCR start ~50 to 100 kb upstream of the β-­globin gene cluster and extend 3′ to varying degrees. Although some of these deletions (such as the Hispanic deletion shown in Fig. 12.12) leave all or some of the genes at the β-­globin locus completely intact, they ablate expression from the entire cluster to cause (εγδβ)0-­thalassemia. Such variants demonstrate the total dependence of gene expression from the β-­globin gene cluster on the integrity of the LCR (see Fig. 12.4). A second group of large β-­globin gene cluster deletions of medical significance are those that leave at least one of the γ genes intact (such as the English deletion in Fig. 12.12). Individuals carrying such variants have one of two clinical manifestations, depending on the deletion: either δβ0-­thalassemia, or a benign condition called hereditary persistence of fetal hemoglobin (HPFH) that is due to disruption of the perinatal switch from γ-­globin to β-­globin synthesis. Homozygotes with either of these conditions are viable because the remaining γ gene(s) are still active after birth, instead of switching off as would normally occur. As a result, Hb F (α2γ2) synthesis continues postnatally at a high level and compensates for the absence of Hb A. The clinically innocuous nature of HPFH that results from the substantial production of γ chains is due to a higher level of Hb F in heterozygotes (17–­35% Hb F) than is generally seen in δβ0-­thalassemia heterozygotes (5–­18% Hb F). Because the deletions that cause δβ0-­ thalassemia overlap with those that cause HPFH (see Fig. 12.12), it is not clear why patients with HPFH have higher levels of γ gene expression. One possibility is that some HPFH deletions bring enhancers closer to the γ-­globin genes. Insight into the role of regulators of Hb F expression, such as BCL11A and MYB (see earlier discussion), has been partly derived from the study of individuals with complex deletions of the β-­globin gene cluster. For example, the study of several individuals with HPFH due to rare deletions of the β-­globin gene cluster identified a 3.5 kb region, near the 5′ end of the δ-­globin gene, that contains binding sites for BCL11A, the critical silencer of Hb F expression in the adult. Public Health Approaches to Preventing Thalassemia Large-­Scale Population Screening. The clinical severity of many forms of thalassemia, combined with their high frequency, imposes a tremendous health burden on many societies. To reduce the high incidence of the disease in some parts of the world, governments have introduced successful thalassemia control programs based on offering or requiring thalassemia carrier screening of individuals of childbearing age in the population (see Box 12.1). As a result of such programs, in many parts of the Mediterranean, the birth rate of affected newborns has been reduced by as much as 90%, through programs of education directed both to the general population and to health care providers. African American Indian Sicilian HPFH Turkish Thai (δβ)0 thalassemia German Italian (Αγδβ)0 thalassemia Hispanic English (εγδβ)0 thalassemia 5' HS ε δ β Aγ Gγ LCR –20 5' 3' –10 0 10 20 30 40 50 60 70 120 130 140 150 kb Chromosome 11p15 Figure 12.12 Location and size of deletions of various (εγδβ)0-thalassemia, (δβ)0-­thalassemia, (Aγδβ)0-thalassemia, and HPFH mutants. Note that deletions of the locus control region (LCR) abrogate the expression of all genes in the β-­globin cluster. The deletions responsible for δβ-­thalassemia, Aγδβ-­thalassemia, and HPFH overlap (see text). HPFH, Hereditary persistence of fetal hemoglobin; HS, hypersensitive sites.
CHAPTER 12 — The Molecular Basis of Genetic Disease 247 of β-­globin and one or more of the other β-­like chains. These disorders are named according to the genes deleted (e.g., [δβ]0-­thalassemia or...
Ch12 · Pt18 248 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Screening Restricted to Extended Families. The initiation of screening programs for thalassemia can be a major economic and logistical challenge. However, work in Pakistan and Saudi Arabia has demonstrated the effectiveness of a screening strategy that may be broadly applicable in countries where consanguineous marriages are common. In the Rawalpindi region of Pakistan, β-­thalassemia was found to be largely restricted to a specific group of families that came to attention because there was an identifiable index case (see Chapter 7). In 10 extended families with such an index case, testing of almost 600 persons established that ~8% of the married couples examined consisted of two carriers; outside of these 10 families, no couple at risk was identified among 350 randomly selected pregnant people and their partners. All carriers reported that the information provided was used to avoid further pregnancy if they already had two or more healthy children or, for couples with only one or no healthy children, for prenatal diagnosis. Although the long-­term impact of this program must be established, extended family screening of this type may contribute importantly to the control of recessive diseases in parts of the world where a cultural preference for consanguineous marriage is present. In other words, because of consanguinity, disease gene variants are trapped within extended families, so that an affected child indicates an extended family at high risk for the disease. The initiation of carrier testing and prenatal diagnosis programs for thalassemia requires not only the education of the public and of physicians but the establishment of skilled central laboratories and the consensus of the population to be screened (see Box). Whereas population-­wide programs to control thalassemia are inarguably less expensive than the cost of lifetime care for a large population of affected individuals, the temptation for governments or physicians to pressure individuals into accepting such programs must be avoided. The autonomy of the individual in reproductive decision making—a bedrock of modern bioethics, and the cultural and religious views of their communities must be respected. GENERAL REFERENCES Higgs DR, Engel JD, Stamatoyannopoulos G: Thalassaemia, Lancet 379:373–­383, 2012. Higgs DR, Gibbons RJ: The molecular basis of α-­thalassemia: a model for understanding human molecular genetics, Hematol Oncol Clin North Am 24:1033–­1054, 2010. Mc Cavit TL: Sickle cell disease, Pediatr Rev 33:195–­204, 2012. Roseff SD: Sickle cell disease: a review, Immunohematol 25:67–­74, 2009. Taher AT, Musallam KM, Cappellini MD: β-­thalassemias, N Engl J Med 384:727–­743, 2021. Weatherall DJ: The role of the inherited disorders of hemoglobin, the first “molecular diseases,” in the future of human genetics, Annu Rev Genomics Hum Genet 14:1–­24, 2013. REFERENCES FOR SPECIFIC TOPICS Bauer DE, Orkin SH: Update on fetal hemoglobin gene regulation in hemoglobinopathies, Curr Opin Pediatr 23:1–­8, 2011. Ingram VM: Gene mutations in human haemoglobin: the chemical difference between normal and sickle cell haemoglobin, Nature 180:326–­328, 1957. BOX 12.1 ETHICAL AND SOCIAL ISSUES RELATED TO POPULATION SCREENING FOR β-­THALASSEMIAa Worldwide, approximately 70,000 infants are born each year with β-­thalassemia, at high economic cost to health care systems and at great emotional cost to affected families. To identify individuals and families at increased risk for the disease, screening is done in many countries. National and international guidelines recommend that screening not be compulsory and that education and genetic counseling should inform decision making. Widely differing cultural, religious, economic, and social factors significantly influence the adherence to guidelines. For example: In Sardinia, a program initiated in 1975 involves voluntary screening, followed by testing of the extended family once a carrier is identified. In Greece, screening is voluntary, is available both premaritally and prenatally, requires informed consent, is widely advertised by the mass media and in military and school programs, and is accompanied by genetic counseling for carrier couples. In Iran and Turkey, these practices differ only in that screening is mandatory premaritally (but in all countries with mandatory screening, carrier couples have the right to marry if they wish). Major obstacles to more effective population screening for β-­thalassemia. The principal obstacles include the facts that pregnant individuals may feel overwhelmed by the array of tests offered to them, many health professionals have insufficient knowledge of genetic disorders, appropriate education and counseling are costly and time consuming, it is commonly misunderstood that informing an individual about a test is equivalent to obtaining consent, and the effectiveness of mass education varies greatly, depending on the community or country. The effectiveness of well-­executed β-­thalassemia screening programs. In populations where β-­thalassemia screening has been effectively implemented, the reduction in the incidence of the disease has been striking. For example, in Sardinia, screening between 1975 and 1995 reduced the incidence from 1 per 250 to 1 per 4000 individuals. Similarly, in Cyprus, the incidence of affected births fell from 51 in 1974 to none up to 2007. a Based on Cousens NE, Gaff CL, Metcalfe SA, et al: Carrier screening for β-­thalassaemia: A review of international practice, Eur J Hum Genet 18:1077–­1083, 2010.
248 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Screening Restricted to Extended Families. The initiation of screening programs for thalassemia can be a major economic and logistical chall...
Ch12 · Pt19 CHAPTER 12 — The Molecular Basis of Genetic Disease 249 Ingram VM: Specific chemical difference between the globins of normal human and sickle-­cell anaemia haemoglobin, Nature 178:792–­ 794, 1956. Kervestin S, Jacobson A: NMD, a multifaceted response to premature translational termination, Nat Rev Mol Cell Biol 13:700–­712, 2012. Pauling L, Itano HA, Singer SJ, et al: Sickle cell anemia, a molecular disease, Science 110:543–­548, 1949. Sankaran VG, Lettre G, Orkin SH, et al: Modifier genes in mendelian disorders: the example of hemoglobin disorders, Ann N Y Acad Sci 1214:47–­56, 2010. Steinberg MH, Sebastiani P: Genetic modifiers of sickle cell disease, Am J Hematol 87:795–­803, 2012. Weatherall DJ: The inherited diseases of hemoglobin are an emerging global health burden, Blood 115:4331–­4336, 2010. PROBLEMS 1. A newborn female dies of hydrops fetalis attributable to α-thalassemia. Draw a pedigree with genotypes illustrating to the biological parents the genetic basis of this disease. Explain why a Melanesian couple, whom they met in the hematology clinic and who both also have α-thalassemia trait, are unlikely to have a similarly affected child. 2. Why are most individuals with β-thalassemia compound heterozygotes for causal DNA variants in the β-globin gene? In what situation(s) might you anticipate that an individual with β-thalassemia would likely have two identical β-globin alleles (i.e., to be homozygous for the causal DNA variant)? 3. Tony, a young male of self-identified Italian ancestry, is found to havehas non-transfusion-dependent β-thalassemia, with a hemoglobin concentration of 7 g/d L (normal, amounts are 10 to 13 g/d L). When you perform a Northern blot of his reticulocyte RNA, you unexpectedly find three β-globin mRNA bands, one of normal size, one larger than normal, and one smaller than normal. What variant mechanism(s) could account for the presence of three bands like this observation in an individual with β-thalassemia? In this patient, the fact that the anemia is mild suggests that a significant fraction of normal β-globin mRNA is being made. What type(s) of variants would allow this to occur? 4. A man is heterozygous for Hb M Saskatoon, a hemoglobin missense structural alteration in which the normal amino acid His is replaced by Tyr at position 63 of the β chain. His mate is heterozygous for Hb M Boston, in which His is replaced by Tyr at position 58 of the α chain. Heterozygosity for either of these mutant variant alleles produces methemoglobinemia. Outline the possible genotypes and phenotypes of their offspring. 5. A child has a paternal uncle and a maternal aunt with sickle cell disease; both of her parents do not have sickle cell disease. What is the probability that the child has sickle cell disease? 6. A woman has sickle cell trait, and her mate is heterozygous for Hb C. What is the probability that their child has no abnormal hemoglobin? 7. Match the following: 8. Exome sequencing is organized for a child with unexplained intellectual disability, who also has non-transfusion-dependent β-thalassemia. Although this reveals a genetic cause for the intellectual disability, the molecular basis of the β-thalassemia phenotype is not elucidated, as only a single heterozygous pathogenic variant in the β-globin gene is identified. List possible explanations. 9. What are some possible explanations for the fact that thalassemia control programs, such as the successful one in Sardinia, have not reduced the birth rate of newborns with severe thalassemia to zero? For example, in Sardinia from 1999 to 2002, approximately two to five such infants were born each year.
___________ complex β-thalassemia ___________ β+-thalassemia ___________ number of α-globin genes missing in Hb H disease ___________ two different variant alleles at a locu ___________ ATR-X syn...
Select a segment to play