🧬 Part 2: Genetic Diversity and Clinical Cytogenetics English

← Back to Index
⚠️ For personal study use only. Commercial use is prohibited.
0%
0 / 33 listened Reset
Part 1 Part 2 Part 3 Part 4 Part 5 Part 6 Part 7 Part 8 Part 9 Part 10

Chapter 4: Human Genetic Diversity Genomic Variation

Ch4 · Pt1 chapter 4 Human Genetic Diversity Genomic Variation Stephen W. Scherer
Ada Hamosh The study of DNA variation is the conceptual cornerstone for genetics in medicine and for the broader field of human genetics. During the course of evolution, the steady influx of new varia...
Ch4 · Pt2 46 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE (rather than mutation). An exception is for the use of single nucleotide polymorphism (SNP) in the context of microarrays, where it is strongly entrenched in the lexicon. The Concept of Variation In this chapter we begin by exploring the nature of genomic variation, ranging from the change of a single nucleotide to alterations of an entire chromosome. To recognize a change means that there has to be a gold standard, compared to which the variant shows a difference. As we saw in Chapter 2, there is no single individual whose genome sequence could serve as such a definitive standard for the human species, and thus one arbitrarily designates the most common sequence or arrangement in a population at any one position in the genome as the so-­called reference sequence (see Fig. 2.6). As more and more genomes from individuals around the globe are sampled (and thus as more and more variation is detected among the currently 7.9 billion genomes that make up our species), this reference genome is subject to constant evaluation and change. Indeed, a number of international collaborations share and update data on the nature and frequency of DNA variation in different populations in the context of the reference human genome sequence and make the data available through publicly accessible databases that serve as essential resources for scientists, physicians, and other health care professionals (Table 4.1). As we learn more about variation and, in particular, as long-­read sequencing allows us to fill holes in the reference genome, updated genome TABLE 4.1 Useful Databases of Information on Human Genetic Diversity Description URL The Human Genome Project, completed in 2003, was an international collaboration to map and sequence the genome of our species. The draft sequence of the genome was released in 2001, and the “essentially complete” reference genome assembly was published in 2004. https://­www.genome.gov/­human-­genome-­project http://­genome.ucsc.edu/­cgi-­bin/­hg Gateway http://­www.ensembl.org/­Homo_­sapiens/­Info/­Index The Single Nucleotide Polymorphism Database (db SNP) and the Structural Variation Database (db Var) are databases of small-­scale and large-­scale variations, including single nucleotide variants, microsatellites, indels, and CNVs. ncbi.nlm.nih.gov/­snp/­ ncbi.nlm.nih.gov/­dbvar/­ The 1000 Genomes Project created a catalogue of common human genetic variation, using openly consented samples from people who declared themselves to be healthy. All data are publicly available. The International Genome Sample Resource (IGSR) maintains and shares the human genetic variation resource. www.internationalgenome.org The Genome Aggregation Database (gnom AD) reports variants from 125,748 exomes and 15,708 genomes (141,456 unrelated individuals) aligned on GRCh 37 in v 2.1 and 76,156 genomes from unrelated individuals aligned on GRCh 38 in v 3.0. gnomad.broadinstitute.org Clin Var is a freely accessible, public archive of reports of the relationships among human variants and phenotypes, with supporting evidence. www.ncbi.nlm.nih.gov/­clinvar The Human Gene Mutation Database is a comprehensive collection of published germline variants associated with or causing human inherited disease (currently including over >210,000 mutations in 8519 genes). www.hgmd.cf.ac.uk/­ac/­index.php The Database of Genomic Variants is a curated catalogue of structural variation in the human genome. As of 2023, the database contains over 8 million entries. dgv.tcag.ca CNV, Copy number variant; SNV, single nucleotide variant. Updated from Willard HF: The human genome: a window on human genetics, biology and medicine. In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine, ed 3, New York, 2016, Elsevier. builds are released by the human genome reference committee; the current reference is h GRC38. Because errors are corrected and new sequences added, it is very important to always specify the build used to annotate a genomic variant. Variants are sometimes classified by the size of the altered DNA sequence and, at other times, by the functional effect of the change on gene expression. Although classification by size is somewhat arbitrary, it can be helpful conceptually to recognize the spectrum of changes at three different levels: Variation in chromosome number that leaves chromosomes intact but changes the number of chromosomes in a cell (aneuploidy) Alterations that change only a portion of a chromosome and might involve an unbalanced change of a subchromosomal segment or a structural rearrangement involving parts of one or more chromosomes (regional variation or copy number variation [CNV]) Alterations of the sequence of DNA, involving the substitution, deletion, or insertion of DNA, range from an SNV through small repetitive units (such as trinucleotide repeats) and insertion-­deletion variants (indels) up to an arbitrarily set (and evolving) limit of approximately 1 kb where such a change becomes a CNV. The basis for and consequences of this third type of variation are the principal focus of this chapter, whereas both chromosome and regional variation will be presented at length in Chapters 5 and 6. The functional consequences of DNA mutations, even those that change a single base pair, run the gamut from being completely innocuous to causing serious
46 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE (rather than mutation). An exception is for the use of single nucleotide polymorphism (SNP) in the context of microarrays, where it is strong...
Ch4 · Pt3 CHAPTER 4 — Human Genetic Diversity 47 illness, all depending on the location, nature, and size of the resulting variant. For example, even a change within a coding exon of a gene may have no effect on how a gene is expressed if the change does not alter the primary amino acid sequence of the polypeptide product; even if it does, the resulting change in the encoded amino acid sequence may not alter the functional properties of the protein. Not all variants, therefore, manifest in a clinical phenotype, though they will be reflected as DNA sequence variants. The Concept of Common Variants The DNA sequence of a given region of the genome is remarkably similar among chromosomes carried by many different individuals from around the world. In fact, any randomly chosen segment of human DNA of ~1000 bp in length, on average, will differ by only one base pair between the homologous segments inherited from that individual’s parents (assuming the parents are unrelated). However, across all human populations, hundreds of millions of single nucleotide differences and over a million more complex variants have been identified and catalogued. Because of limited sampling, these figures are likely to underestimate the true extent of genetic diversity in our species. Many populations have yet to be adequately studied. Even in those that have been well studied, the number of individuals examined is too small to reveal most variants with minor allele frequencies below 1% to 2%. Thus, as more people are included in variant discovery projects, additional (and rarer) variants will certainly continue to be uncovered. Whether a variant is formally considered common or not depends entirely on whether its frequency in a population exceeds a certain threshold, such as 1% of the alleles in that population. It does not depend on what kind of mutation caused it, how large a segment of the genome is involved, or whether it has a demonstrable effect on the individual. Although most common sequence variants are located between genes or within introns and are most often inconsequential to the functioning of any gene, others may be located in the coding sequence of genes themselves and result in different protein variants that may lead in turn to distinctive differences in human populations. Still, others are in regulatory regions and may have important effects on transcription or RNA stability. One might expect that deleterious variants that cause rare monogenic diseases are unlikely to become considered common variants. Although it is true that the alleles responsible for most clearly inherited clinical conditions are rare, some alleles that have a profound effect on health—­such as alleles of genes encoding enzymes that metabolize drugs (e.g., sensitivity to abacavir in some individuals infected with human immunodeficiency virus) (Case 1), the sickle cell allele in African populations and others of African and Mediterranean ancestry (see Chapter 12) (Case 42), or the p. Phe 508del variant in CFTR that causes cystic fibrosis (see Chapter 13) (Case 12)—­are relatively common. Nonetheless, these are exceptions. As more and more genetic variation is discovered and catalogued, it is clear that the vast majority of variants in the genome—­whether common or rare—­reflect differences in DNA sequence that have no overt significance to health. Common variants are key elements for the study of human and medical genetics. The ability to distinguish different inherited forms of a gene or different segments of the genome provides critical tools for a wide array of applications, both in research and in clinical practice (see Box 4.1). BOX 4.1 INHERITED VARIATION IN HUMAN AND MEDICAL GENETICS Allelic variants can be used as markers for tracking the inheritance of the corresponding segment of the genome in families and in populations. Such variants can be used as follows: As powerful research tools for mapping a gene to a particular region of a chromosome by linkage analysis or by allelic association (see Chapter 11) For prenatal diagnosis of genetic disease and for detection of carriers of deleterious alleles (see Chapter 18) In blood banking and tissue typing for transfusions and organ transplantation In forensic applications such as identity testing for determining paternity, identifying remains of crime victims, or matching DNA from a crime investigation to that of a perpetrator To provide genomic-­based precision medicine (see Chapter 19), medical care is tailored, for example, to whether an individual carries variants that increase or decrease the risk for common adult disorders (such as coronary heart disease, cancer, and diabetes; see Chapter 9) or that influence the efficacy or safety of particular medications (see Chapter 19) INHERITED COMMON VARIATION IN DNA The original Human Genome Project and the subsequent study of many millions of individuals worldwide have provided vast DNA sequence information. With this information in hand, one can begin to characterize the types and frequencies of common variation found in the human genome and to generate catalogues of the world’s human DNA sequence diversity. Such variants can be classified according to how the DNA sequence differs among the different alleles (Table 4.2 and Figs. 4.1 and 4.2). Single Nucleotide Variants The simplest and most common of all variants are SNVs. Those that occur at a high population frequency (typically defined as >1% or >5%) have been called SNPs, but more recently, common SNVs. A polymorphic locus characterized by a common SNV usually has
CHAPTER 4 — Human Genetic Diversity 47 illness, all depending on the location, nature, and size of the resulting variant. For example, even a change within a coding exon of a gene may have no effect o...
Ch4 · Pt4 48 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE C – G – SNV Indel A Indel B G G G G A T T T T T C T C... A G C A T G C A... A A Allele 1 Allele 2 G G G G G G G G A A T T T T T T T T T C C T T C C...... A A G G C C A A T T G G C C A A...... A A A A Allele 1 Allele 2 G G G G G G G G A A T T T T T T T T C C T T C C...... A A G G C C A A T T G G C C A A...... A A A A Allele 1 Allele 2 G G G G G G G G A A T T T T T T T T T C C T C C A...... A A G A C T A C T G G C C T A G... A... A T A A Reference sequence 5 20 15 10 Figure 4.1 Three polymorphisms in genomic DNA from the segment of the human genome reference assembly shown at the top (see also Figure 2.6). The single nucleotide variation (SNV) at position 8 has two alleles, one with a T (corresponding to the reference sequence) and one with a C. There are two indels in this region. At indel A, allele 2 has an insertion of a G between positions 11 and 12 in the reference sequence (allele 1). At indel B, allele 2 has a 2 ­bp deletion of positions 5 and 6 in the reference sequence. A B C D E F G H ABCDEFGH ABCDEFGFGFGH ABCDEFGH ABEDCFGH Allele 1 Allele 2 Allele 3 G G G G G G C C C A A A A A A T T T A A A T T T T T T A A A C C C C C C A A A......... A A A G C C A A A A A A G A A A C G T A A A A G C A T T G A A T C C G A G A T T A C C C A A C T G T G... A C G G... T A C... A G C C C A A A Allele 1 Allele 2 Allele 1 Allele 2 Microsatellite polymorphism Mobile element insertion polymorphism Copy number variant LINE Allele 1 Allele 2 Inversion polymorphism Figure 4.2 Examples of variation in the human genome larger than single nucleotide variants. Clockwise from upper right: The microsatellite locus has three alleles, with four, five, or six copies of a CAA trinucleotide repeat. The inversion variant has two alleles corresponding to the two orientations (indicated by the arrows) of the genomic segment shown in green; such inversions can involve regions up to many megabases of DNA. Copy number variants involve deletion or duplication of hundreds of kilobase pairs to over a megabase of genomic DNA. In the example shown, allele 1 contains a single copy, whereas allele 2 contains three copies of the chromosomal segment containing the F and G genes; other possible alleles with zero, two, four, or more copies of F and G are not shown. The mobile element insertion variant has two alleles, one with and one without insertion of a ~6-­kb LINE repeated retroelement; the insertion of the mobile element changes the spacing between the two genes and may alter gene expression in the region. TABLE 4.2 Common Variation in the Human Genome Type of Variation Size Range (Approx.) Basis for the Variant Number of Alleles Single nucleotide variant 1 bp Substitution of one or another base pair at a particular location in the genome Usually 2 Insertion/­deletions (indels) 1 bp–­1 kb Simple: Presence or absence of a short segment of DNA 1–­1000 bp in length Microsatellites: Generally, a 2-­, 3-­, or 4-­nucleotide unit repeated in tandem 5–­25 times Simple: 2 Microsatellites: typically ≥5 Copy number variant 1 kb–­> ≅ 3 Mb Typically the presence or absence of 1-­kb to 1.5-­Mb segments of DNA, although tandem duplication of 2, 3, 4, or more copies can also occur ≥2 Inversions Few bp–­>1 Mb A DNA segment present in either of two orientations with respect to the surrounding DNA 2 bp, Base pair; kb, kilobase pair; Mb, megabase pair.
48 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE C – G – SNV Indel A Indel B G G G G A T T T T T C T C... A G C A T G C A... A A Allele 1 Allele 2 G G G G G G G G A A T T T T T T T T T C C T...
Ch4 · Pt5 CHAPTER 4 — Human Genetic Diversity 49 only two alleles, corresponding to two different bases at that particular location (see Fig. 4.1). Common SNVs are observed, on average, once every 1000 bp. However, their distribution is uneven around the genome; many more are found in noncoding parts of the genome, in introns and in sequences that are some distance from protein-­coding genes. Nonetheless, a significant number of SNVs, both common and rare, occur in genes and other known functional elements in the genome. Approximately half of these do not alter the predicted amino acid sequence of the encoded protein and thus are termed synonymous, whereas those that do alter the amino acid sequence are called nonsynonymous. Other SNVs are candidates to have significant functional consequences, as they introduce or change a stop codon (see Table 3.1), or alter a known splice site. The significance for health of the vast majority of common SNVs is unknown and is the subject of ongoing research. The fact that these variants are common does not mean that they are without detrimental or protective effect on health or longevity. What it does mean is that any effect of common SNVs is likely to involve a relatively subtle altering of disease susceptibility rather than be a direct cause of serious illness. Insertion-­Deletion Variants A second class of variants result from insertion or deletion (indels) of segments that range from a single base pair up to ~1 kb. Over a million indels have been described among human genomes, numbering in the hundreds of thousands for any one individual. Approximately half of all indels are referred to as simple because they have only two alleles—­that is, the presence or absence of the inserted or deleted segment (see Fig. 4.1). Microsatellite Variants Other indels, however, are multiallelic due to variable numbers of a segment of DNA in tandem at a particular location. The term satellite comes from the early observation that this fraction of DNA has a different density, causing separation during centrifugation. Sometimes called variable number of tandem repeats, these microsatellites are highly vulnerable to mutation. They consist of DNA cassettes composed of units of several nucleotides—­such as TG, CAA, or AAAT—­ repeated in tandem between one and a few dozen times (see Fig. 4.2). The numbers of repeated units determine the different alleles, sometimes also referred to as short tandem repeats (STRs). A microsatellite locus often has many alleles (repeat lengths) that can be rapidly evaluated by standard laboratory procedures to distinguish different individuals and to infer familial relationships (Fig. 4.3). Many tens of thousands of microsatellite loci are known throughout the human genome. Microsatellites are particularly useful for genetic mapping. Determining the alleles at multiple microsatellite loci is currently the method of choice for DNA fingerprinting used for identity testing. For example, the US Federal Bureau of Investigation (FBI) currently uses 20 STRs for its DNA fingerprinting panel. Two individuals (other than monozygotic twins) are so unlikely to have exactly the same alleles at all 20 loci that the panel will allow effectively definitive determination of whether samples came from the same individual. The information is stored in the FBI’s Combined DNA Index System (CODIS). Mobile Element Insertion Variants Nearly half of the human genome consists of dispersed families of repetitive elements (see Chapter 2). Although most of the copies of these repeats are stationary, some of them are mobile and contribute to human genetic diversity through the process of retrotransposition. As introduced in Chapter 3 in the context of processed pseudogenes, this involves transcription into an RNA, reverse transcription into a DNA sequence, and insertion (i.e., transposition) into another site in the genome. The two most common mobile element families are the Alu and long interspersed nuclear elements (LINE) families of repeats, and nearly 10,000 mobile element insertion variants have been Unrelated Individuals Allele length Family Members Mother Father Child 1 Child 2 Child 3 7 6 5 4 3 2 1 Figure 4.3 A schematic of a hypothetical microsatellite marker in human DNA. The different-­sized alleles (numbered 1–­7) correspond to fragments of genomic DNA containing different numbers of copies of a microsatellite repeat, and their relative lengths are determined by separating them by gel electrophoresis. The shortest allele (allele 1) migrates toward the bottom of the gel, whereas the longest allele (allele 7) remains closest to the top. Left, For this multiallelic microsatellite, each of the six unrelated individuals has two different alleles. Right, Within a family, the inheritance of alleles can be followed from each parent to each of the three children.
CHAPTER 4 — Human Genetic Diversity 49 only two alleles, corresponding to two different bases at that particular location (see Fig. 4.1). Common SNVs are observed, on average, once every 1000 bp. Howe...
Ch4 · Pt6 50 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE described in different populations. Each polymorphic locus consists of two alleles, one with and one without the inserted mobile element (see Fig. 4.2). Mobile element variants are found on all human chromosomes; although most are found in nongenic regions, a small proportion of them are found within genes. For many of these loci the insertion allele has a frequency of greater than 10% in various populations. Copy Number Variants Another important type of human polymorphism includes CNVs, which are conceptually related to indels and microsatellites but involve larger segments of the genome, operationally defined as from 1000 bp to ~3 million bp (i.e., the span between limits of sequencing detection and cytogenetic analysis, respectively). In the general population, variants larger than 500 kb are found in 5% to 10% of individuals, and those encompassing more than 1 Mb in 1% to 2%. The largest CNVs are sometimes in regions of the genome characterized by repeated blocks of homologous sequences called segmental duplications (or segdups). The importance of these regions in mediating duplication and deletion of the corresponding segments is discussed further in Chapter 6 in the context of various chromosomal syndromes. As with indels, smaller CNVs may have only two alleles (i.e., the presence or absence of a segment). Some large CNVs have multiple alleles due to the presence of different numbers of tandem copies of a DNA segment (see Fig. 4.2). In terms of genome diversity, the amount of DNA involved in CNVs vastly exceeds the amount that differs because of SNVs. Compared to the reference genome, the content of any given individual’s genome can differ by as much as 30 Mb because of copy number and indel differences. Notably, since their variable segments can include from one to several dozen genes, CNV loci are frequently implicated in traits that involve altered gene dosage. When a CNV is frequent enough, it represents a background of common variation that must be understood to properly interpret alterations in copy number for medical purposes. As with all DNA variation, the significance of different CNV alleles in health and disease susceptibility is the subject of intensive investigation. Inversions A final group of structural variants is inversions. These regions of the genome, from a few base pairs up to several Mb, are found in either of two orientations (see Fig. 4.2). Most inversions are characterized by regions of sequence homology at the edges of the inverted segment, implicating a process of homologous recombination in their origin. Regardless of orientation, an inversion that does not involve a gain or loss of DNA is balanced. Some can achieve substantial frequencies in the general population. However, anomalous recombination can result in the duplication or deletion of DNA located between the regions of homology—­a process associated with clinical disorders that we will explore further in Chapters 5 and 6. THE ORIGIN AND FREQUENCY OF DIFFERENT TYPES OF MUTATION Along the spectrum of diversity from rare to common variants, the different kinds of mutation occur in the context of such fundamental processes of cell division as DNA replication, DNA repair, DNA recombination, and chromosome segregation in mitosis or meiosis. The frequency of mutation per locus per cell division is a basic measure of how error prone these processes are, which is of fundamental importance for genome biology and evolution. However, of greatest importance to medical geneticists is the frequency of mutation per disease locus per generation, rather than the overall mutation rate across the genome per cell division. Measuring disease-­causing mutation rates can be difficult, however, because many mutations cause early embryonic lethality before the result can be recognized in a fetus or newborn. Further, some people with a disease-­causing variant may manifest the condition only late in life or may never show signs of the disease. Despite these limitations, we have made great progress in determining the overall frequency—­sometimes referred to as the genetic load—­of all mutations affecting the human species. These major types of mutation occur at appreciable frequencies in many different cells in the body. In the practice of genetics, we are principally concerned with inherited genome variation; however, all such variation had to originate as a new (de novo or spontaneous) change in a germ cell. From this unique start in the population, the ultimate frequency of each variant over time depends on chance and on the principles of inheritance and population genetics (see Chapter 10). Although the original mutation would have occurred only in the DNA of a cell in the germline, any progeny derived from that cell would then carry it as a constitutional variant in essentially all the cells of the body. In contrast, somatic mutations, depending on when they arise, occur in different proportions of cells throughout the body, but they cannot be transmitted to the next generation of individuals (unless they involved a germline cell). Given the rate of mutation (see later in this section), one would predict that every cell in an individual has a slightly different version of the genome, depending on the number of cell divisions that have occurred since conception. Such genomic heterogeneity is particularly likely to be apparent in highly proliferative tissues, such as intestinal epithelia or hematopoietic cells. However, most such variants are not typically detected because, in clinical testing, one usually sequences DNA from many millions of cells, among which the base sequence present at conception will predominate, and rare somatic mutations will be largely invisible and unascertained. Such variants,
50 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE described in different populations. Each polymorphic locus consists of two alleles, one with and one without the inserted mobile element (see...
Ch4 · Pt7 CHAPTER 4 — Human Genetic Diversity 51 however, can be of clinical importance in disorders associated with somatic mosaicism, caused by mutation in only a subset of cells in certain tissues (see Chapter 7). While somatic mutations will typically remain undetected within any multicell DNA sample, cancer provides the major exception. The mutational basis for the origins of cancer and the clonal nature of tumor evolution drive certain somatic changes to be present in essentially all the cells of a tumor. Indeed, 1000 to 10,000 somatic variants (and sometimes many more) are readily found in the genomes of most adult tumors, with mutation frequencies and patterns specific to different cancer types (see Chapter 16). Chromosome Alterations Events that produce a change in chromosome number because of chromosome missegregation are among the most common sources of variation seen in humans, with a rate of one event per 25 to 50 meiotic cell divisions. This estimate is clearly minimal because the developmental consequences of many such events are likely so severe that the resulting embryos are aborted spontaneously shortly after conception without being detected (see Chapters 5 and 6). Structural Variation Alterations affecting the structure or regional organization of chromosomes can arise in a number of different ways. Duplications, deletions, and inversions of a segment of a single chromosome are predominantly the result of homologous recombination between DNA segments with high sequence homology at more than one chromosomal site. Not all structural mutations are the result of homologous recombination, however. Others, such as chromosome translocations and some inversions, can occur at the sites of spontaneous double-­stranded DNA breaks. Once breakage occurs at two places anywhere in the genome, the two broken ends can be joined together, even without any obvious sequence homology between the two ends (a process termed nonhomologous end-­joining repair). Examples of such mutations will be discussed in Chapter 6. Mutation in Genes Gene or DNA variants, including base pair substitutions, insertions, and deletions (Fig. 4.4), can originate by either of two basic mutational mechanisms: errors introduced during DNA replication or arising from a etc. etc. C G A T G C G C T A T A A T C G A T C G C G A T T A C G C G C G A T G C T A T A T A A T C G A T C G C G A T T A C G C G C G A T G C T A T A G C A T C G A T C G C G A T T A C G C G C G A T G C T A T A T A G C A T C G A T C G C G T A C G C G Reference sequence Substitution Deletion Insertion T A G C Figure 4.4 Examples of mutations in a portion of a hypothetical gene with five codons shown (delimited by the dotted lines). The first base pair of the second codon in the reference sequence (shaded in blue) is mutated by a base substitution, deletion, or insertion. The base substitution of a G for the T at this position leads to a codon change (shaded in green) and, assuming that the upper strand is the sense or coding strand, a predicted nonsynonymous change from a serine to an alanine in the encoded protein (see genetic code in Table 3.1); all other codons remain unchanged. Both the single base pair deletion and insertion lead to a frameshift mutation in which the translational reading frame is altered for all subsequent codons (shaded in green), until a termination codon is reached.
CHAPTER 4 — Human Genetic Diversity 51 however, can be of clinical importance in disorders associated with somatic mosaicism, caused by mutation in only a subset of cells in certain tissues (see Chapt...
Ch4 · Pt8 52 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE failure to properly repair DNA after damage. Many such mutation events are spontaneous, arising during the normal (but imperfect) processes of DNA replication and repair, whereas others are induced by physical or chemical agents called mutagens. DNA Replication Errors Typically, the process of DNA replication (see Fig. 2.4) is highly accurate. Most replication errors (i.e., bases other than the complementary bases inserted into the double helix) are rapidly removed from the DNA and corrected by a series of DNA repair enzymes. A process termed DNA proofreading first recognizes which strand in the newly synthesized double helix contains the incorrect base and then replaces it with the proper complementary base. DNA replication needs to be a remarkably accurate process; otherwise, the burden of mutation on the organism and the species would be intolerable. The enzyme, DNA polymerase, faithfully duplicates the two strands of the double helix based on strict base-­pairing rules (A pairs with T, C with G) but errs about once in every 10 million bp. Additional proofreading then corrects more than 99.9% of these errors of DNA replication. Thus the overall mutation rate per base as a result of replication errors is a remarkably low 1 × 10−10 per cell division—­fewer than one mutation per genome per cell division. Repair of DNA Damage In addition to replication errors, about 10,000 to 1,000,000 nucleotides are damaged per human cell per day by (1) spontaneous chemical processes such as depurination, demethylation, or deamination, (2) reaction with chemical mutagens (natural or otherwise) in the environment, or (3) exposure to ultraviolet or ionizing radiation. Some but not all of this damage is repaired. Even if such damage is recognized and excised, the repair machinery may introduce incorrect bases. Thus in contrast to replication-­related DNA changes, which are usually corrected through proofreading mechanisms, nucleotide changes introduced by DNA damage and repair are often permanent. A particularly common spontaneous mutation is the substitution of T for C (or A for G on the other strand). The explanation for this observation comes from considering the major form of epigenetic modification in the human genome: DNA methylation, introduced in Chapter 3. Spontaneous deamination of 5-­methylcytosine to thymine (compare the structures of cytosine and thymine in Fig. 2.2) in the Cp G doublet gives rise to C to T or G to A mutations (depending on which strand the 5-­methylcytosine is deaminated). Such spontaneous mutations may not be recognized by the DNA repair machinery, thus becoming established in the genome after the next round of DNA replication. More than 30% of all single nucleotide substitutions are of this type, and they occur at a rate 25 times greater than those of other single nucleotide mutations. Thus the Cp G doublet represents a true hot spot for mutation in the human genome. Overall Rate of DNA Mutation The rate of DNA mutation at specific loci has been estimated using a variety of approaches. The impact of replication and repair errors on the occurrence of new variants throughout the genome can now be determined directly by whole genome sequencing (WGS), using trios consisting of a child and both parents, looking for new sequences in the child that are not present in either parent. The overall rate of new mutations, averaged between maternal and paternal gametes, is ~1.2 × 10−8 per base pair per generation. This rate, however, varies from gene to gene and perhaps from population to population, or even individual to individual. This rate of change, combined with considerations of population growth and dynamics, predicts that there must be an enormous number of relatively new (and thus very rare) variants among the current worldwide population of 7.9 billion individuals. As might be predicted, the vast majority of these will be single nucleotide variants in noncoding portions of the genome and will probably have little or no functional significance. Nonetheless, at the level of populations, the potential collective impact of these new mutation changes on genes of medical importance should not be overlooked. In the United States, for example, with over 4 million live births each year, ~6 million new changes will occur in coding sequences; thus even for a single protein-­coding gene of average size, we can anticipate several hundred newborns each year with a new variant in the coding sequence of that gene. Conceptually similar studies have determined the rate of mutation for CNVs, where the generation of a new length variant depends on recombination, rather than on errors in DNA synthesis. The measured rate of formation of new CNVs (≈1.2 × 10−2 per locus per generation) is orders of magnitude higher than that of base substitutions. Rate of Disease-­Causing Variations The most direct way of estimating the rate of disease-­ causing mutation, resulting in a pathogenic variant, for a given locus is to measure the incidence of new cases of a genetic condition that is clearly recognizable in all neonates who have a particular genetic alteration. Achondroplasia, a condition of reduced bone growth leading to short stature (Case 2), is a condition that meets these requirements. In one series of 242,257 consecutive births, 7 children with achondroplasia were born to parents of average stature; because achondroplasia always manifests when a pathogenic variant is present, all were considered to represent new mutations. Thus the new mutation rate at this locus can be calculated to be 7 new mutations in a total of 2 × 242,257 copies of the relevant gene, or ~1.4 × 10−5 pathogenic
52 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE failure to properly repair DNA after damage. Many such mutation events are spontaneous, arising during the normal (but imperfect) processes o...
Ch4 · Pt9 CHAPTER 4 — Human Genetic Diversity 53 variants per locus per generation. This high mutation rate is particularly striking because virtually all cases of achondroplasia are due to the identical variant: a G to A transition that changes a glycine codon to an arginine in the encoded protein. The rate of pathogenic mutation has been estimated for a number of other disorders in which the occurrence of a new variant was identified by the appearance of a detectable disease (Table 4.3). The measured rates for these and other disorders vary over a 1000-­fold range, from 10−4 to 10−7 mutations per locus per generation. The basis for these differences may be related to some or all of the following: the size of different genes, the fraction of all variants in that gene that will lead to the disease, the age and sex of the parent in whom the mutation occurred, the mutational mechanism, and the presence or absence of mutational hot spots in the gene. Indeed, the high rate of the particular site-­specific mutation event in achondroplasia may be partially explained by it being at a hot spot for mutation by deamination, as discussed earlier. Notwithstanding this range of rates among different genes, the median gene mutation rate is ~1 × 10−6. Given that there are at least 5000 genes in the human genome in which variants are currently known to cause a discernible disease or other trait (see Chapter 7), ~1 in 200 persons is likely to receive a new pathogenic variant in a known disease-­associated gene due to mutation in one or the other parent. Sex Differences and Age Effects on Mutation Rates Because the DNA undergoes far more replication cycles in sperm than in ova (see Chapter 2), there is greater opportunity for errors to occur in sperm, suggesting that new variants will be more often paternal than maternal in origin. Indeed, where this has been explored, new variants responsible for certain conditions (e.g., achondroplasia, as just discussed) are usually missense variants that arose nearly always in the paternal germline. Furthermore, the older a man is, the more rounds of replication have preceded the meiotic divisions, thus the frequency of new paternal variants might be expected to increase with the age of the father. Indeed, increased paternal age is correlated with increased incidence of SNVs for a number of disorders (including achondroplasia) and with the incidence of CNVs in autism spectrum disorders (Case 5) and intellectual disability. For other diseases, however, the parent-­of-­origin and age effects on mutational spectra are, for unknown reasons, not as striking. TYPES OF MUTATION AND THEIR CONSEQUENCES In this section we consider the nature of different types of mutation and their effect on the genes involved. Each type of mutation discussed here is illustrated by one or more disease examples. Notably, the specificity of the pathogenic variant found in almost all cases of achondroplasia is the exception rather than the rule, and the variants that underlie a single genetic disease are typically heterogeneous among a group of affected individuals. Different cases of a particular disorder will therefore usually be caused by different underlying pathogenic variants in one gene (allelic heterogeneity), sometimes in different genes (locus heterogeneity). In Chapters 11 and 12 we will turn to the ways in which variants in specific disease-­ associated genes cause these disorders. TABLE 4.3 Estimates of Mutation Rates for Selected Human Disease Genes Disease Locus (Protein) Mutation Ratea Achondroplasia (Case 2) FGFR3 (fibroblast growth factor receptor 3) 1.4 × 10−5 Aniridia PAX6 (Pax 6) 2.9–­5 × 10−6 Duchenne muscular dystrophy (Case 14) DMD (dystrophin) 3.5–­10.5 × 10−5 Hemophilia A (Case 21) F8 (factor VIII) 3.2–­5.7 × 10−5 Hemophilia B (Case 21) F9 (factor IX) 2–­3 × 10−6 Neurofibromatosis, type 1 (Case 34) NF1 (neurofibromin) 4–­10 × 10−5 Polycystic kidney disease, type 1 (Case 37) PKD1 (polycystin) 6.5–­12 × 10−5 Retinoblastoma (Case 39) RB1 (Rb 1) 5–­12 × 10−6 a Expressed as mutations per locus per generation. Based on data in Vogel F, Motulsky AG: Human genetics, ed 4, Berlin, 1997, Springer-­Verlag. TABLE 4.4 Types of Variation in Human Genetic Disease Type of Variation Percentage of Disease-­Causing Variants Nucleotide Substitutions Missense variants (amino acid substitutions) 40% Nonsense variants (premature stop codons) 10% RNA processing variants (destroy consensus splice sites, cap sites, and polyadenylation sites or create cryptic sites) 10% Splice-­site variants leading to frameshift mutations and premature stop codons 10% Long-­range regulatory variants Rare Deletions and Insertions Addition or deletions of a small number of bases 25% Larger gene deletions, inversions, fusions, and duplications (may be mediated by DNA sequence homology either within or between DNA strands) 5% Insertion of a LINE or Alu element (disrupting transcription or interrupting the coding sequence) Rare Dynamic variants (expansion of trinucleotide or tetranucleotide repeat sequences) Rare
CHAPTER 4 — Human Genetic Diversity 53 variants per locus per generation. This high mutation rate is particularly striking because virtually all cases of achondroplasia are due to the identical varian...
Ch4 · Pt10 54 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Nucleotide Substitutions Missense Variants A single nucleotide substitution (or point mutation or SNV) in a gene sequence, such as that in the example of achondroplasia, can alter the code in a triplet of bases and cause the nonsynonymous replacement of one amino acid by another in the gene product (see the genetic code in Table 3.1 and the example in Fig. 4.4). Such events are called missense mutations, creating missense variants because they alter the coding (or sense) strand of the gene to specify a different amino acid. Although not all missense variants lead to an observable change in the function of the protein, the resulting protein may fail to work properly, may be unstable and rapidly degraded, or may fail to localize in its proper intracellular position. In many disorders, such as β-­thalassemia, most of the variants detected in different patients are missense variants (see Chapter 12). Nonsense Variants Point mutation in a DNA sequence that causes the replacement of the normal codon for an amino acid by one of the three termination (or “stop”) codons creates a nonsense variant or premature termination codon (PTC; also called stop gain). Because translation of messenger RNA (mRNA) ceases when a termination codon is reached (see Chapter 3), a variant that converts a coding codon into a termination codon causes translation to stop prematurely. In general, mRNAs harboring a PTC are targeted for rapid degradation through a cellular process known as nonsense-­mediated mRNA decay (NMD), and no translation is possible. Rarely, transcripts harboring a PTC escape NMD, most predictably if the premature stop codon occurs in last 50 bp of the penultimate exon or anywhere in the final exon of a gene. In this circumstance, a nonsense mutation can often give rise to a truncated protein with altered function. Rarely, an SNV can alter the normal termination codon (called a stop loss variant), permitting translation to continue until another termination codon in the mRNA is reached further downstream. Such a variant can lead to an abnormal protein product with additional amino acids at its carboxy-­terminus. Alternatively, access of a translating ribosome into the 3′ untranslated region downstream of the normal stop codon can displace proteins that regulate mRNA stability and/­or translation. Variants Affecting RNA Transcription, Processing, and Translation The normal mechanism by which initial RNA transcripts are made and then converted into mature mRNAs (or final versions of noncoding RNAs) requires a series of modifications, including transcription factor binding, 5′ capping, polyadenylation, and splicing (see Chapter 3). All of these steps in RNA maturation depend on specific sequences within the RNA. In the case of splicing, two general classes of splicing variants have been described. For introns to be excised from unprocessed RNA and the exons spliced together to form a mature RNA requires particular nucleotide sequences located at or near the exon-­intron (5′ donor site) or the intron-­exon (3′ acceptor site) junctions. Variants that substitute the required bases at either the splice donor or acceptor site prevent normal RNA splicing. Substitution of less conserved adjacent bases has a variable impact on splicing efficiency. A second class of splicing variants involves base substitutions that do not affect the donor or acceptor site sequences themselves, but instead create alternative donor or acceptor sites that compete with the normal sites during RNA processing. Activation of these so-­called cryptic splice sites can lead to inappropriate exclusion or inclusion of exonic or intronic sequences, respectively, in the mature mRNA. Thus at least a proportion of the mature mRNA or noncoding RNA in such cases may contain improperly spliced intron sequences. Examples of both types of variation are presented in Chapter 12. For protein-­coding genes, even if the mRNA is made, SNVs in the 5′ and 3′ untranslated regions can contribute to disease by changing mRNA stability or translational efficiency, thereby reducing the amount of protein product. Deletions, Insertions, and Rearrangements Mutation can also involve the insertion, deletion, or rearrangement of DNA sequences. Some deletions and insertions involve only a few nucleotides and are generally most easily detected by direct sequencing of that part of the genome. In other cases, a substantial segment of a gene or an entire gene is deleted, duplicated, inverted, or translocated to create a novel arrangement of gene sequences—­collectively called structural variants. Depending on the exact nature of the deletion, insertion, or rearrangement, a variety of different laboratory approaches can be used to detect the genomic alteration. Some deletions and insertions affect only a small number of base pairs. When such a variant occurs in a coding sequence and the number of bases involved is not a multiple of three (i.e., not an integral number of codons), the reading frame will be altered beginning at the point of the insertion or deletion. The results are called frameshift variants (see Fig. 4.4). From the point of the insertion or deletion, a different sequence of codons is thereby generated that encodes incorrect amino acids followed by a termination codon in the shifted frame. This typically leads to degradation of the altered transcript via activation of NMD or, more rarely,
54 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Nucleotide Substitutions Missense Variants A single nucleotide substitution (or point mutation or SNV) in a gene sequence, such as that in th...
Ch4 · Pt11 CHAPTER 4 — Human Genetic Diversity 55 an altered and truncated protein product. In contrast, if the number of base pairs inserted or deleted is a multiple of three, then no frameshift occurs, and there will be a simple insertion or deletion of the corresponding amino acids in the otherwise normally translated gene product. Larger insertions or deletions can affect multiple exons of a gene and cause major disruptions of the coding sequence. One type of insertion mutation involves insertion of a mobile element, such as those belonging to the LINE family of repetitive DNA (LINE-­1 [L1] elements). Any of the 146 putatively active L1 elements currently recognized in the human genome are capable of movement by retrotransposition (introduced earlier). Such movement not only generates genetic diversity in our species (see Fig. 4.2) but can cause disease by insertional mutagenesis. For example, in some patients with the severe bleeding disorder hemophilia A (Case 21), LINE sequences several kilobases long are found within an exon in the factor VIII gene, interrupting the coding sequence and inactivating the gene. LINE insertions throughout the genome are also common in colon cancer, reflecting retrotransposition in somatic cells (see Chapter 16). As we discussed earlier in this chapter, duplications, deletions, and inversions of a larger segment of a single chromosome are predominantly the result of homologous recombination between DNA segments with high sequence homology (Fig. 4.5). Disorders arising as a result of such exchanges can be due to a change in the dosage of otherwise wild-­type gene products when the homologous segments lie outside the genes themselves (see Chapter 6). Alternatively, such events can lead to a change in the nature of the encoded protein itself when recombination occurs between different genes within a gene family (see Chapter 12) or between genes on different chromosomes (see Chapter 16). Abnormal pairing and recombination between two similar sequences in opposite orientation on a single strand of DNA leads to inversion. For example, nearly half of all cases of hemophilia A are due to recombination that inverts a number of exons, thereby disrupting gene structure and rendering the gene incapable of encoding a normal gene product (see Fig. 4.5). Repeat Expansion Variants The pathogenic variant in some disorders involves amplification of a simple nucleotide repeat sequence. For example, simple repeats such as (CCG)n, (CAG)n, or (CCTG)n—­located in the coding portion of an exon, in an untranslated region of an exon, or even in an intron—­may expand during gametogenesis in repeat expansion or dynamic mutation, and interfere with normal gene expression or protein function. An expanded repeat in a coding region will generate an abnormal protein product; in the untranslated regions or introns of a gene, it may interfere with transcription, mRNA processing, or translation. How repeat expansions occur is not completely understood; they are conceptually similar to microsatellites but expand at a much higher rate. The involvement of simple nucleotide repeat expansions in disease is discussed further in Chapter 7. In such disorders, marked parent-­of-­origin effects are A B 23 22 21 1 23 1 A B 23 22 21 1 Mispairing and recombination Hemophilia A mutation Remainder of gene Upstream of gene Remainder of gene Upstream of gene Factor VIII gene A/B A/B 21 22 Inverted segment within gene Figure 4.5 Inverted homologous sequences, labeled A and B, located 500 kb apart on the X chromosome, one upstream of the factor VIII gene, the other in an intron between exons 22 and 23 of the gene. Intrachromosomal mispairing and recombination results in inversion of exons 1 through 22 of the gene, thereby disrupting the gene and causing severe hemophilia.
CHAPTER 4 — Human Genetic Diversity 55 an altered and truncated protein product. In contrast, if the number of base pairs inserted or deleted is a multiple of three, then no frameshift occurs, and the...
Ch4 · Pt12 56 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE well known and appear characteristic of the specific disease and/­or the particular simple nucleotide repeat involved (see Chapter 13). Such differences may be due to fundamental biologic differences between oogenesis and spermatogenesis but may also result from selection against gametes carrying certain repeat expansions. VARIATION IN INDIVIDUAL GENOMES The most extensive current inventory of the amount and type of variation in any given genome, relative to the (composite) human reference genome sequence (see Chapter 2), comes from the analysis of individual diploid human genome sequences. The first such sequence, that of a male individual, was reported in 2007. Now, hundreds of thousands of individual genomes have been sequenced, some as part of large international research consortia exploring human genetic diversity in health and disease, and others in the context of clinical sequencing to determine the underlying basis of disorders in particular patients. What degree of genome variation does one detect in such studies? Individual human genomes typically carry ~3.5 million SNVs when compared to the reference genome, of which—­depending in part on the population—­currently 1% are novel (i.e., not previously documented) (see Box 4.2). This suggests that the number of SNVs described for our species is still incomplete, although presumably the novel fraction will decrease as more genomes from more populations are sequenced. Within this variation lie variants with clinical impact that are either known, likely, or suspected. Each genome carries 50 to 100 variants that have previously been implicated in known inherited conditions. In addition, each genome carries thousands of nonsynonymous SNVs in protein-­coding genes, some of which would be predicted to alter protein function. Each genome also carries ~200 to 500 likely loss-­of-­ function variants, some of which are present at both alleles of a gene in that individual. Within the clinical setting, this realization has important implications for the interpretation of genome sequence data from patients, particularly when trying to predict the impact of variants in genes of currently unknown function (see Chapter 13). An interesting and unanticipated aspect of individual genome sequencing is that each new genome reveals some sequence that is still undocumented or unannotated in the reference human genome assembly. It is estimated that, once fully elucidated, the complete genome sequence representing the current world population will be 20 to 40 Mb larger than the extant reference assembly. Recently WGS performed on only 94 individuals from 44 African populations revealed 33.6 million SNVs, of which 5.7 million (17%) were Clinical Sequencing Studies In the context of genomic medicine, a key question is the extent to which variation in the sequence and/­or expression of one’s genome influences the likelihood of disease onset, determines the natural history of disease, and/­or provides clues relevant to its management. As just discussed, constitutional genomic variants can have a number of different direct or indirect effects on gene function. Sequencing of entire genomes (WGS, also referred to as genome sequencing) or of the subset comprising all known coding exons (exome sequencing) has been introduced in a number of clinical settings, as will be discussed in greater detail in Chapter 11. Both exome sequencing and WGS have been used to detect de novo changes (both SNVs and CNVs) in a variety of conditions of complex and/­or unknown etiology. These include, for example, various neurodevelopmental or neuropsychiatric conditions, such as autism, schizophrenia, epilepsy, intellectual disability, and developmental delay. Clinical sequencing studies can target either germline or somatic variants. In cancer, especially, various strategies have been used to search for somatic variants in tumor tissue to identify genes potentially relevant to cancer progression (see Chapter 16). BOX 4.2 VARIATION DETECTED IN A TYPICAL HUMAN GENOME Individuals vary greatly in a wide range of biologic functions, determined in part by variation among their genomes. Any individual genome will contain (on average) the following: ≈3.5 million SNVs compared to the reference genome (varies by population) 25,000–­50,000 rare variants (private mutations or seen previously in <0.5% of individuals tested) ≈65 new (de novo) SNVs and indels not detected in parental genomes ≈200,000 indels (1–­50 bp) (varies by population) 500–­1000 deletions 1–­45 kb, overlapping ≈200 genes ≈102 in-­frame indels ≈132 shifts in reading frame 10,000–­12,000 synonymous SNVs >11,000 nonsynonymous SNVs in 4000–­5000 genes 175–­500 rare nonsynonymous variants 1 new nonsynonymous mutation ≈474 premature stop codons or splice site disrupting variants 250–­300 genes with likely loss-­of-­function variants ≈25 genes predicted to be completely inactivated novel. This illustrates the extent to which individuals of descent other than European are underrepresented among sequenced cohorts, creating a significant gap in our knowledge of human genetic diversity and our ability to advance genomic medicine.
56 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE well known and appear characteristic of the specific disease and/­or the particular simple nucleotide repeat involved (see Chapter 13). Such...
Ch4 · Pt13 CHAPTER 4 — Human Genetic Diversity 57 Direct-­to-­Consumer Genomics Access to laboratory genomic testing has moved into the public realm in recent years, no longer with mainstream medicine as its gatekeeper. Tests are marketed directly to consumers for information ranging from genealogy and ancestry tracing to information about personal health and inherited traits. Targeted screening panels are starting to give way to genome sequencing opportunities. A limited number of specific diagnostic tests have been authorized by the US Food and Drug Administration for marketing. There is considerable public appetite for the sort of individual information being offered commercially, and the massive volume of data being collected and stored has enormous research potential. Notwithstanding or minimizing the significant scientific, ethical, and clinical issues that lie ahead, it is certain that individual genome sequences will be an ongoing active part of medical practice and that some of these will be delivered from patient to practitioner, rather than vice versa. IMPACT OF GENOMIC VARIATION Although it will be self-­evident to students of human genetics that new pathogenic or rare variants in the population can have clinical consequences, it may appear less obvious that common variants can be clinically relevant. For the proportion of variation that occurs within protein-­coding genes, such loci can be studied by examining variation in the proteins encoded by the different alleles. Any one individual is likely to carry two distinct alleles determining structurally differing polypeptides at an estimated 20% of protein-­coding loci; when individuals from different ancestral groups are studied, an even greater fraction of proteins has been found to exhibit detectable polymorphism. In addition, even when the gene product is unaltered, the levels of expression of that product may be very different among individuals, determined by a combination of genetic and epigenetic variation, as we saw in Chapter 3. Thus a striking degree of biochemical individuality exists within the human species in its makeup of enzymes and other gene products. Furthermore, the products of many of the encoded biochemical and regulatory pathways interact in functional and physiologic networks. Each individual, therefore—­regardless of state of health—­has a unique, genetically determined chemical makeup and responds in a unique manner to environmental, dietary, and pharmacologic influences. This concept of chemical individuality first put forward over a century ago by Archibald Garrod, the remarkably prescient British physician introduced in Chapter 1, remains true today. The broad question of what is normal—­an essential concept in human biology and in clinical medicine—­remains very much an open and controversial one when it comes to the human genome. The following chapters will explore this concept of individuality in detail, first in the context of structural genome and chromosome variants (Chapters 5 and 6) and then in terms of intragenic variants that determine the inheritance of genetic disease (Chapter 7) and influence its likelihood in families and populations (Chapter 10). Assessing the Clinical Significance of a Gene Variant The American College of Medical Genetics and Genomics and the Association for Molecular Pathology recommend that all variants detected during sequencing of genes for monogenic disease (whether from targeted, exome, or genome sequencing) be classified on a five-­level scale, spanning pathogenic, likely pathogenic, of uncertain significance, likely benign, and benign variants. Specialists in molecular diagnostics, human genomics, and bioinformatics have developed criteria for assessing where a variant sits among these five categories. None of these criteria are definitive; they must be considered together to provide an overall assessment of the evidence for pathogenicity. These criteria include the following: Population frequency—­If a variant has been seen frequently in a sizable fraction of the general population, beyond what is expected based on the prevalence of the disease, it is considered less likely to be disease causing. Being frequent, however, is no guarantee that a variant is benign. Autosomal recessive conditions result from homozygosity for disease-­causing variants that may be surprisingly common, largely harbored by asymptomatic heterozygous carriers. Conversely, rare variants are not necessarily pathogenic; most variants found in an exome or genome sequence are individually rare. In silico assessment—­Computational algorithms can evaluate how likely a missense variant is to be damaging to the protein, by using information such as whether the amino acid at that position is conserved in orthologous proteins (in other species), the structural location of the variant, and machine-­learning algorithms. Such tools are limited in their accuracy for predicting functional impact and therefore can never be used alone to definitively determine pathogenicity. They are, however, improving with time, and their contribution to variant assessment may strengthen. Other bioinformatics tools assess the pathogenicity of other types of variants, such as potential splice site variants and other noncoding sequence variants. Functional data—­If a particular variant adversely affects in vitro biochemical activity, a function in cultured cells, or the health of a model organism, then it is less likely to be benign. However, it remains possible
CHAPTER 4 — Human Genetic Diversity 57 Direct-­to-­Consumer Genomics Access to laboratory genomic testing has moved into the public realm in recent years, no longer with mainstream medicine as its gat...
Ch4 · Pt14 58 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE that a particular variant will appear benign by these criteria and still be disease causing in humans because of a prolonged human life span, environmental triggers, or compensatory genes present in the model organism but not in humans. Conversely, functional effects demonstrated in systems that do not fully represent the human biologic state may falsely implicate a variant as pathogenic. Caution must be exercised to ensure adequate validation of these assays with variants determined to be pathogenic or benign through other types of evidence. Segregation data—­If a particular variant is coinherited with a disease in one or more families or, conversely, does not track with a disease in the family under investigation, then it is more or less likely to be pathogenic. Of course, when only a few individuals are affected, the variant and disease may appear to track by random chance; to be considered strong evidence for pathogenicity, the number of times a variant and disease must be coinherited is generally accepted to be in at least five informative meioses. Finding affected individuals in the family who do not carry the variant would be strong evidence against the variant being pathogenic, but finding unaffected individuals who do carry the variant is less persuasive if the disorder is known to have reduced penetrance. De novo variant—­The appearance of a severe disorder in a child along with a new variant in a coding exon that neither parent carries (de novo variant) is additional evidence for that variant to be pathogenic. However, between one and two new changes occur in the coding regions of genes in every child (see earlier). Only de novo variants in genes that are associated with the individual’s phenotype are considered evidence for pathogenicity, given a lower prior probability of de novo mutation for a small, targeted set of genes. Variant characterization—­A variant may be synonymous, missense, nonsense, a frameshift with a premature termination downstream, or cause a highly conserved splice site change. The impact on the function of the gene can be inferred but, once again, is not definitive. For example, a synonymous change that does not alter an amino acid codon might be thought to be benign but may have deleterious effects on normal splicing and be pathogenic (see examples in Chapter 12). Conversely, one might assume that premature termination or frameshift variants are always deleterious and disease causing; however, such an alteration at the far 3′ end of a gene may produce a truncated protein that is still functional and, therefore, be a benign change. Prior occurrence—­Having been seen multiple times among collections of patients with a similar disorder is important additional evidence that a variant is pathogenic. Even if a missense variant is novel (i.e., never described before) it is more likely to be pathogenic if it occurs at the same position in the protein as other known pathogenic missense variants. ACKNOWLEDGMENT We thank Miriam Reuter, Heidi Rehm and Jeff Mac Donald for contributing to this chapter. GENERAL REFERENCES Olson MV: Human genetic individuality, Ann Rev Genomics Hum Genet 13:1–­27, 2012. Strachan T, Read A, editors: Human molecular genetics ed 5, New York, 2018, Garland Science. The 1000 Genomes Project Consortium: An integrated map of genetic variation from 1,092 human genomes, Nature 491:56–­65, 2012. Trost B, Loureiro LO, Scherer SW: Discovery of genomic variation across a generation, Human Molecular Genetics, Volume 30, Issue R2, 15 October 2021, Pages R174–R186, https://­doi. org/­10.1093/­hmg/­ddab 209. Willard HF: The human genome: a window on human genetics, biology and medicine. In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine ed 3, New York, 2016, Elsevier. REFERENCES FOR SPECIFIC TOPICS Alkan C, Coe BP, Eichler EE: Genome structural variation discovery and genotyping, Nature Rev Genet 12:363–­376, 2011. Bagnall RD, Waseem N, Green PM, et al: Recurrent inversion breaking intron 1 of the factor VIII gene is a frequent cause of severe hemophilia A, Blood 99:168–­174, 2002. Crow JF: The origins, patterns and implications of human spontaneous mutation, Nature Rev Genet 1:40–­47, 2000. Fan S, Kelley DE, Beltrame MH, et al: African evolutionary history inferred from whole genome sequence data of 44 indigenous African populations, Genome Biol 20:82, 2019. Gardner RJ: A new estimate of the achondroplasia mutation rate, Clin Genet 11:31–­38, 1977. Karczewski KJ, Francioli LC, Tiao G, et al: The mutational constraint spectrum quantified from variation in 141,456 humans, Nature 581:434–­443, 2020. https://­doi.org/­10.1038/­s 41586-­020-­2308-­7 Kong A, Frigge ML, Masson G, et al: Rate of de novo mutations and the importance of father’s age to disease risk, Nature 488:471–­475, 2012. Lappalainen T, Sammeth M, Friedlander MR, et al: Transcriptome and genome sequencing uncovers functional variation in humans, Nature 501:506–­511, 2013. Mac Arthur DG, Balasubramanian S, Rrankish A, et al: A systematic survey of loss-­of-­function variants in human protein-­coding genes, Science 335:823–­828, 2012. Mc Bride CM, Wade CH, Kaphingst KA: Consumers’ view of direct-­ to-­consumer genetic information, Ann Rev Genomics Hum Genet 11:427–­446, 2010. Richards S, Aziz N, Bale S, et al: Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology, Genet Med 17(5):405–­424, 2015. https://doi.org/10.1038/gim.2015.30 Stewart C, Kural D, Stromberg MP, et al: A comprehensive map of mobile element insertion polymorphisms in humans, PLo S Genet 7:e 1002236, 2011. Sun JX, Helgason A, Masson G, et al: A direct characterization of human mutation based on microsatellites, Nature Genet 44:1161–­ 1165, 2012.
58 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE that a particular variant will appear benign by these criteria and still be disease causing in humans because of a prolonged human life span,...
Ch4 · Pt15 CHAPTER 4 — Human Genetic Diversity 59 PROBLEMS 1. Variation can arise from a variety of mechanisms, with different consequences. Describe and contrast the types of variation that can have the following effects: a. A change in dosage of a gene or genes b. A change in the sequence of multiple amino acids in the product of a protein-­coding gene c. A change in the final structure of an RNA produced from a gene d. A change in the order of genes in a region of a chromosome e. No obvious effect 2. Aniridia is an eye disorder characterized by the complete or partial absence of the iris and is always present when a pathogenic variant occurs in the responsible gene. In one population, 41 children diagnosed with aniridia were born to parents of normal vision among 4.5 million births during a period of 40 years. Assuming that these cases were due to new mutation events, what is the estimated mutation rate at the aniridia locus? On what assumptions is this estimate based, and why might this estimate be either too high or too low? 3. Which of the following types of variation would be most effective for distinguishing two individuals from the general population: a single nucleotide variant (SNV), a simple indel, or a microsatellite? Explain your reasoning. 4. Compare the likely impact of each of the following on the overall rate of mutation detected in any given genome: age of the parents, hot spots of mutation, intrachromosomal homologous recombination, genetic variation in the parental genomes.
CHAPTER 4 — Human Genetic Diversity 59 PROBLEMS 1. Variation can arise from a variety of mechanisms, with different consequences. Describe and contrast the types of variation that can have the follow...

Chapter 5: Principles of Clinical Cytogenetics and Genome Analysis

Ch5 · Pt1 chapter 5 Principles of Clinical Cytogenetics and Genome Analysis Dimitri J. Stavropoulos Clinical cytogenetics is the study of chromosomes, their structure, and their inheritance, as applied to the practice of medicine. It has been apparent for over 50 years that chromosome abnormalities—microscopically visible changes in the number or structure of ­chromosomes— could account for a number of clinical conditions that are thus referred to as chromosome disorders. With their focus on the complete set of genetic material, cytogeneticists were the first to bring a genome-­wide perspective to the practice of medicine. Today, chromosome ­analysis—with increasing resolution and precision at both the cytologic and genomic levels—is an important diagnostic procedure in numerous areas of clinical medicine. Current genome analyses that use approaches to be explored in this chapter, including chromosomal microarrays and whole genome sequencing (WGS), represent impressive improvements in capacity and resolution but ones that are conceptually similar to microscopic methods focusing on chromosomes (Fig. 5.1). Chromosome disorders form a major category of genetic disease. They account for a large proportion of all spontaneous pregnancy losses, congenital malformations, and intellectual disability and play an important role in the pathogenesis of cancer. Specific cytogenetic disorders are responsible for hundreds of distinct syndromes that collectively are more common than all the single-­gene diseases together. Cytogenetic abnormalities are present in nearly 1% of live births, in ~2% of pregnancies in women older than 35 years who undergo prenatal diagnosis, and in half of all spontaneous, first-­ trimester pregnancy losses. The spectrum of analysis from microscopically visible changes in chromosome number and structure to anomalies of genome structure and sequence detectable at the level of WGS encompasses literally the entire field of medical genetics (see Fig. 5.1). In this chapter we present the general principles of chromosome and genome analysis and focus on the chromosome variants and structural variants introduced in the previous chapter. We restrict our discussion to disorders due to genomic imbalance—either for the hundreds to thousands of genes found on individual chromosomes or for smaller numbers of genes located within a particular chromosome region. Application of these principles to some of the most common and best-­known chromosomal and genomic disorders will then be presented in Chapter 6. INTRODUCTION TO CYTOGENETICS AND GENOME ANALYSIS The general morphology and organization of human ­chromosomes, as well as their molecular and genomic composition, were introduced in Chapters 2 and 3. Chromosome analysis can be performed for clinical purposes by obtaining peripheral blood and stimulating T lymphocytes to prepare short-­term cultures. After a few days, the dividing cells are arrested in metaphase with chemicals that inhibit the mitotic spindle, and chromosomes are fixed to glass slides and stained by one of several techniques, depending on the particular diagnostic procedure being performed. They are then ready for analysis. Although ideal for rapid clinical analysis, cell cultures prepared from peripheral blood have the disadvantage of being short lived (3–­4 days). Long-­term cultures suitable for permanent storage or further studies can be derived from a variety of other tissues. Skin biopsy, a minor surgical procedure, can provide samples of tissue that in culture produce fibroblasts, which can be used for a variety of biochemical and molecular studies as well as for chromosome and genome analysis. White blood cells can also be transformed in culture to form lymphoblastoid cell lines that are potentially immortal. Bone marrow has the advantage of containing a high proportion of dividing cells so that little if any culturing is required; however, it can be obtained only by the relatively invasive procedure of marrow biopsy. Its main use is in the diagnosis of suspected hematologic malignancies. Fetal cells derived from amniotic fluid (amniocytes) or obtained by chorionic villus biopsy can also be cultured successfully for cytogenetic, genomic, biochemical, or molecular analysis. Chorionic villus cells can also be analyzed directly after biopsy, without the need for culturing. Remarkably, small amounts of cell-­free fetal DNA are found in the maternal plasma and can be tested by WGS (see Chapter 18 for further discussion).
chapter 5 Principles of Clinical Cytogenetics and Genome Analysis Dimitri J. Stavropoulos Clinical cytogenetics is the study of chromosomes, their structure, and their inheritance, as applied to the p...
Ch5 · Pt2 62 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Molecular analysis of the genome, including WGS, can be carried out on any appropriate clinical material, provided that good-­quality DNA can be obtained. Cells need not be dividing for this purpose, and thus it is possible to study DNA from tissue and tumor samples, for example, as well as from peripheral blood. Which approach is most appropriate for a particular diagnostic or research purpose is a rapidly evolving area as the resolution, sensitivity, and ease of chromosome and genome analysis increase (see Box 5.1). Chromosome Identification The different chromosomes in the genome can be identified cytologically by their characteristic banding patterns after applying specific staining procedures. The most common of these, Giemsa banding (G-­banding), was developed in the early 1970s and was the first widely used whole genome analytic tool for research and clinical diagnosis that may still apply. It has been the gold standard for the detection and characterization of structural and numerical genomic abnormalities in clinical diagnostic settings for both constitutional (postnatal or prenatal) and acquired (cancer) disorders.
Chromosome/ Genome Variation Interchromosomal translocations Ring chromosomes, isochromosomes Marker chromosomes Aneuploidy Aneusomy Segmental aneusomy Chromosomal inversions Intrachromosomal transloc...
Ch5 · Pt3 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 63 BOX 5.1 CLINICAL INDICATIONS FOR CHROMOSOME AND GENOME ANALYSIS Chromosome analysis is indicated as a routine diagnostic procedure for a number of specific conditions encountered in medicine, and some general clinical indications include: Problems of early growth and development. Failure to thrive, developmental delay, dysmorphic facies, multiple malformations, short stature, ambiguous genitalia, and intellectual disability are frequent findings in children with chromosome abnormalities. Unless there is a definite nonchromosomal diagnosis, genomic analysis should be performed to detect diagnostic genome-­wide copy number and sequence variants for patients presenting with any combination of such problems (see Chapter 11). Stillbirth and neonatal death. The incidence of chromosome abnormalities is much higher among stillbirths (up to ~10%) than among live births (~0.7%). It is also elevated among infants who die in the neonatal period (~10%). For unexplained stillbirths and neonatal deaths, genome-­wide copy number and sequence analysis may serve to reveal a genetic etiology. These analyses may provide important information for prenatal or preimplantation genetic diagnosis (see Chapter 18) in future pregnancies. Fertility problems. Chromosome studies by G-­banded karyotype are indicated for women with amenorrhea and for couples with a history of infertility or recurrent miscarriage. A chromosome abnormality is seen in one or the other parent in 3–­6% of cases in which there is infertility or two or more miscarriages. Structural characterization of genomic imbalances and family follow-­up studies. Copy number variations (CNV) identified by genome-­wide copy number analysis may require additional studies by G-­banding karyotype or metaphase fluorescence in situ hybridization (FISH) to characterize the structure of the alteration. A known unbalanced chromosome abnormality may have resulted from a parental balanced rearrangement, which will have implications for future pregnancies and potential prenatal diagnosis, as well as potentially for other family members of the carrier parent. Neoplasia. Almost all cancers are associated with one or more chromosome abnormalities (see Chapter 16). Chromosome and genome evaluation in the tumor itself, or in bone marrow for hematologic malignant neoplasms, can offer diagnostic or prognostic information. Pregnancy. Several prenatal risk factors, including advanced maternal age, biochemical markers, and ultrasound findings, are associated with chromosome abnormalities (see Chapter 18). Fetal genome-­wide copy number and sequence analysis should be offered as a routine part of prenatal care in such pregnancies. Noninvasive prenatal screening using whole genome sequencing of cell-­free DNA in maternal blood is also available to screen for the most common chromosome disorders. only a single arm, does not occur in the normal human karyotype, but it is occasionally observed in chromosome rearrangements. The human acrocentric chromosomes (13, 14, 15, 21, and 22) have small, distinctive masses of chromatin known as satellites attached to their short arms by narrow stalks (called secondary constrictions). The stalks of these five chromosome pairs contain hundreds of copies of genes for ribosomal RNA (the major component of ribosomes; see Chapter 3) as well as a variety of repetitive sequences. The standard G-­banded karyotype at a 400-­ to 550-­ band stage of resolution, as seen in a typical metaphase preparation, allows detection of deletions and duplications greater than ~5 to 10 Mb (see Fig. 5.1). However, the sensitivity of G-­banding at this resolution may be lower in regions of the genome in which the banding patterns are less specific. High-­resolution banding (also called prometaphase banding) can achieve 850 or more bands in a haploid set by staining chromosomes that have been obtained at an early stage of mitosis (prophase or prometaphase), when they are still in a relatively uncondensed state (see Chapter 2). Development of high-­resolution chromosome analysis in the early 1980s allowed the discovery of a number of new microdeletion and microdyuplication syndrome, caused by smaller genomic rearrangements in the 2-to-3 Mb size range (see Fig. 5.1). However, the time-­consuming and technically difficult nature of this cytogenetic method precludes its routine use for whole genome analysis. In addition to changes in banding pattern, nonstaining gaps, called fragile sites, are heritable variants that can be observed at particular chromosome sites that are prone to regional genomic instability induced by stress on DNA replication. Over 100 common and rare (population frequency <5%) fragile sites are documented. Common fragile sites are postulated to drive genomic instability in cancer cells, and a small proportion of rare fragile sites are associated with specific clinical disorders. For example, the rare fragile site located at Xq 27.3 is caused by an expansion of CGG repeats and is observed in patients with fragile X syndrome (Case 17). Fluorescence In Situ Hybridization (FISH) Targeted high-­resolution chromosome banding was largely replaced in the early 1990s by FISH, a method for detecting the presence or absence of a particular DNA sequence or for evaluating the number or structural organization of a chromosome or chromosomal region in situ (literally, “in place”) in the cell. This convergence of genomic and cytogenetic approaches—variously termed molecular cytogenetics or cytogenomics—dramatically expanded both the scope and precision of chromosome analysis in routine clinical practice. FISH technology takes advantage of ordered collections of recombinant large-­insert DNA clones containing DNA from virtually any locus in the genome. Clones containing specific human DNA sequences can be labeled with a fluorescent dye and used as probes to detect the corresponding region of the genome in
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 63 BOX 5.1 CLINICAL INDICATIONS FOR CHROMOSOME AND GENOME ANALYSIS Chromosome analysis is indicated as a routine diagnostic procedur...
Ch5 · Pt4 64 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE p 1 q p q 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 X Y 36.2 35 34.2 33 31 24 22 16 14 12 24 26 22 13.1 13.3 22 24 26.1 26.3 28 14 12 13 13 22 24 26 28 31.2 32 34 12 14 12 21 23 14 15.2 12 14 16 12 22 24 26 22 21.2 24 31 21 12 33 21 14 23 21.3 21.1 12 24.2 22 12 33 31 21 12 23 21 12 23 21 14 12 25 14 12 14 12 22 24 14 12 23 12 21 24.2 35 32 34 15.3 15.1 12 14.1 14.3 22 24 32 34 36 21 12 12 22 24 31 41 43 13 21 31 33 12 21 23 31 12 21 23 21.2 25 14 11.2 21 23 13.2 12 26.2 12 22 24 12 12 22 11.3 12 12 13.2 12 21 12 12 13.2 11.22 12 22.2 21 11.3 12 11.3 21 23 25 27 13.4 13.2 Figure 5.2 Ideogram showing G-­banding patterns for human chromosomes at metaphase, with ~400 bands per haploid karyotype. As drawn, chromosomes are typically represented with the sister chromatids so closely aligned that they are not recognized as distinct entities. Centromeres are indicated by the primary constriction and narrow dark gray regions separating the p and q arms. For convenience and clarity, only the G-­dark bands are numbered. For examples of full numbering scheme, see Fig. 5.3. (Redrawn from Shaffer LG, Mc Gowan-­ Jordan J, Schmid M, editors: ISCN 2013: an international system for human cytogenetic nomenclature, Basel, 2013, Karger.) chromosome preparations or in interphase nuclei for a variety of research and diagnostic purposes (Fig. 5.4). Although FISH technology provides much higher resolution and specificity than G-­banded chromosome analysis, it does not allow for efficient analysis of the entire genome. Its use is limited to targeting a specific genomic region based on a clinical diagnosis or suspicion, structural characterization of genomic imbalances and family follow-up studies. Multiplex Ligation-­Dependent Probe Amplification (MLPA) MLPA is a targeted copy number assay used to detect exon-­level deletions and duplications in a gene or targeted chromosome region. This method uses multiplex polymerase chain reaction (PCR) to amplify DNA sequences from multiple exons simultaneously in one PCR. The relative quantity of DNA sequence generated from each exon is then compared to amplification of control regions with normal copy number (two copies). Since the total quantity of amplified PCR product is directly proportional to the copy number of each targeted exon in the individual DNA sample, a heterozygous deletion (one copy) will produce approximately half as much PCR product when compared to regions with normal copy number (two copies), and a heterozygous duplication (three copies) will generate ~50% more PCR product when compared to regions with normal copy number. This method is limited by the number of targeted regions that can be included in one PCR assay and is not amenable to genome-­wide copy number analysis. It is used to investigate a specific gene (e.g., DMD) or known recurrent microdeletion/­ microduplication syndrome region (e.g., 22q11.2). MLPA can be combined with methylation analysis
64 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE p 1 q p q 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 X Y 36.2 35 34.2 33 31 24 22 16 14 12 24 26 22 13.1 13.3 22 24 26.1 26.3 28...
Ch5 · Pt5 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 65 (MS-­MLPA) to specifically amplify targeted chromosome regions that are methylated, to determine imprinting status. As described in Chapter 8, a subset of the genome is differentially imprinted depending on the parent of origin; and abnormal methylation of these regions can be identified by MS-­MPLA to confirm a diagnosis of imprinting disorder (e.g., 15q11.2q13 in Prader-­Willi and Angelman syndromes; see [Case 38]). Genome Analysis Using Microarrays Chromosome microarray analysis (CMA) has replaced G-­banded karyotype as the frontline diagnostic test to detect genome-­wide copy number imbalances for most clinical applications. CMA simultaneously queries the whole genome on a glass slide containing regularly spaced DNA probes that represent loci across the entire genome. This technology detects relative copy number gains and losses in a genome-­wide manner by hybridizing equal amounts of control and subject DNA to the DNA probes and calculating the ratio of each DNA sample hybridized to each probe. Microarray probes showing equal ratio of subject and control DNA indicate normal copy number at the respective genomic loci. An excess of subject DNA indicates copy number gain, whereas underrepresentation of subject DNA indicates copy number loss at the genomic loci represented by the microarray probes (Fig. 5.5). Microarray platforms may contain copy number probes (see earlier). Alternatively, they may comprise single nucleotide polymorphism (SNP) probes that contain versions of sequences corresponding to the variant alleles (as introduced in Chapter 4). The data from SNP probes can be plotted on an allele difference plot, which indicates whether a specific SNP locus is homozygous for the A allele (AA), homozygous for the B allele (BB), or heterozygous (AB). Normal copy number across a chromosome typically shows the three allele combinations of AA, AB, and BB along its length (see Fig. 5.5). A genomic region of homozygosity (ROH), with AA and BB but no AB track, can be observed when the chromosome region is identical by decent (due to parental consanguinity) or when there is uniparental disomy (UPD) with both copies of the chromosome inherited from one parent. The identification of several ROHs across the genome of a patient and involving multiple chromosomes suggests parental consanguinity and raises the possibility of a recessive disorder. While an ROH affecting only one chromosome raises the possibility of UPD, parental genotype analysis is required to confirm that the ROH is in fact due to UPD. For routine clinical testing of suspected chromosome disorders, probe spacing on the array provides a resolution as high as 100 kb over the entire unique portion of the human genome. A higher density of probes can be used to achieve even higher resolution (20 kb) over regions of particular clinical interest, such as those p q 5 6 7 15.3 15.2 15.1 14 13.3 13.1 11 11.1 11.2 12 13.1 13.2 13.3 14 15 21 22 23.1 23.3 31.1 31.2 31.3 32 33.1 33.3 34 35.1 35.3 35.2 33.2 23.2 12 25 24 23 22.3 22.1 21.3 11.1 11 12 13 14 15 16.1 16.3 21 22.1 22.3 23.1 23.3 24 25.1 25.3 26 27 21.2 21.1 12 22.2 11.2 16.2 22.2 23.2 25.2 22 21 15.3 11.1 11.1 15.1 14 13 12 15.2 11.2 11.22 11.23 21.1 21.2 21.3 22 31.1 31.2 31.3 32 33 34 35 36 13.2 11.21 Figure 5.3 Examples of G-­banding patterns for chromosomes 5, 6, and 7 at the 550-­band stage of condensation. Band numbers permit unambiguous identification of each G-­dark or G-­light band. The banding nomenclature indicates the chromosome number (1–­22,X,Y), the short arm (p) or long arm (q), the region, band, and subband. For example, chromosome 5p15.2 is pronounced as “5-­p-­one-­five-­point-­2.” (Redrawn from Shaffer LG, Mc Gowan-­ Jordan J, Schmid M, editors: ISCN 2013: an international system for human cytogenetic nomenclature, Basel, 2013, Karger.) Metaphase Locus-specific probes Satellite DNA probes Interphase Figure 5.4 Fluorescence in situ hybridization to human chromosomes at metaphase and interphase, with different types of DNA probe. (Top) Single-­copy DNA probes specific for sequences within bands 4q12 (red fluorescence) and 4q31.1 (green fluorescence). (Bottom) Repetitive α-­satellite DNA probes specific for the centromeres of chromosomes 18 (aqua), X (green), and Y (red) used to count the number of each chromosome in this individual. (Images courtesy M. Katharine Rudd, Emory Genetics Laboratory, Atlanta, Georgia.)
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 65 (MS-­MLPA) to specifically amplify targeted chromosome regions that are methylated, to determine imprinting status. As described...
Ch5 · Pt6 66 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE associated with known developmental disorders or congenital anomalies (see Chapter 6). This approach is being used in clinical laboratories to provide high-­ resolution analysis of targeted clinically significant genes and lower resolution backbone coverage across the rest of the genome. Microarrays have been used successfully to identify chromosome and genome abnormalities in children with unexplained developmental delay, intellectual disability, or birth defects, revealing a number of pathogenic genomic alterations that were not detectable by conventional G-­banding. Based on this significantly increased yield (1–­3% from karyotype versus 15–­20% from microarray), genome-­wide arrays have replaced the G-­banded karyotype as the routine frontline test for these patient populations. Two important limitations of this technology bear mentioning, however. First, array-­based methods measure only the relative copy number of DNA sequences but not whether they have been translocated or rearranged from their normal position(s) in the genome. Thus further characterization of copy number variants (CNVs) by karyotyping or FISH is important to determine the nature of an abnormality and thus its risk for recurrence for other family members. Second, high-­resolution genome analysis can reveal variants in particular, small differences in copy number, that are of uncertain clinical significance. An increasing number of such variants are being documented and catalogued even from the general population. As we saw in Chapter 4, many are likely to be benign CNVs. Their existence underscores the unique nature of an individual’s genome and emphasizes the diagnostic challenge of assessing what is considered normal and what is likely to be pathogenic. Genome Analysis by Whole Genome Sequencing On the same spectrum as cytogenetic and microarray analysis, the ultimate resolution for clinical tests to detect chromosomal and genomic disorders would be to sequence genomes in their entirety. Indeed, as the efficiency of WGS has increased and its costs have fallen, it is becoming increasingly practical to sequence samples in a clinical setting (see Fig. 5.1). The most widely used WGS approach generates millions of short-­sequence reads that range between 100 and 500 bp in length, depending on the sequencing platform. A B Log 2 Ratio AA AB BB Region of Homozygosity Log 2 Ratio Allele Difference Allele Difference AA AB BB Deletion Figure 5.5 Chromosome microarray analysis to detect copy number variants and regions of homology. (A) Chromosome 17: G-banding ideogram, followed by an example of copy number and single nucleotide polymorphism (SNP) microarray output, showing the Log 2 ratio of fluorescence intensity and allele difference plots. DNA probes (blue dots) with a Log 2 ratio of 0 indicate diploid copy number. In chromosome region 17p11.2, consecutive probes with a Log 2 ratio of −­1 indicate a heterozygous deletion of ~3.7 Mb, associated with Smith-­Magenis syndrome. (B) Chromosome 18: SNP microarray output plot showing a region of homozygosity of ~6.052 Mb. The total copy number is unaffected, but the allele difference plot shows a stretch of only homozygous genotypes (AA or BB) with no heterozygous genotypes (AB). (Microarray images courtesy of Genome Diagnostics, The Hospital for Sick Children.)
66 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE associated with known developmental disorders or congenital anomalies (see Chapter 6). This approach is being used in clinical laboratories t...
Ch5 · Pt7 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 67 An individual’s genome is represented by overlapping sequence reads, with typically 30 to 40 reads corresponding to any particular segment of the genome. A genomic region or chromosome with an abnormally low or high representation of those sequence reads is likely to have a numeric or structural abnormality of that genomic region. To detect numeric abnormalities of an entire chromosome it is generally not necessary to sequence a genome to completion; even a limited number of sequences that align to a particular chromosome of interest should reveal whether those sequences are found in the expected number (e.g., equivalent to two copies per diploid genome for an autosome) or whether they are significantly overrepresented or underrepresented (Fig. 5.6). This concept is now being applied to the prenatal diagnosis of fetal chromosome imbalance (see Chapter 18). To detect balanced rearrangements of the genome, however, in which DNA in the genome is neither gained nor lost, a more complete genome sequence is required. Here, instead of sequence reads that align perfectly to the reference human genome sequence, one finds rare sequence reads that align to two different and normally noncontiguous regions in the reference sequence (whether on the same chromosome or on different chromosomes) (see Fig. 5.6). This approach has been used to identify the specific genes involved in some cancers, and in children with various congenital defects due to translocations, involving the juxtaposition of sequences that are normally located on different chromosomes. More recently, bioinformatics algorithms have been developed to estimate the size of trinucleotide repeat expansions and provide the opportunity to assay known clinically significant loci (such as those involved in fragile X syndrome or Huntington disease) as part of the WGS diagnostic test. Clinical laboratories are beginning to implement WGS to accurately detect sequence-­level variants (single nucleotide variants); insertions/­deletions (indels) up to 50 bp and CNVs for genetically heterogeneous disorders; however, CMA and whole exome sequencing have so far been the predominant tests used for this purpose due to their lower cost. Exome sequencing (ES) generates sequence reads from protein-­coding exons, which represent ~1.5% of the genome. This provides accurate detection of exonic sequence–­level variants; however, detection of CNVs is less accurate by ES than by WGS because the number of sequence reads generated for each exon can be less consistent, and there is considerable uncertainty of CNV breakpoints due to the large unsequenced chromosome regions between exons. In addition, ES cannot detect balanced rearrangements and noncoding variants. As the cost of WGS continues to decrease, it will replace ES and CMA in genomic diagnostics, providing a much more complete representation of all the variants within an individual’s genome. Although WGS short-­read technologies provide a considerable improvement over microarray and ES, the short-­read lengths limit the ability to resolve complex A B Reference sequence of individual chromosomes Sequence reads: Patient with duplication Reference sequence of individual chromosomes Sequence reads: Patient with translocation Translocation in patient genome Sequences overrepresented in patient genome Figure 5.6 Strategies for detection of numeric and structural chromosome abnormalities by whole genome sequence analysis. Although only a small number of reads are illustrated schematically here, in practice, many millions of sequence reads are analyzed and aligned to the reference genome to obtain statistically significant support for a diagnosis of aneuploidy or a structural chromosome abnormality. (A) Alignment of sequence reads from a patient’s genome to the reference sequence of three individual chromosomes. Overrepresentation of sequences from the red chromosome indicates that the patient is aneuploid for this chromosome. (B) Alignment of sequence reads from a patient’s genome to the reference sequence of two chromosomes reveals a number of reads that contain contiguous sequences from both chromosomes. This indicates a translocation in the patient’s genome involving the blue and orange chromosomes at the positions designated by the dotted lines.
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 67 An individual’s genome is represented by overlapping sequence reads, with typically 30 to 40 reads corresponding to any particula...
Ch5 · Pt8 68 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE structural variation, repetitive regions, and genes with homologous sequence in other regions of the genome (e.g., pseudogenes). The emergence of long-­read sequencing technologies that generate sequence reads greater than 10 kb has made it possible to begin to address many of these challenges that can potentially involve clinically relevant genes. Notably, this technology is able to (1) sequence genes without interference of homologous sequence from pseudogenes, (2) provide haplotypes and phase variants across large stretches of DNA (>10 kb), (3) sequence large repeat expansions and identify intervening sequence that may affect phenotype and repeat stability (e.g. involved in DMPK, the genes for myotonic dystrophy), and (4) identify balanced and unbalanced translocations, insertions, deletions, duplications, and inversions. CHROMOSOME ABNORMALITIES Abnormalities of chromosomes may be either numeric or structural and may involve one or more autosomes, sex chromosomes, or both simultaneously. The overall incidence of chromosome abnormalities is ~1 in 154 live births (Fig. 5.7), and their impact is therefore substantial, both in clinical medicine and for society. By far the most common type of clinically significant chromosome abnormality is aneuploidy, an abnormal chromosome number due to an extra or missing chromosome. An aneuploid karyotype is typically associated with physical or neurodevelopmental abnormalities, or both. Structural abnormalities (rearrangements involving one or more chromosomes) are also relatively common (see Fig. 5.7). Depending on whether a structural rearrangement leads to an imbalance of genomic content, disruption of coding, or regulatory sequence, these may or may not have a phenotypic effect. However, as explained later in this chapter, even individuals with benign balanced chromosome abnormalities may be at an increased risk for abnormal offspring in the subsequent generation. Chromosome abnormalities are described by a standard set of abbreviations and nomenclature that indicate the nature of the abnormality and (in the case of analyses performed by FISH or microarrays) the technology used. Some of the more common abbreviations and examples of abnormal karyotypes and abnormalities are listed in Table 5.1. Gene Dosage, Balance, and Imbalance For chromosome and genomic disorders, it is primarily the quantitative aspects of gene expression that underlie disease, in contrast to single-­gene disorders in which pathogenesis often reflects qualitative aspects of a gene’s function. The clinical consequences of any particular chromosome alteration will depend on the resulting imbalance of parts of the genome, the specific genes contained in or affected by the alteration, and the likelihood of its transmission to the next generation. The central concept for thinking about chromosome and genomic disorders is that of gene dosage and its balance or imbalance. As we shall see in later chapters, this same concept applies generally to considering some single-­ gene disorders and their underlying etiology. It takes on uniform importance, however, for chromosome abnormalities, where we are generally more concerned with the dosage of genes within the relevant chromosomal region than with the actual sequence of those genes. Most genes in the human genome are present in two doses and are expressed from both copies. Some genes, however, are expressed from only a single copy (e.g., imprinted genes and X-­linked genes subject to X inactivation; see Chapter 3). Extensive analysis of clinical cases has demonstrated that the relative dosage of these genes is critical for normal development. One or three doses instead of two are generally not conducive to normal function for a dosage-­sensitive gene or set of genes that is typically expressed from two copies. Similarly, abnormalities of genomic imprinting or X inactivation that cause the anomalous expression of two copies or no expression of a gene or set of genes, instead of one, can lead to clinical disorders. Predicting clinical outcomes for chromosomal and genomic disorders can be an enormous challenge for genetic counseling, particularly in the prenatal setting. Many such diagnostic dilemmas will be presented throughout this section and in Chapters 6 and 17, but there are a number of general principles that should be kept in mind as we explore specific types of chromosome abnormality in the sections that follow (see Box 5.2). Incidence among newborns (%) Aneuploidy Structural abnormalities Total Total Balanced Unbalanced Xand Y chromosomes Autosomes All chromosome abnormalities 0 0.25 0.5 0.75 1.0 1/154 1/263 1/475 1/375 1/500 1/1600 1/700 Figure 5.7 Incidence of chromosome abnormalities in newborn surveys, based on chromosome analysis of over 68,000 newborns. (Data summarized from Hsu LYF: Prenatal diagnosis of chromosomal abnormalities through amniocentesis. In Milunsky A, editor: Genetic disorders and the fetus, ed 4, Baltimore, 1998, Johns Hopkins University Press, pp 179–­248.)
68 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE structural variation, repetitive regions, and genes with homologous sequence in other regions of the genome (e.g., pseudogenes). The emergenc...
Ch5 · Pt9 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 69 Abnormalities of Chromosome Number A human chromosome complement with any number other than 46 is said to be heteroploid. An exact multiple of the haploid chromosome number (n) is called euploid, and any other chromosome number is aneuploid. Triploidy and Tetraploidy In addition to the diploid (2n) number characteristic of normal somatic cells, two other euploid chromosome complements, triploid (3n) and tetraploid (4n), are occasionally observed in clinical material. Both triploidy and tetraploidy have been seen in fetuses. Triploidy is observed in 1% to 3% of recognized conceptions; triploid infants can be liveborn, although they do not survive long. Among the few that survive at least to the end of the first trimester of pregnancy, most result from fertilization of an egg by two sperm (dispermy). Other cases result from failure of one of the meiotic divisions in either sex, resulting in a diploid egg or sperm. The phenotypic manifestation of a triploid karyotype depends on the source of the extra chromosome set; triploids with an extra set of maternal chromosomes are typically aborted spontaneously early in pregnancy, whereas those with an extra set of paternal chromosomes typically have an abnormal TABLE 5.1 Some Abbreviations Used for Description of Chromosomes and Their Abnormalities, With Representative Examples Abbreviation Meaning Example Example’s Condition 46,XX Normal female karyotype 46,XY Normal male karyotype arr Microarray arr(X,1-­22)x 2 Normal female arr(X,Y)x 1,(1-­22)x 2 Normal male arr[GRCh 38] 8p23.3(835185_­1242591)x 1 Deletion in 8p23.3 at genomic position 835185 to 1242591 using Genome Build GRCh 38 cen Centromere del Deletion 46,XX,del(5)(q 13) Female with terminal deletion of one chromosome 5 distal to band 5q13 der Derivative chromosome der(1) Translocation chromosome derived from chromosome 1 and containing the centromere of chromosome 1 dic Dicentric chromosome dic(X;Y) Translocation chromosome containing the centromeres of both the X and Y chromosomes dup Duplication inv Inversion inv(3)(p 25q21) Pericentric inversion of chromosome 3 mar Marker chromosome 47,XX,+mar Female with an extra, unidentified chromosome mat Maternal origin arr[GRCh 38] 7p22.3(580556_­1191665)x 3 mat Maternally inherited duplication in 7p22.3 genomic position 580556 to 1191665 using genome build GRCh 38 p Short arm of chromosome pat Paternal origin q Long arm of chromosome r Ring chromosome 46,X,r(X) Female with ring X chromosome rob Robertsonian translocation 45,XX,rob(14;21)(q 10;q 10) Female with balanced Robertsonian translocation in which breakage and reunion have occurred at band 14q10 and band 21q10 in the centromeric regions of chromosomes 14 and 21; however either rob or der may be used. t Translocation 46,XX,t(2;8)(q 22;p 21) Female with balanced translocation between chromosomes 2 and 8, with breakpoints in bands 2q22 and 8p21 + Gain of 47,XX,+21 Female with trisomy 21 –­ Loss of 45,XY,–­22 Male with monosomy 22 /­ Mosaicism mos 47,XX,+21[20]/46,XX[10] Female with two populations of cells, one with trisomy 21 observed in 20 cells, and one with a normal karyotype observed in 10 cells Abbreviations from Mc Gowan-­Jordan J, Hastings RJ, Adelaide SM editors: ISCN 2020: an international system for human cytogenetic nomenclature, Basel, 2020, Karger. BOX 5.2 UNBALANCED KARYOTYPES AND GENOMES IN LIVEBORNS: GENERAL GUIDELINES FOR COUNSELING Monosomies are more deleterious than trisomies. Complete monosomies are generally not viable, except for monosomy for the X chromosome. Complete trisomies are viable for chromosomes 13, 18, 21, X, and Y. The phenotype in partial (subchromosomal) aneuploidy depends on a number of factors, including the size of the unbalanced segment, which regions of the genome are affected and which genes are involved, and whether the imbalance is monosomic or trisomic. Risk in cases of inversions depends on the location of the inversion with respect to the centromere and on the size of the inverted segment. For inversions that do not involve the centromere (paracentric inversions), there is a very low risk for an abnormal phenotype in the next generation. But, for inversions that do involve the centromere (pericentric inversions), the risk for birth defects in offspring may be significant and increases with the size of the inverted segment. For a mosaic karyotype involving any chromosome abnormality, the results from testing one tissue may not accurately represent the extent of mosaicism in other tissues of the body. Counseling is particularly challenging because the degree of mosaicism in relevant tissues or relevant stages of development is generally unknown. Thus there is uncertainty about the severity of the phenotype.
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 69 Abnormalities of Chromosome Number A human chromosome complement with any number other than 46 is said to be heteroploid. An exac...
Ch5 · Pt10 70 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE degenerative placenta (resulting in a partial hydatidiform mole), with a small fetus. Tetraploids are always 92,XXXX or 92,XXYY and likely result from failure of completion of an early cleavage division of the zygote. Aneuploidy Aneuploidy is the most common and clinically significant type of human chromosome disorder, occurring in at least 5% of all clinically recognized pregnancies. Most aneuploid individuals have either trisomy (three instead of the normal pair of a particular chromosome) or, less often, monosomy (only one representative of a particular chromosome). Either trisomy or monosomy can have severe phenotypic consequences. Trisomy can involve part of a chromosone, but trisomy for a whole chromosome is only occasionally compatible with life. By far the most common type of trisomy in liveborn infants is trisomy 21, the chromosome constitution seen in 95% of patients with Down syndrome (karyotype 47,XX,+21 or 47,XY,+21) (Fig. 5.8). Other trisomies observed in liveborns include trisomy 18 and trisomy 13. It is notable that these autosomes (13, 18, and 21) are the three with the lowest number of genes; presumably, trisomy for autosomes with a greater number of genes is lethal in most instances. Monosomy for an entire chromosome is almost always lethal; an important exception is monosomy for the X chromosome, as seen in Turner syndrome (Case 47). These conditions are considered in greater detail in Chapter 6. Although the causes of aneuploidy are not fully understood, the most common chromosomal mechanism is meiotic nondisjunction. This refers to the failure of a pair of chromosomes to disjoin properly during one of the two meiotic divisions, usually during meiosis I. The genomic consequences of nondisjunction during meiosis I and meiosis II are different (Fig. 5.9). If the error occurs during meiosis I, the gamete with 24 chromosomes contains both the paternal and the maternal members of the pair. If it occurs during meiosis II, the gamete with the extra chromosome contains both copies of either the paternal or the maternal chromosome. (Strictly speaking, these statements refer only to the paternal or maternal centromere because recombination between homologous chromosomes has usually taken place in the preceding meiosis I, resulting in some genetic differences between the chromatids and thus between the corresponding daughter chromosomes; see Chapter 2.) Proper disjunction of a pair of homologous chromosomes in meiosis I appears relatively straightforward (see Fig. 5.9). In reality, however, it involves a feat of complex engineering that requires precise temporal and spatial control over alignment of the two homologues, their tight connections to each other (synapsis), their interactions with the meiotic spindle, and, finally, their release and subsequent movement to opposite poles and to different daughter cells. The propensity for non-­disjunction of a chromosome pair has been strongly associated with aberrations in the frequency or placement, (or both), of recombination events in meiosis I, which are critical for maintaining proper synapsis. A chromosome pair with too few (or even no) recombinations, or with recombination too close to the centromere or telomere, may be more susceptible to nondisjunction than a chromosome pair with a more typical number and distribution of recombination events. In some cases, aneuploidy can result from premature separation of sister chromatids in meiosis I instead of meiosis II. If this happens, the separated chromatids may by chance segregate to the oocyte or to the polar body, leading to an unbalanced gamete. Nondisjunction can also occur in a mitotic division after formation of the zygote. If this happens at an early cleavage division, clinically significant mosaicism may result (see later section). In some malignant cell lines and some cell cultures, mitotic nondisjunction can lead to highly abnormal karyotypes. Abnormalities of Chromosome Structure Structural rearrangements result from chromosome breakage, recombination, or exchange, followed by reconstitution into an abnormal combination. Whereas rearrangements can take place in many ways, they are collectively less common than aneuploidy; overall, structural abnormalities are present in ~1 in 375 newborns (see Fig. 5.7). Like numeric abnormalities, structural rearrangements may be present in all cells of a person or in mosaic form. Structural rearrangements are classified as balanced if the genome has the normal complement of chromosomal material or unbalanced if material is additional or missing. Clearly, these designations depend on the resolution of the method(s) used to analyze a particular rearrangement (see Fig. 5.1); some that appear balanced at the level of high-­resolution banding, for example, may be seen as unbalanced when studied with chromosomal microarrays or by DNA sequence analysis. Chromosome rearrangements can be stable, capable of passing through mitotic and meiotic cell divisions unaltered, whereas others are unstable. Some of the more common types of structural rearrangements observed in human chromosomes are illustrated schematically in Fig. 5.10. Unbalanced Rearrangements Unbalanced rearrangements are detected in ~1 in 1600 live births (see Fig. 5.7); the phenotype is likely to be abnormal because of deletion or duplication of multiple genes, or (in some cases) both. Duplication of part of a chromosome leads to partial trisomy for the genes within that segment; deletion leads to partial monosomy. As a general concept, any change that disturbs normal gene dosage balance can result in abnormal development; a broad range of phenotypes can result, depending on the nature of the specific genes whose dosage is altered in a particular case.
70 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE degenerative placenta (resulting in a partial hydatidiform mole), with a small fetus. Tetraploids are always 92,XXXX or 92,XXYY and likely re...
Ch5 · Pt11 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 71 Large structural rearrangements involving imbalance of at least a few megabases can be detected at the level of routine chromosome banding. Detection of smaller changes, however, generally requires higher resolution analysis, involving FISH or chromosomal microarray analysis. B A C D 21 21 Mean fluorescence ratio (Log R) 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 22 X Y -2.5 -2.0 -1.5 -1.0 -.5 0.5 1.0 1.5 2.0 2.5 Normalized sequence representation 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 22 0.6 0.8 1.0 1.2 1.4 1.6 Figure 5.8 Chromosomal and genomic approaches to the diagnosis of trisomy 21. (A) Karyotype from a male patient with Down syndrome, showing three copies of chromosome 21. (B) Interphase fluorescence in situ hybridization analysis using locus-­specific probes from chromosome 21 (red, three spots) and from a control autosome (green, two spots). (C) Detection of trisomy 21 in a female patient by whole genome chromosome microarray. Increase in the fluorescence ratio for sequences from chromosome 21 is indicated by the red arrow. (D) Detection of trisomy 21 by whole genome sequencing and overrepresentation of sequences from chromosome 21. Normalized sequence representation for individual chromosomes (± SD) in chromosomally normal samples is indicated by the gray-shaded region. A normalized ratio of ~1.5 indicates three copies of chromosome 21 sequences instead of two, consistent with trisomy 21. (A, Courtesy Center for Human Genetics Laboratory, University Hospitals of Cleveland; B, courtesy M. Katharine Rudd, Emory Genetics Laboratory; C, courtesy Daynna J. Wolff, Medical University of South Carolina; D, original data from Dan S, Chen F, Choy KW, et al: Prenatal detection of aneuploidy and imbalanced chromosomal arrangements by massively parallel sequencing. PLo S ONE 7:e 27835, 2012.)
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 71 Large structural rearrangements involving imbalance of at least a few megabases can be detected at the level of routine chromosom...
Ch5 · Pt12 72 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE A D E F B C Terminal deletion Isochromosome Robertsonian translocation Reciprocal translocation Interstitial deletion Duplication Ring Figure 5.10 Structural rearrangements of chromosomes, described in the text. (A) Terminal and interstitial deletions, each generating an acentric fragment that is typically lost. (B) Duplication of a chromosomal segment, leading to partial trisomy. (C) Ring chromosome with two acentric fragments. (D) Generation of an isochromosome for the long arm of a chromosome. (E) Robertsonian translocation between two acrocentric chromosomes, frequently leading to a pseudodicentric chromosome. Robertsonian translocations are nonreciprocal, and the short arms of the acrocentrics are lost. (F) Translocation between two chromosomes, with reciprocal exchange of the translocated segments. Nondisjunction Nondisjunction Normal MEIOSIS I MEIOSIS II Normal Normal Normal Normal Normal Normal Figure 5.9 The different consequences of nondisjunction at meiosis I (center) and meiosis II (right), compared with normal disjunction (left). If the error occurs at meiosis I, the gametes either contain a representative of both members of the chromosome 21 pair or lack a chromosome 21 altogether. If nondisjunction occurs at meiosis II, the abnormal gametes contain two copies of one parental chromosome 21 (and no copy of the other) or lack a chromosome 21. Deletions and Duplications. Deletions involve loss of a chromosome segment, resulting in chromosome imbalance (see Fig. 5.10). A carrier of a chromosomal deletion (with one normal homologue and one deleted homologue) is monosomic for the genetic information on the corresponding segment of the normal homologue. The clinical consequences generally reflect haploinsufficiency (literally, the inability of a single copy of the genetic material to carry
72 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE A D E F B C Terminal deletion Isochromosome Robertsonian translocation Reciprocal translocation Interstitial deletion Duplication Ring Figure...
Ch5 · Pt13 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 73 out the functions normally performed by two copies), and, where examined, their severity reflects the size of the deleted segment and the number and function of the specific genes that are deleted. Cytogenetically visible autosomal deletions have an incidence of ~1 in 7000 live births. Smaller, submicroscopic deletions detected by CMA or WGS are much more common, but as mentioned earlier, the clinical significance of many such variants has yet to be determined. A deletion may occur at the end of a chromosome (terminal) or within a chromosome arm (interstitial). Deletions may originate simply from chromosome breakage and loss of the acentric segment. Numerous deletions have been identified in the course of prenatal diagnosis or in the investigation of dysmorphic patients or those with intellectual disability; specific examples of such cases will be discussed in Chapter 6. In general, duplication appears to be less harmful than deletion. However, duplication in a gamete results in chromosomal imbalance (i.e., partial trisomy), and the chromosome breaks that generate it may disrupt genes, and can lead to some phenotypic abnormality. Marker and Ring Chromosomes. Very small, unidentified chromosomes, called marker chromosomes, are occasionally seen in chromosome preparations, frequently in a mosaic state. They are usually in addition to the normal chromosome complement and are thus also referred to as supernumerary chromosomes or extra structurally abnormal chromosomes. The prenatal frequency of de novo supernumerary marker chromosomes has been estimated to be ~1 in 2500 pregnancies. Because of their small and indistinctive appearance, higher resolution genome analysis (e.g., FISH and/or CMA) is usually required for precise identification. Larger marker chromosomes contain genomic material from one or both chromosome arms, creating an imbalance for whatever genes are present. Depending on the origin of the marker chromosome, the risk for a fetal abnormality can range from very low to 100%. For reasons not fully understood, a relatively high proportion of such markers derive from chromosome 15 and from the sex chromosomes. Many marker chromosomes lack telomeres and are ring chromosomes that are formed when a chromosome undergoes two breaks, and the broken ends of the chromosome reunite in a ring structure (see Fig. 5.10). Some rings experience difficulties at mitosis, when the two sister chromatids become tangled in their attempt to disjoin at anaphase. There may be breakage of the ring followed by fusion, and larger and smaller rings may thus be generated. Because of this mitotic instability it is not uncommon for ring chromosomes to be found in only a proportion of cells. Isochromosomes. An isochromosome is a chromosome in which one arm is missing and the other duplicated in a mirror-­image fashion (see Fig. 5.10). A person with 46 chromosomes carrying an isochromosome therefore has a single copy of the genetic material of one arm (partial monosomy) and three copies of the genetic material of the other arm (partial trisomy). Although isochromosomes for a number of autosomes have been described, the most common isochromosome involves the long arm of the X chromosome—designated i(X)(q 10)—in a proportion of individuals with Turner syndrome (see Chapter 6, Case 47). Isochromosomes are also frequently seen in karyotypes of both solid tumors and hematologic malignant neoplasms (see Chapter 16). Dicentric Chromosomes. A dicentric chromosome is a rare type of abnormal chromosome in which two chromosome segments, each with a centromere, fuse end to end. Dicentric chromosomes, despite their two centromeres, can be mitotically stable if one of the two centromeres is inactivated epigenetically or if the two centromeres always coordinate their movement to one or the other pole during anaphase. Such chromosomes are formally called pseudodicentric. The most common pseudodicentrics involve the sex chromosomes or the acrocentric chromosomes (so-­called Robertsonian translocations; see later). Balanced Rearrangements Balanced chromosomal rearrangements are found in as many as 1 in 500 individuals (see Fig. 5.7) and do not usually lead to a phenotypic effect because all the genomic material is present, even though it is arranged differently (see Fig. 5.10). As noted earlier, it is important to distinguish here between truly balanced rearrangements and those that appear balanced cytogenetically but are really unbalanced at the molecular level. Because of the high frequency of common copy number variants around the genome (see Chapter 4), collectively adding up to differences of many megabases between genomes of unrelated individuals, the concept of what is balanced or unbalanced is subject to ongoing investigation and continual refinement. Even when structural rearrangements are truly balanced, they can pose a threat to the subsequent generation because carriers are likely to produce unbalanced gametes and therefore have an increased risk for abnormal offspring with unbalanced karyotypes. There is also a possibility that one of the chromosome breaks will disrupt a gene, leading to a pathogenic variant. Especially with the use of WGS to examine the nature of apparently balanced rearrangements in patients who present with significant phenotypes, this is an increasingly well-­ documented cause of disorders in carriers of balanced translocations; such translocations can be a useful clue to the identification of the gene responsible for a particular genetic disorder.
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 73 out the functions normally performed by two copies), and, where examined, their severity reflects the size of the deleted segment...
Ch5 · Pt14 74 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Translocations. Translocation involves the movement of chromosome segments between two chromosomes. There are two main types: reciprocal and nonreciprocal. Reciprocal Translocations. This type of rearrangement results from breakage or recombination involving nonhomologous chromosomes, with reciprocal exchange of the broken-­off or recombined segments (see Fig. 5.10). Usually, only two chromosomes are involved, and because the exchange is reciprocal, the total chromosome number and content is unchanged. Such translocations are usually without phenotypic effect; however, like other balanced structural rearrangements, they are associated with a high risk for unbalanced gametes and abnormal progeny. They come to attention either during prenatal diagnosis or when the parents of a clinically abnormal child with an unbalanced translocation are karyotyped. Balanced translocations are more common in couples who have had two or more spontaneous pregnancy losses and in infertile males than in the general population. Translocations present challenges for the process of chromosome pairing and homologous recombination during meiosis (see Chapter 2). When the chromosomes of a carrier of a balanced reciprocal translocation pair at meiosis (Fig. 5.11), they must form a quadrivalent to ensure proper alignment of homologous sequences (rather than the typical bivalents seen with normal chromosomes). In typical segregation, two of the four chromosomes in the quadrivalent go to each pole at anaphase; however, the chromosomes can segregate from this configuration in several ways, depending on which chromosomes go to which pole. Alternate segregation, the usual type of meiotic segregation, produces balanced gametes that have either a normal chromosome complement or contain the two reciprocal chromosomes. Other segregation patterns, however, always yield unbalanced gametes (see Fig. 5.11). Robertsonian Translocations. Robertsonian translocations are the most common type of chromosome rearrangement observed in our species. They involve two acrocentric chromosomes that fuse near the centromere region with loss of the short arms (see Fig. 5.10). Such translocations are nonreciprocal, and the resulting karyotype has only 45 chromosomes, including the translocation chromosome, which in effect is made up of the long arms of two acrocentric chromosomes. Because, as noted earlier, the short arms of all five pairs of acrocentric chromosomes consist largely of various classes of satellite DNA, as well as hundreds of copies of ribosomal RNA genes, loss of the short arms of two B Quadrivalent formation in meiosis A Chromosomes A B der(A) der(B) C Segregation and gametes Adjacent-1 Unbalanced Unbalanced Unbalanced Unbalanced Normal Balanced Adjacent-2 Alternate Figure 5.11 (A) Diagram illustrating a balanced translocation between two chromosomes, involving a reciprocal exchange between the distal long arms of chromosomes A and B. (B) Formation of a quadrivalent in meiosis is necessary to align the homologous segments of the two derivative chromosomes and their normal homologues. (C) Patterns of segregation in a carrier of the translocation, leading to either balanced or unbalanced gametes, shown at the bottom. Adjacent-­1 segregation (in red, top chromosomes to one gamete, bottom chromosomes to the other) leads only to unbalanced gametes. Adjacent-­2 segregation (in green, left chromosomes to one gamete, right chromosomes to the other) also leads only to unbalanced gametes. Only alternate segregation (in gray, upper left/­lower right chromosomes to one gamete, lower left/­upper right to the other) can lead to balanced gametes.
74 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE Translocations. Translocation involves the movement of chromosome segments between two chromosomes. There are two main types: reciprocal and...
Ch5 · Pt15 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 75 acrocentric chromosomes is not deleterious; thus the karyotype is considered to be balanced, despite having only 45 chromosomes. Robertsonian translocations are typically, although not always, pseudodicentric (see Fig. 5.10), reflecting the location of the breakpoint on each acrocentric chromosome. Although Robertsonian translocations can involve all combinations of the acrocentric chromosomes, two—designated rob(13;14)(q 10;q 10) and rob(14;21) (q 10;q 10)—are relatively common. The translocation involving 13q and 14q is found in ~1 in 1300 persons and is thus by far the most common chromosome rearrangement in our species. Rare individuals with two copies of the same type of Robertsonian translocation have been described; these phenotypically normal individuals have only 44 chromosomes and lack any normal copies of the acrocentrics involved, replaced by two copies of the translocation. Although a carrier of a Robertsonian translocation does not present with any obvious clinical phenotype, there is a risk for unbalanced gametes and, therefore, for unbalanced offspring. The risk for unbalanced offspring varies according to the particular Robertsonian translocation and the sex of the carrier parent; carrier females in general have a higher risk for transmitting the translocation to an affected child. The chief clinical importance of this type of translocation is that carriers of a Robertsonian translocation involving chromosome 21 are at risk for producing a child with translocation Down syndrome, as will be explored further in Chapter 6. Insertions. An insertion is another type of nonreciprocal translocation that occurs when a segment removed from one chromosome is inserted into a different chromosome or in a different location within the same chromosome, either in its usual orientation with respect to the centromere or inverted. Because they require three chromosome breaks, insertions are relatively rare. Segregation in an insertion carrier can produce offspring with duplication or deletion of the inserted segment, as well as normal offspring and balanced carriers. The average risk for producing an affected child can be up to 50%, and prenatal diagnosis is therefore indicated. Inversions. An inversion occurs when a single chromosome undergoes two breaks and is reconstituted with the segment between the breaks inverted. Inversions are of two types (Fig. 5.12): paracentric, in which both breaks occur in one arm (Greek para, beside the centromere); and pericentric, in which there is a break in each arm (Greek peri, around the centromere). Pericentric inversions can be easier to identify cytogenetically when they change the proportion of the chromosome arms as well as the banding pattern. An inversion does not usually cause an abnormal phenotype in carriers because it is a balanced rearrangement. Its medical significance is for the progeny; a carrier of either type of inversion is at risk for producing unbalanced gametes and offspring. When an inversion is present, a loop needs to form to allow alignment and pairing of homologous segments of the normal and inverted chromosomes in meiosis I (see Fig. 5.12). When recombination occurs within the loop, gametes with balanced chromosome complements (either normal or with the inversion) and gametes with unbalanced complements are formed, depending on the location of recombination events. When the inversion is paracentric, the unbalanced recombinant chromosomes are acentric or dicentric and typically do not lead to viable offspring (see Fig. 5.12); thus the risk that a carrier of a paracentric inversion will have a liveborn child with an abnormal karyotype is very low. A pericentric inversion, on the other hand, can lead to the production of unbalanced gametes with both duplication and deletion of chromosome segments (see Fig. 5.12). The duplicated and deleted segments are those distal to the inversion. Overall, the risk for the child of a carrier of a pericentric inversion to have an unbalanced karyotype is estimated to be 5% to 10%. Each pericentric inversion, however, is associated with a particular risk, typically reflecting the size and content of the duplicated and deficient segments. Mosaicism for Chromosome Abnormalities Sometimes, two or more different chromosome complements are present among the cells in an individual; this situation is called chromosomal mosaicism. Such mosaicism is typically detected by conventional karyotyping, interphase FISH analysis, or chromosomal microarrays. A common cause of mosaicism is nondisjunction in an early postzygotic mitotic division. For example, a zygote with an additional chromosome 21 might lose the extra chromosome in a mitotic division and continue to develop as a 47,+21/46 mosaic. The effects of mosaicism on development vary with the timing of the nondisjunction event, the nature of the chromosome abnormality, the proportions of the different chromosome complements present, and the tissues affected. It is often believed that individuals who are mosaic for a given trisomy, such as mosaic Down syndrome or mosaic Turner syndrome, are less severely affected than nonmosaic individuals. When detected in lymphocytes, in cultured cell lines or in prenatal samples, it can be difficult to assess the significance of mosaicism, especially if it is identified prenatally. The proportions of the different chromosome complements seen in the tissue being analyzed (e.g., cultured amniocytes or lymphocytes) may not necessarily reflect the proportions present in other tissues or in the embryo during its early developmental stages. Mosaicism can also arise in cells in culture after they were taken from the individual; thus cytogeneticists
CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 75 acrocentric chromosomes is not deleterious; thus the karyotype is considered to be balanced, despite having only 45 chromosomes....
Ch5 · Pt16 76 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE attempt to differentiate between true mosaicism, present in the individual, and pseudomosaicism, which has occurred in the laboratory. The distinction between these types is not always easy or certain and can lead to major interpretive difficulties in prenatal diagnosis (see Box 2 earlier and Chapter 18). Incidence of Chromosome Anomalies Visible by Karyotype Analysis The incidence of different types of chromosomal aberration has been measured in a number of large population surveys and was summarized earlier in Fig. 5.7. The major number disorders of chromosome observed in liveborns are three autosomal trisomies (21, 18, and 13) and four types of sex chromosomal aneuploidy: Turner syndrome (usually 45,X), Klinefelter syndrome (47,XXY), 47,XYY, and 47,XXX (see Chapter 6). Triploidy and tetraploidy account for only a small percentage of cases, typically in spontaneous abortions. The classification and incidence of chromosomal defects measured in these surveys can be used to consider the fate of 10,000 conceptuses (Table 5.2). TABLE 5.2 Outcome of 10,000 Pregnanciesa Outcome Pregnancies Spontaneous Abortions (%) Live Births Total 10,000 1500 (15) 8500 Normal chromosomes 9200 750 (8) 8450 Abnormal chromosomes 800 750 (94) 50 Specific Abnormalities Triploid or tetraploid 170 170 (100) 0 45,X 140 139 (99) 1 Trisomy 16 112 112 (100) 0 Trisomy 18 20 19 (95) 1 Trisomy 21 45 35 (78) 10 Trisomy, other 209 208 (99.5) 1 47,XXY, 47,XXX, 47,XYY 19 4 (21) 15 Unbalanced rearrangements 27 23 (85) 4 Balanced rearrangements 19 3 (16) 16 Other 39 37 (95) 2 a These estimates are based on observed frequencies of chromosome abnormalities in spontaneous pregnancy losses and in liveborn infants. It is likely that the frequency of chromosome abnormalities in all conceptuses is much higher than this because many undergo spontaneous pregnancy loss before they are recognized clinically. A B A B C Invert D A C B D Paracentric A B C Invert D A C B D Pericentric A B C D A B C A D B C D A C B D A B C D D B C D Inviable Balanced Unbalanced Balanced A C B D A B C A Figure 5.12 Crossing over within inversion loops formed at meiosis I in carriers of a chromosome with segment B-­C inverted. (A) Paracentric inversion. Gametes formed after the second meiotic division usually contain either a normal (A-­B-­C-­D) or a balanced (A-­C-­B-­D) copy of the chromosome because the acentric and dicentric products of the crossover are inviable. (B) Pericentric inversion. Gametes formed after the second meiotic division may be balanced (normal or inverted) or unbalanced. Unbalanced gametes contain a copy of the chromosome with a duplication or a deletion of the material flanking the inverted segment (A-­B-­C-­A or D-­B-­C-­D).
76 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE attempt to differentiate between true mosaicism, present in the individual, and pseudomosaicism, which has occurred in the laboratory. The di...
Ch5 · Pt17 CHAPTER 5 — Principles of Clinical Cytogenetics and Genome Analysis 77 Live Births The overall incidence of chromosome abnormalities in newborns is ~1 in 154 births (0.65%) (see Fig. 5.7). Most of the autosomal abnormalities can be diagnosed at birth, but most sex chromosome abnormalities, with the exception of Turner syndrome, are not recognized clinically until puberty (see Chapter 6). Unbalanced rearrangements are likely to come to clinical attention because of abnormal appearance and neurodevelopmental abnormalities in the affected individual. In contrast, balanced rearrangements are rarely identified clinically unless a carrier of a rearrangement has a child with an unbalanced chromosome complement and family studies are undertaken. Spontaneous Pregnancy Loss The frequency of chromosome abnormalities in spontaneous pregnancy loss is at least 40% to 50%, and the kinds of abnormalities differ in a number of ways from those seen in liveborns. Somewhat surprisingly, the single most common abnormality in pregnancy loss is 45,X (the same abnormality found in Turner syndrome), which accounts for nearly 20% of chromosomally abnormal spontaneous pregnancy losses but less than 1% of chromosomally abnormal live births (see Table 5.2). Another difference is the distribution of kinds of trisomy; for example, trisomy 16 is not seen at all in live births but accounts for approximately one-­ third of trisomies in pregnancy losses. CHROMOSOME AND GENOME ANALYSIS IN CANCER We have focused in this chapter on constitutional chromosome abnormalities that are seen in most or all of the cells in the body and derive from changes in chromosome structure or number that have been transmitted from a parent (either inherited or occurring de novo in the germline of a parent) or that have occurred in the zygote in early mitotic divisions. However, such chromosome abnormalities also occur in somatic cells throughout life and are a hallmark of cancer, both in hematologic neoplasias (e.g., leukemias and lymphomas) and in the context of solid tumor progression. An important area in cancer research is the delineation of chromosomal and genomic changes in specific forms of cancer and the relation of the breakpoints of the various structural rearrangements to the process of oncogenesis. The chromosome and genomic changes seen in cancer cells are numerous and diverse. The association of cytogenetic and genome analysis with tumor type and with the effectiveness of therapy is already an important part of the management of patients with cancer; these are discussed further in Chapter 16. ACKNOWLEDGMENT We thank Mary Ann George and Mary Shago for contributing to this chapter. GENERAL REFERENCES Gardner RJM, Armor D: Gardner and Sutherland’s chromosome abnormalities and genetic counseling, ed 5, New York, 2018, Oxford University Press. Feuk L, Carson AR, Scherer SW: Structural variation in the human genome, Nat Rev Genet 7(2):85–­97, 2006. Mc Gowan-­Jordan J, Hastings RJ, Adelaide SM, editors: ISCN 2020: an international system for human cytogenetic nomenclature, Basel, 2020, Karger. REFERENCES FOR SPECIFIC TOPICS Baldwin EK, May LF, Justice AN, et al: Mechanisms and consequences of small supernumerary marker chromosomes: from Barbara Mc Clintock to modern genetic-counseling issues, Am J Hum Genet 82:398–­410, 2008. Coulter ME, Miller DT, Harris DJ, et al: Chromosomal microarray testing influences medical management, Genet Med 13:770–­776, 2011. Dan S, Chen F, Choy KW, et al: Prenatal detection of aneuploidy and imbalanced chromosomal arrangements by massively parallel sequencing, PLo S ONE 7:e 27835, 2012. Feng W, Chakraborty A: Fragility extraordinaire: unsolved mysteries of chromosome fragile sites, Adv Exp Med Biol 1042:489–­526, 2017. Firth HV, Richards SM, Bevan AP, et al: DECIPHER: database of chromosomal imbalance and phenotype in humans using Ensembl resources, Am J Hum Genet 84:524–­533, 2009. Green RC, Rehm HL, Kohane IS: Clinical genome sequencing. In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine ed 2, New York, 2013, Elsevier, pp 102–­122. Higgins AW, Alkuraya FS, Bosco AF, et al: Characterization of apparently balanced chromosomal rearrangements from the Developmental Genome Anatomy Project, Am J Hum Genet 82:712–­722, 2008. Ledbetter DH, Riggs ER, Martin CL: Clinical applications of whole-­ genome chromosomal microarray analysis. In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine ed 2, New York, 2013, Elsevier, pp 133–­144. Lee C: Structural genomic variation in the human genome. In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine ed 2, New York, 2013, Elsevier, pp 123–­132. Miller DT, Adam MP, Aradhya S, et al: Consensus statement: chromosomal microarray is a first-­tier clinical diagnostic test for individuals with developmental disabilities or congenital anomalies, Am J Hum Genet 86:749–­764, 2010. Nagaoka SI, Hassold TJ, Hunt PA: Human aneuploidy: mechanisms and new insights into an age-­old problem, Nat Rev Genet 13:493–­ 504, 2012. Reddy UM, Page GP, Saade GR, et al: Karyotype versus microarray testing for genetic abnormalities after stillbirth, N Engl J Med 367:2185–­2193, 2012. Riggs ER, Andersen EF, Cherry AM, et al: Technical standards for the interpretation and reporting of constitutional copy-­number variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource (Clin Gen), Genet Med 22(2):245–­257, 2020. South ST, Lee C, Lamb AN, et al: ACMG standards and guidelines for constitutional cytogenomic microarray analysis, including postnatal and prenatal applications: revision, Genet Med 15:901-909, 2013. Talkowski ME, Ernst C, Heilbut A, et al: Next-­generation sequencing strategies enable routine detection of balanced chromosome rearrangements for clinical diagnostics and genetic research, Am J Hum Genet 88:469–­481, 2011. Trost B, Loureiro LO, Scherer SW: Discovery of genomic variation across a generation, Hum Mol Genet 30(R2):R174–­R186, 2021. https://­doi.org/­10.1093/­hmg/­ddab 209
.
Ch5 · Pt18 78 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE PROBLEMS 1. You send a blood sample from an infant with multiple congenital anomalies to the chromosome laboratory for analysis. The laboratory identifies two copy number variants by chromosome microarray, arr[GRCh 38] 7q33(136240808_159345973)x 3,18q12.3(45466214_ 80373285)x 1. G-banding analysis indicates the child’s karyotype is 46,XY,der(18)t(7;18)(q 33;q 12.3) a. What do these results mean? b. The laboratory asks for blood samples from the clinically normal parents for analysis. Why? c. The laboratory reports the mother's karyotype as 46,XX and the father's karyotype as 46,XY,t(7;18)(q 33;q 12.3). What does the latter karyotype mean? Referring to the normal chromosome ideograms in Figure 5.2, sketch the translocation chromosome or chromosomes in the father and in his son. Sketch these chromosomes in meiosis in the father. What kinds of gametes can he produce? 2. A spontaneous pregnancy loss is found to have trisomy 18. a. What proportion of pregnancies with trisomy 18 are spontaneously lost? b. What is the risk that the parents will have a liveborn child with trisomy 18 in a future pregnancy? 3. A newborn child with Down syndrome, when karyotyped, is found to have two cell lines: 70% of her cells have a 47,XX,+21 karyotype, and 30% are normal 46,XX. When did the nondisjunction event likely occur? What is the prognosis for this child? 4. Which of the following persons is or is not expected to be phenotypically normal? For questions 4c and 4d, what are the reproductive risks associated with the chromosome rearrangements assuming the other parent is chromosomally normal? a. A female with 47 chromosomes, including a small supernumerary marker chromosome derived from the centromeric region of chromosome 15 b. A female with the karyotype 47,XX,+13 c. A person with a balanced reciprocal translocation d. A person with a pericentric inversion of chromosome 6 5. For each of the following, state whether chromosome/ genome analysis is indicated or not. For which family members, if any? For what kind of chromosome abnormality might the family in each case be at risk? a. A pregnant 29-year-old woman and her 41-year-old husband, with no history of genetic defects b. A pregnant 41-year-old woman and her 29-year-old husband, with no history of genetic defects c. A couple whose only child has Down syndrome d. A couple whose only child has cystic fibrosis e. A couple who has two boys with global developmental delay and severe intellectual disability 6. Explain the nature of the chromosome abnormality and the method of detection indicated by the following nomenclature. a. 46,XX,inv(X)(q 21q26) b. 46,XX,del(1)(p 36.2) c. 46,XX.ish del(15)(q 11.2q11.2)(SNRPN−,D15S10−) d. 46,XX,del(15)(q 11.2q13).ish del(15)(SNRPN−,D15 S10−) e. 46,XX.arr[GRCh 38] 1p36.33p36.32(1755688_263 3531)x 3 f. 47,XY,+mar.ish der(8)(D8Z1+) g. 46,XX,der(13;21)(q 10;q 10),+21 h. 45,XY,der(13;21)(q 10;q 10) 7. A laboratory performs microarray on a 4 year old male with learning disabilities, ventricular septal defect and immune deficiency. This analysis identifies a deletion affecting chromosome 22 at position 18874431 to 20348930 using genome build GRCh 38, in chromosome band 22q11.21. Refer to the ISCN microarray nomenclature in Table 5.1 to describe this deletion.
78 THOMPSON AND THOMPSON GENETICS AND GENOMICS IN MEDICINE PROBLEMS 1. You send a blood sample from an infant with multiple congenital anomalies to the chromosome laboratory for analysis. The laborat...
Select a segment to play