1: Genes are DNA
These lecture notes are based on Chapter 1 of Lewin’s Genes XII, figures are credited to Lewin’s Genes XII.
The genome
We call all the hereditary information in an organism its genome. In eukaryotes, the genome is made up of chromosomal DNA, but additionally contains DNA in organelles such as mitochondria and chloroplasts. Such DNA is termed organellar DNA.
In prokaryotes, such as bacteria, the genome is often found in a single, circular chromosome, but often also includes extra, smaller, pieces of DNA known as plasmids.
The genome is the carrier of information, and does not itself perform an active role in the development of the organism. Rather, cellular machinery acts on specific sequences in the DNA (which we will later learn are known as genes) to produce products. The complex series of interactions between the genome and the cellular machinery is known as gene expression.
Recall that DNA is a double helix, made up of two strands of nucleotides. Each nucleotide contains a sugar, a phosphate group, and a nitrogenous base (adenine, guanine, cytosine, or thymine) joined by covalent bonds; hydrogen bonds form between complementary bases on the two DNA strands.
Through the process of gene expression, the DNA sequence directs the production of all the ribonucleic acids (RNAs) in the cell, which in turn are used to produce the proteins of the organism. The process of information flow (which underpins gene expression), from DNA to RNA to protein, is known as the central dogma of molecular biology.
This is a tightly regulated process. Genes must be switched on and off in response to the needs of the cell. Therefore the genome must also carry information about the regulation of gene expression. We will learn about such regulatory elements later in the course.
Hierarchical organisation of the genome
The genome may be physically divided into units known as chromosomes. These are discrete units of the genome carrying many genes. Each consists of a very long molecule of duplex DNA and an approximately equal mass of proteins (in eukaryotes). It is visible as a morphological entity only during cell division. Therefore the ultimate definition of a genome is the sequence of the DNA in each chromosome.
- While all genes are made up of DNA, not all DNA is part of a gene. About 98% of the human genome does not encode protein; this non-coding DNA includes introns and non-coding RNA genes as well as intergenic DNA.
- Further to this, while all genes code for an RNA, not all RNA codes for a protein! We call these non-coding RNAs. These RNAs typically have many different functions within the cell.
Further to this, the genome can be functionally split into genes (Figure 1). A gene is a sequence of DNA that encodes a single type of RNA. That RNA may encode a polypeptide chain.
Each discrete chromosome may contain a large number of genes. E. coli, for example, has a genome comprised of a single circular chromosome of approximately 4.6 million base pairs, and contains 4,000-5,500 genes (depending on the strain). Haploid human cells on the other hand contain 23 chromosomes, with a total of 3.2 billion base pairs, comprising approximately 20,000 protein-coding genes (the number of non-protein coding genes, those genes which only code for an RNA, is still not known).
While you may be tempted to think that the number of genes, or the size of a genome is correlated with the complexity of an organism, this is not the case. Take the example of rice, it has a genome of approximately 400 million base pairs (far smaller than the human genome), but contains 50,000 to 60,000 genes.
Common units for the size of DNA/RNA are:
- 1 kb (kilobase) = 1,000 bases.
- 1 Mb (Megabase) = 1,000,000 bases.
- 1 Gb (Gigabase) = 1,000,000,000 bases.
| Organism | Genome Size (Mbp) | Number of Chromosomes | Number of Genes |
|---|---|---|---|
| E. coli | 4.6 | 1 | 4,000-5,500 |
| Human | 3,200 | 23 | 20,000+ |
| Rice | 400 | 12 | 50,000-60,000 |
Genes can also exist in multiple forms within a genome, known as alleles. An allele is a variant form of a gene. In diploid organisms (those with two sets of chromosomes), such as humans, each gene has two alleles, one inherited from each parent.
Each chromosome is a linear array of genes, one after the other, and the particular location on the chromosome at which a gene is located is known as its locus. The alleles of a gene are the different forms of that gene found at that gene’s particular locus. In other words, there are generally two alleles per locus in a diploid organism. This pair of alleles is found at the same locus on the chromosome pair (which are known as homologous chromosomes).
Genes
Genes are the functional units of the genome, and determine the characteristics of an organism, such as its physical appearance, susceptibility to disease, and its behaviour. Genes can be classified by the type of product they encode:
- protein-coding genes encode RNAs that are translated into polypeptides. The older term structural gene is sometimes used in these notes for a protein-coding gene.
- non-protein-coding genes encode RNAs that function without being translated into polypeptides.
Separately, regulator genes encode products that regulate the expression of other genes. A regulator gene may therefore also be protein-coding, or it may encode a non-coding RNA.
To fully understand gene expression, we must always keep in mind the central dogma of molecular biology (Figure 2). But be sure to keep in mind two key points, which should help avoid confusion in the future:
- Not all genes code for a protein. Some only code for an RNA.
- Therefore all genes code for an RNA, but not all genes code for a protein.
- Recall that DNA is double stranded. Only one of the strands on a gene is used to code for the RNA.
- The final products of gene expression, be they RNAs or polypeptides, may themselves regulate the expression of other gene products.
In these notes, I will try to be explicit when I refer to a gene, calling it either a protein coding gene or a non-protein coding gene.
DNA is the Genetic Material of Bacteria and Viruses
Evidence for DNA as the Genetic Material of Bacteria
In 1928, Frederick Griffith, a British microbiologist, conducted an experiment that provided evidence for DNA as the genetic material of bacteria. He was studying the bacterium Streptococcus pneumoniae, which kills mice by causing pneumonia. The virulence (ability to cause damage) of the bacterium is determined by its capsular polysaccharide, which is found on the surface of the bacterium.
If the cell has a specific type of this surface polysaccharide, so called the “S” type (for its smooth appearance), it is able to evade destruction by the mouse’s immune system and cause damage. These cells are virulent. If the cell has the “R” type (so called for its rough appearance), it is avirulent and does not kill the mouse.
Griffith found that when he killed “S” type bacteria by heat treatment, and then injected the cells into the mouse, they were no longer able to harm the animal. However, when he mixed dead heat-killed “S” type bacteria with living “R” type bacteria, and jointly injected the cells into the mouse, it would die from a pneumonia infection. What he then found was that live “S” type bacteria could be isolated from the mouse’s blood (Figure 3).
This finding showed that some property of the dead “S” bacteria could transform the live “R” bacteria into “S” bacteria. This was known as the transforming principle, and showed that genetic material could be transferred from one bacterial strain to another. It was then proven that DNA is the carrier of the genetic information by isolating it from “S” bacteria, and adding it to live “R” bacteria and inspecting their cell surface (Figure 4).
Evidence for DNA as the Genetic Material of Viruses
Having shown that DNA is the genetic material of bacteria, we now turn our attention to a specific class of viruses known as bacteriophages (often abbreviated to phages). These are viruses that infect bacteria.
Bacteriophages are the most abundant organisms on Earth. There are an estimated trillion phages for every grain of sand in the world.
We will focus on a phage known as Phage T2, a virus that infects Escherichia coli bacteria. When phage particles are added to the bacteria, they attach to the outside surface, and some material enters the cell. Approximately 20 minutes later, the cell bursts open, releasing new phage particles, which go on to infect other cells. This is known as the lytic cycle (Figure 5).
In 1952, Martha Chase and Alfred Hershey were interested in understanding which material injected into the cell was responsible for the infection and reproduction of the viral particles. In other words, they wanted to know which molecule was the genetic information carrier for the virus.
To do this, they used a technique called radiolabelling, wherein they replaced the phosphorous atoms in the T2 phage’s DNA with radioactive phosphorus-32, and the sulfur atoms in the T2 phage’s protein with radioactive sulfur-35.
They infected bacteria with these radiolabelled T2 phages and allowed the infection to proceed for a short time. Blending then detached the adsorbed phage coats from the bacteria; subsequent centrifugation placed the bacteria in the pellet and the detached phage material in the supernatant.
Chase and Hershey found that the lysed fraction contained 80% of the sulfur-35 label, while the intact fraction contained 70% of the phosphorus-32 label. What they also found was that the progeny phages contained 30% of the phosphorus-32 label, but only 1% of the sulfur-35 label. What this neatly shows is that only DNA is injected into the bacterium, and can therefore become part of the progeny phages, proving that DNA is the genetic information carrier for these viruses (Figure 6).
Evidence for DNA as the Genetic Material of Eukaryotes
When DNA is added to eukaryotic cells growing in culture, it can enter the cells, and some of this DNA may be expressed and produce new proteins. When a particular gene is used, its uptake into the cells would result in the production of the protein associated with that gene. This process is termed transfection. Transfected DNA may be expressed transiently without integration; in stable transfection, selected DNA is maintained through cell division, usually after genomic integration. Expression of this new DNA can also result in a new phenotype within the cell.
Figure 7 shows that cells deficient for the thymidine kinase gene (which is required for DNA synthesis) can be rescued by the addition of a plasmid containing the thymidine kinase gene. Under selection, the surviving cells can proliferate, demonstrating that introduced genetic information can be maintained through cell division in this stable-transfection experiment.
While transfection is analogous to transformation, for historical reasons it was named differently, and that name stuck. To this day, we term the process of adding DNA to bacterium as transformation, and adding DNA to a eukaryotic cell as transfection.
The structure of nucleotides
The basic building block of nucleic acids (DNA and RNA) is the nucleotide, which has three components:
- A nitrogenous base
- A sugar
- One or more phosphate groups
The nitrogenous base is a purine or pyrimidine ring, and the base is linked to the 1’ (pronounced one-prime) carbon on pentose sugar by a glycosidic bond. (from the N of pyrimidines, or the N of purines). The pentose sugar linked to a nitrogenous base is called a nucleoside.
Nucleic acids are named for the type of sugar in their backbone. DNA has a 2’-deoxyribose sugar backbone, and RNA has a 2’-ribose sugar backbone. The only difference is the addition of a hydroxyl group (-OH) on the 2’ carbon of the ribose sugar in RNA.
A nucleoside linked to a phosphate at the 5’ carbon is called a nucleotide. Figure 8 shows the arrangement of all three components which make up a nucleotide.
Each nucleic acid contains four types of nitrogenous bases. The same two purines, adenine (A) and guanine (G), are present in both DNA and RNA. The two pyrimidines in DNA are cytosine (C) and thymine (T); in RNA, uracil (U) is found instead of thymine.
The terminal nucleotide at one end of the chain has a free 5′ phosphate group, whereas the terminal nucleotide at the other end has a free 3′ hydroxyl group. It is conventional to write nucleic acid sequences in the 5′ to 3′ direction—that is, from the 5′ terminus at the left to the 3′ terminus at the right. Successive nucleotides joined together are termed a polynucleotide chain (shown in Figure 9).
DNA is a Double Helix
In 1953, James Watson and Francis Crick proposed a model for the structure of DNA. They proposed that DNA is a double helix, made up of two strands of nucleotides. The work by Rosalind Franklin and her student Raymond Gosling provided the crucial evidence for this model, when they captured the X-ray diffraction pattern of DNA.
Watson and Crick had proposed that the two polynucleotide chains in the double helix associate by hydrogen bonding between the nitrogenous bases. Normally, G can hydrogen-bond most stably with C, whereas A can bond most stably with T . This hydrogen bonding between bases is described as base pairing, and the paired bases (G forming three hydrogen bonds with C, or A forming two hydrogen bonds with T) are said to be complementary. The Watson–Crick model has the two polynucleotide chains running in opposite directions, so they are said to be antiparallel. This is shown in Figure 10.
The sugar–phosphate backbones are on the outside of the double helix and carry negative charges on the phosphate groups. When DNA is in solution in vitro, the charges are neutralised by the binding of metal ions, typically Na. In the cell, positively charged proteins provide some of the neutralising force. These proteins play important roles in determining the organisation of DNA in the cell.
The base pairs are on the inside of the double helix. They are flat and lie perpendicular to the axis of the helix. Using the analogy of the double helix as a spiral staircase, the base pairs form the steps (Figure 11). Proceeding up the helix, bases are stacked on one another like a pile of plates.
Each base pair is rotated about 36° around the axis of the helix relative to the next base pair, so approximately 10 base pairs make a complete turn of 360°. The twisting of the two strands around each other forms a double helix with a minor groove that is about 12 Å (1.2 nm) across and a major groove that is about 22 Å (2.2 nm) across (Figure 12). In B-DNA (the form in which DNA is found in the cell), the double helix is said to be “right-handed”; the turns run clockwise as viewed along the helical axis.
DNA Replication is Semiconservative
DNA’s structure is such that it can be replicated with high fidelity. Because the two strands of DNA are joined by only relatively weak hydrogen bonds, they are able to separate. The specificity of base pairing suggests that each of the two separated strands can act as the template for the synthesis of the other. Therefore the structure of DNA provides the means for its own replication (Figure 13).
The top part of Figure 13 shows an unreplicated parental duplex with the original two parental strands. The lower part shows the two daughter duplexes produced by complementary base pairing. Each of the daughter duplexes is identical in sequence to the original parent duplex, containing one parental strand and one newly synthesized strand. The structure of DNA carries the information needed for its own replication. The consequences of this mode of replication, called semiconservative replication, are that parental duplex is replicated to form two daughter duplexes, each of which consists of one parental strand and one newly synthesized daughter strand. The unit conserved from one generation to the next is one of the two individual strands comprising the parental duplex.
The key point to note about semiconservative replication is that a single strand of DNA, once created, continues to exist in the daughter duplexes. None of the DNA from one strand ever becomes part of the other strand.
Evidence for Semiconservative Replication
In 1958, Meselson and Stahl showed that DNA replication is semiconservative by using a technique called gradient centrifugation. They created “heavy” parental DNA by growing E. coli in a medium containing a heavy isotope of nitrogen, then switched the growth medium to one with “light” nitrogen. Their findings were as follows:
- At generation 0 (the parental DNA), before growth in the light medium, all of the DNA was heavy.
- After one generation in the light medium, all of the DNA was “medium” weight, or hybrid, with one heavy strand and one light strand.
- After two generations, there was a fraction of light DNA and a fraction of hybrid DNA, with no intermediate weights. This conclusively showed that the nucleotides in the two strands of DNA did not mix (shown in Figure 14).
These results also ruled out two other models of replication. Had replication been conservative, keeping the parental duplex whole beside an entirely new one, one generation would have given separate heavy and light bands. Had it been dispersive, mixing old and new DNA along each strand, two generations would have given a single band between hybrid and light.
Polymerases Act on Separated DNA Strands at the Replication Fork
Since we’ve established that DNA replication is semiconservative, the two strands need to be separated. This process is called denaturation, and is the disruption of the duplex. Since the two strands are held together by base-specific pairing, the process is reversible, and is called renaturation.
Only a small region of the duplex DNA is denatured at any time during replication, however. This region is known as the replication fork. The fork moves along the parental DNA duplex, exposing both parental strands as templates for the synthesis of two daughter strands (Figure 15).
The synthesis of DNA is aided by enzymes called DNA polymerases which recognise the template strand and catalyse the addition of nucleotide subunits to the polynucleotide chain that is being synthesised. These enzymes are assisted by a number of other proteins, such as DNA helicase, which unwinds the DNA.
DNA polymerase adds nucleotides only to a 3’ end, so a new strand grows only in the 5’ to 3’ direction. Because the two parental strands are antiparallel, one daughter strand, the leading strand, is made continuously as the fork moves, while the other, the lagging strand, is made in short pieces, called Okazaki fragments, that are later joined.
Genetic Information can be Provided by DNA or RNA
Until now we have spoken about genomes in the context of duplex DNA only. However, in some viruses, the genomes may be made up of:
- double stranded DNA
- single stranded DNA
- double stranded RNA
- single stranded RNA
No cellular organism has been found to have a genome made of anything other than double stranded DNA.
The process by which RNA is synthesised by reading a DNA template is called transcription, and is carried out by RNA polymerase. The process by which a polypeptide is synthesised by reading an RNA template is called translation, and is carried out by the ribosome. These two processes, along with DNA replication, form the foundation of the central dogma of molecular biology. Most RNA viruses replicate RNA from RNA using an RNA-dependent RNA polymerase; retroviruses instead copy their RNA genome into DNA by reverse transcription, carried out by reverse transcriptase.
However, one process which does not occur in nature is the transfer of information from the polypeptide level and back to the nucleotide level.
Nucleic Acids Hybridise by Base Pairing
The concept of base pairing is central to all processes involving nucleic acids. Disruption of the hydrogen bonds which form base pairs is crucial to the function of a double-stranded nucleic acid, whereas the ability to form base pairs is essential for the activity of a single-stranded nucleic acid.
Base pairing is generally intercompatible between complementary strands of the same nucleic acid (Figure 17), but single stranded DNA and single stranded RNA can also base pair to form double stranded structures known as DNA-RNA heteroduplexes
Mutations Change the Sequence of DNA
Mutations provide decisive evidence that DNA is the genetic material of life. When a change in the sequence of DNA causes an alteration in the sequence of a gene, the resulting gene expression product, for example a protein, could be altered. This could result in a change in the phenotype of the organism.
All organisms experience a certain number of mutations as the result of normal cellular operations, or random interactions with the environment, these are called spontaneous mutations. The rate at which these occur is different among species, and can be different among tissue types within the same species.
Mutations are rare events, and, of course, those that have deleterious effects are selected against during evolution. It is therefore difficult to observe large numbers of spontaneous mutants from natural populations.
Mutation rates can be increased by exposure to radiation, chemicals, and other agents, these are called mutagens, and the changes they cause are called induced mutations. Most mutagens work by modifying either a particular base of the DNA, or by becoming incorporated into the nucleic acid.
Parent-offspring genome sequencing directly estimates human germline mutation rates. One large trio study found about 63 autosomal de novo single-nucleotide variants per child on average, with the count varying especially with parental age.
In bacteria, the rate of mutations is approximately 10 per locus per generation. This corresponds an average of approximately - mutations per cell per base pair per generation.
Mutations can affect single base pairs, or long sequences
Any base pair on the DNA can be mutated. A point mutation (Figure 19) is one that changes only a single base pair, and can be caused by two types of events:
- Chemical modification of the DNA directly converting one base into another.
- An error during the replication of DNA causes the wrong base to be inserted into a polynucleotide.
Point mutations can be further divided into two types (Figure 20), depending on the nature of the base substitution:
- Transition: the conversion of one pyrimidine into another pyrimidine, or one purine into another purine. This would result in a G-C pair being replaced by an A-T pair, or vice versa. This is the most common type of point mutation.
- Transversion: the conversion of one pyrimidine into a purine, or vice versa. This would result in an A-T pair being replaced with a T-A pair OR a C-G pair. This is less common.
Point mutations were thought for a long time to be the principal means of change in individual genes. We now know, though, that insertions of short sequences are quite frequent. Often, the insertions are the result of transposable elements, which are sequences of DNA with the ability to move from one site to another.
An insertion within a coding region usually abolishes the activity of the gene because it can alter the reading frame; such an insertion is a frameshift mutation. (Similarly, a deletion within a coding region is usually a frameshift mutation.) Insertions of transposable elements can subsequently result in deletion of part or all of the inserted material, and sometimes of the adjacent regions.
Some mutations can be reversed
Figure 21 shows that the possibility of reversion mutations, or revertants, is an important characteristic that distinguishes point mutations and insertions from deletions:
- A point mutation can revert either by restoring the original sequence (a true reversion) or by gaining a compensatory mutation elsewhere in the gene (a second-site reversion).
- An insertion can revert by deletion of the inserted sequence.
- A deletion of a sequence cannot revert in the absence of some mechanism to restore the lost sequence.