← Genes are DNA
CONTENTS

1: Genes are DNA

Author

Georgeos Hardo

These lecture notes are based on Chapter 1 of Lewin’s Genes XII, figures are credited to Lewin’s Genes XII.

The genome

We call all the hereditary information in an organism its genome. In eukaryotes, the genome is made up of chromosomal DNA, but additionally contains DNA in organelles such as mitochondria and chloroplasts. Such DNA is termed organellar DNA.

In prokaryotes, such as bacteria, the genome is often found in a single, circular chromosome, but often also includes extra, smaller, pieces of DNA known as plasmids.

The genome is the carrier of information, and does not itself perform an active role in the development of the organism. Rather, cellular machinery acts on specific sequences in the DNA (which we will later learn are known as genes) to produce products. The complex series of interactions between the genome and the cellular machinery is known as gene expression.

TipRecap: The structure of DNA

Recall that DNA is a double helix, made up of two strands of nucleotides. Each nucleotide contains a sugar, a phosphate group, and a nitrogenous base (adenine, guanine, cytosine, or thymine) joined by covalent bonds; hydrogen bonds form between complementary bases on the two DNA strands.

Through the process of gene expression, the DNA sequence directs the production of all the ribonucleic acids (RNAs) in the cell, which in turn are used to produce the proteins of the organism. The process of information flow (which underpins gene expression), from DNA to RNA to protein, is known as the central dogma of molecular biology.

This is a tightly regulated process. Genes must be switched on and off in response to the needs of the cell. Therefore the genome must also carry information about the regulation of gene expression. We will learn about such regulatory elements later in the course.

Hierarchical organisation of the genome

The genome may be physically divided into units known as chromosomes. These are discrete units of the genome carrying many genes. Each consists of a very long molecule of duplex DNA and an approximately equal mass of proteins (in eukaryotes). It is visible as a morphological entity only during cell division. Therefore the ultimate definition of a genome is the sequence of the DNA in each chromosome.

WarningA note on gene expression
  • While all genes are made up of DNA, not all DNA is part of a gene. About 98% of the human genome does not encode protein; this non-coding DNA includes introns and non-coding RNA genes as well as intergenic DNA.
  • Further to this, while all genes code for an RNA, not all RNA codes for a protein! We call these non-coding RNAs. These RNAs typically have many different functions within the cell.

Further to this, the genome can be functionally split into genes (Figure 1). A gene is a sequence of DNA that encodes a single type of RNA. That RNA may encode a polypeptide chain.

Figure 1: Each chromosome consists of a single, long molecule of DNA within which are the sequences of individual genes.

Each discrete chromosome may contain a large number of genes. E. coli, for example, has a genome comprised of a single circular chromosome of approximately 4.6 million base pairs, and contains 4,000-5,500 genes (depending on the strain). Haploid human cells on the other hand contain 23 chromosomes, with a total of 3.2 billion base pairs, comprising approximately 20,000 protein-coding genes (the number of non-protein coding genes, those genes which only code for an RNA, is still not known).

While you may be tempted to think that the number of genes, or the size of a genome is correlated with the complexity of an organism, this is not the case. Take the example of rice, it has a genome of approximately 400 million base pairs (far smaller than the human genome), but contains 50,000 to 60,000 genes.

TipDNA/RNA size units

Common units for the size of DNA/RNA are:

  • 1 kb (kilobase) = 1,000 bases.
  • 1 Mb (Megabase) = 1,000,000 bases.
  • 1 Gb (Gigabase) = 1,000,000,000 bases.
Organism Genome Size (Mbp) Number of Chromosomes Number of Genes
E. coli 4.6 1 4,000-5,500
Human 3,200 23 20,000+
Rice 400 12 50,000-60,000

Genes can also exist in multiple forms within a genome, known as alleles. An allele is a variant form of a gene. In diploid organisms (those with two sets of chromosomes), such as humans, each gene has two alleles, one inherited from each parent.

Each chromosome is a linear array of genes, one after the other, and the particular location on the chromosome at which a gene is located is known as its locus. The alleles of a gene are the different forms of that gene found at that gene’s particular locus. In other words, there are generally two alleles per locus in a diploid organism. This pair of alleles is found at the same locus on the chromosome pair (which are known as homologous chromosomes).

Genes

Genes are the functional units of the genome, and determine the characteristics of an organism, such as its physical appearance, susceptibility to disease, and its behaviour. Genes can be classified by the type of product they encode:

  • protein-coding genes encode RNAs that are translated into polypeptides. The older term structural gene is sometimes used in these notes for a protein-coding gene.
  • non-protein-coding genes encode RNAs that function without being translated into polypeptides.

Separately, regulator genes encode products that regulate the expression of other genes. A regulator gene may therefore also be protein-coding, or it may encode a non-coding RNA.

To fully understand gene expression, we must always keep in mind the central dogma of molecular biology (Figure 2). But be sure to keep in mind two key points, which should help avoid confusion in the future:

WarningKey points about gene expression
  • Not all genes code for a protein. Some only code for an RNA.
  • Therefore all genes code for an RNA, but not all genes code for a protein.
  • Recall that DNA is double stranded. Only one of the strands on a gene is used to code for the RNA.
  • The final products of gene expression, be they RNAs or polypeptides, may themselves regulate the expression of other gene products.

In these notes, I will try to be explicit when I refer to a gene, calling it either a protein coding gene or a non-protein coding gene.

Figure 2: The central dogma of molecular biology: a gene encodes an RNA, which may encode a polypeptide

DNA is the Genetic Material of Bacteria and Viruses

Evidence for DNA as the Genetic Material of Bacteria

In 1928, Frederick Griffith, a British microbiologist, conducted an experiment that provided evidence for DNA as the genetic material of bacteria. He was studying the bacterium Streptococcus pneumoniae, which kills mice by causing pneumonia. The virulence (ability to cause damage) of the bacterium is determined by its capsular polysaccharide, which is found on the surface of the bacterium.

If the cell has a specific type of this surface polysaccharide, so called the “S” type (for its smooth appearance), it is able to evade destruction by the mouse’s immune system and cause damage. These cells are virulent. If the cell has the “R” type (so called for its rough appearance), it is avirulent and does not kill the mouse.

Griffith found that when he killed “S” type bacteria by heat treatment, and then injected the cells into the mouse, they were no longer able to harm the animal. However, when he mixed dead heat-killed “S” type bacteria with living “R” type bacteria, and jointly injected the cells into the mouse, it would die from a pneumonia infection. What he then found was that live “S” type bacteria could be isolated from the mouse’s blood (Figure 3).

Figure 3: Neither heat-killed S-type nor live R-type bacteria can kill mice, but simultaneous injection of both can kill mice just as effectively as the live S type.

This finding showed that some property of the dead “S” bacteria could transform the live “R” bacteria into “S” bacteria. This was known as the transforming principle, and showed that genetic material could be transferred from one bacterial strain to another. It was then proven that DNA is the carrier of the genetic information by isolating it from “S” bacteria, and adding it to live “R” bacteria and inspecting their cell surface (Figure 4).

Figure 4: The DNA of S-type bacteria can transform R-type bacteria into the same S type.

Evidence for DNA as the Genetic Material of Viruses

Having shown that DNA is the genetic material of bacteria, we now turn our attention to a specific class of viruses known as bacteriophages (often abbreviated to phages). These are viruses that infect bacteria.

NotePhage abundance

Bacteriophages are the most abundant organisms on Earth. There are an estimated trillion phages for every grain of sand in the world.

We will focus on a phage known as Phage T2, a virus that infects Escherichia coli bacteria. When phage particles are added to the bacteria, they attach to the outside surface, and some material enters the cell. Approximately 20 minutes later, the cell bursts open, releasing new phage particles, which go on to infect other cells. This is known as the lytic cycle (Figure 5).

Figure 5: The lytic cycle of a bacteriophage. The phage infects the bacteria, replicates its DNA, which is used to produce new phages. These then lyse the bacteria, releasing new phages.

In 1952, Martha Chase and Alfred Hershey were interested in understanding which material injected into the cell was responsible for the infection and reproduction of the viral particles. In other words, they wanted to know which molecule was the genetic information carrier for the virus.

To do this, they used a technique called radiolabelling, wherein they replaced the phosphorous atoms in the T2 phage’s DNA with radioactive phosphorus-32, and the sulfur atoms in the T2 phage’s protein with radioactive sulfur-35.

They infected bacteria with these radiolabelled T2 phages and allowed the infection to proceed for a short time. Blending then detached the adsorbed phage coats from the bacteria; subsequent centrifugation placed the bacteria in the pellet and the detached phage material in the supernatant.

Chase and Hershey found that the lysed fraction contained 80% of the sulfur-35 label, while the intact fraction contained 70% of the phosphorus-32 label. What they also found was that the progeny phages contained 30% of the phosphorus-32 label, but only 1% of the sulfur-35 label. What this neatly shows is that only DNA is injected into the bacterium, and can therefore become part of the progeny phages, proving that DNA is the genetic information carrier for these viruses (Figure 6).

Figure 6: The genetic material of phage T2 is DNA.

Evidence for DNA as the Genetic Material of Eukaryotes

When DNA is added to eukaryotic cells growing in culture, it can enter the cells, and some of this DNA may be expressed and produce new proteins. When a particular gene is used, its uptake into the cells would result in the production of the protein associated with that gene. This process is termed transfection. Transfected DNA may be expressed transiently without integration; in stable transfection, selected DNA is maintained through cell division, usually after genomic integration. Expression of this new DNA can also result in a new phenotype within the cell.

Figure 7: Eukaryotic cells can acquire a new phenotype as the result of transfection by added DNA.

Figure 7 shows that cells deficient for the thymidine kinase gene (which is required for DNA synthesis) can be rescued by the addition of a plasmid containing the thymidine kinase gene. Under selection, the surviving cells can proliferate, demonstrating that introduced genetic information can be maintained through cell division in this stable-transfection experiment.

NoteTransfection

While transfection is analogous to transformation, for historical reasons it was named differently, and that name stuck. To this day, we term the process of adding DNA to bacterium as transformation, and adding DNA to a eukaryotic cell as transfection.

The structure of nucleotides

The basic building block of nucleic acids (DNA and RNA) is the nucleotide, which has three components:

The nitrogenous base is a purine or pyrimidine ring, and the base is linked to the 1’ (pronounced one-prime) carbon on pentose sugar by a glycosidic bond. (from the N1_1 of pyrimidines, or the N9_9 of purines). The pentose sugar linked to a nitrogenous base is called a nucleoside.

Nucleic acids are named for the type of sugar in their backbone. DNA has a 2’-deoxyribose sugar backbone, and RNA has a 2’-ribose sugar backbone. The only difference is the addition of a hydroxyl group (-OH) on the 2’ carbon of the ribose sugar in RNA.

A nucleoside linked to a phosphate at the 5’ carbon is called a nucleotide. Figure 8 shows the arrangement of all three components which make up a nucleotide.

Figure 8: A nucleoside is a nitrogenous base linked to a pentose sugar. A nucleotide is a nucleoside linked to a phosphate group.

Each nucleic acid contains four types of nitrogenous bases. The same two purines, adenine (A) and guanine (G), are present in both DNA and RNA. The two pyrimidines in DNA are cytosine (C) and thymine (T); in RNA, uracil (U) is found instead of thymine.

The terminal nucleotide at one end of the chain has a free 5′ phosphate group, whereas the terminal nucleotide at the other end has a free 3′ hydroxyl group. It is conventional to write nucleic acid sequences in the 5′ to 3′ direction—that is, from the 5′ terminus at the left to the 3′ terminus at the right. Successive nucleotides joined together are termed a polynucleotide chain (shown in Figure 9).

Figure 9: A polynucleotide chain consists of a series of 5′–3′ sugar–phosphate links that form a backbone from which the bases protrude.

DNA is a Double Helix

In 1953, James Watson and Francis Crick proposed a model for the structure of DNA. They proposed that DNA is a double helix, made up of two strands of nucleotides. The work by Rosalind Franklin and her student Raymond Gosling provided the crucial evidence for this model, when they captured the X-ray diffraction pattern of DNA.

Watson and Crick had proposed that the two polynucleotide chains in the double helix associate by hydrogen bonding between the nitrogenous bases. Normally, G can hydrogen-bond most stably with C, whereas A can bond most stably with T . This hydrogen bonding between bases is described as base pairing, and the paired bases (G forming three hydrogen bonds with C, or A forming two hydrogen bonds with T) are said to be complementary. The Watson–Crick model has the two polynucleotide chains running in opposite directions, so they are said to be antiparallel. This is shown in Figure 10.

Figure 10: The double helix maintains a constant width because purines always face pyrimidines in the complementary A-T and G-C base pairs.

The sugar–phosphate backbones are on the outside of the double helix and carry negative charges on the phosphate groups. When DNA is in solution in vitro, the charges are neutralised by the binding of metal ions, typically Na+^+. In the cell, positively charged proteins provide some of the neutralising force. These proteins play important roles in determining the organisation of DNA in the cell.

The base pairs are on the inside of the double helix. They are flat and lie perpendicular to the axis of the helix. Using the analogy of the double helix as a spiral staircase, the base pairs form the steps (Figure 11). Proceeding up the helix, bases are stacked on one another like a pile of plates.

Figure 11: Flat base pairs lie perpendicular to the sugar–phosphate backbone.

Each base pair is rotated about 36° around the axis of the helix relative to the next base pair, so approximately 10 base pairs make a complete turn of 360°. The twisting of the two strands around each other forms a double helix with a minor groove that is about 12 Å (1.2 nm) across and a major groove that is about 22 Å (2.2 nm) across (Figure 12). In B-DNA (the form in which DNA is found in the cell), the double helix is said to be “right-handed”; the turns run clockwise as viewed along the helical axis.

Figure 12: The two strands of DNA form a double helix.

DNA Replication is Semiconservative

DNA’s structure is such that it can be replicated with high fidelity. Because the two strands of DNA are joined by only relatively weak hydrogen bonds, they are able to separate. The specificity of base pairing suggests that each of the two separated strands can act as the template for the synthesis of the other. Therefore the structure of DNA provides the means for its own replication (Figure 13).

Figure 13: Base pairing provides the mechanism for replicating DNA.

The top part of Figure 13 shows an unreplicated parental duplex with the original two parental strands. The lower part shows the two daughter duplexes produced by complementary base pairing. Each of the daughter duplexes is identical in sequence to the original parent duplex, containing one parental strand and one newly synthesized strand. The structure of DNA carries the information needed for its own replication. The consequences of this mode of replication, called semiconservative replication, are that parental duplex is replicated to form two daughter duplexes, each of which consists of one parental strand and one newly synthesized daughter strand. The unit conserved from one generation to the next is one of the two individual strands comprising the parental duplex.

The key point to note about semiconservative replication is that a single strand of DNA, once created, continues to exist in the daughter duplexes. None of the DNA from one strand ever becomes part of the other strand.

Evidence for Semiconservative Replication

In 1958, Meselson and Stahl showed that DNA replication is semiconservative by using a technique called gradient centrifugation. They created “heavy” parental DNA by growing E. coli in a medium containing a heavy isotope of nitrogen, then switched the growth medium to one with “light” nitrogen. Their findings were as follows:

  • At generation 0 (the parental DNA), before growth in the light medium, all of the DNA was heavy.
  • After one generation in the light medium, all of the DNA was “medium” weight, or hybrid, with one heavy strand and one light strand.
  • After two generations, there was a fraction of light DNA and a fraction of hybrid DNA, with no intermediate weights. This conclusively showed that the nucleotides in the two strands of DNA did not mix (shown in Figure 14).
Figure 14: The results of the Meselson and Stahl experiment. Generation 0 contained only heavy duplex DNA; after one generation in light medium, all DNA was hybrid; after two generations, there was a mixture of light and hybrid DNA.

These results also ruled out two other models of replication. Had replication been conservative, keeping the parental duplex whole beside an entirely new one, one generation would have given separate heavy and light bands. Had it been dispersive, mixing old and new DNA along each strand, two generations would have given a single band between hybrid and light.

Polymerases Act on Separated DNA Strands at the Replication Fork

Since we’ve established that DNA replication is semiconservative, the two strands need to be separated. This process is called denaturation, and is the disruption of the duplex. Since the two strands are held together by base-specific pairing, the process is reversible, and is called renaturation.

Only a small region of the duplex DNA is denatured at any time during replication, however. This region is known as the replication fork. The fork moves along the parental DNA duplex, exposing both parental strands as templates for the synthesis of two daughter strands (Figure 15).

Figure 15: The replication fork is the region of DNA in which there is a transition from the unwound parental duplex to the newly replicated daughter duplexes.

The synthesis of DNA is aided by enzymes called DNA polymerases which recognise the template strand and catalyse the addition of nucleotide subunits to the polynucleotide chain that is being synthesised. These enzymes are assisted by a number of other proteins, such as DNA helicase, which unwinds the DNA.

DNA polymerase adds nucleotides only to a 3’ end, so a new strand grows only in the 5’ to 3’ direction. Because the two parental strands are antiparallel, one daughter strand, the leading strand, is made continuously as the fork moves, while the other, the lagging strand, is made in short pieces, called Okazaki fragments, that are later joined.

Genetic Information can be Provided by DNA or RNA

Until now we have spoken about genomes in the context of duplex DNA only. However, in some viruses, the genomes may be made up of:

No cellular organism has been found to have a genome made of anything other than double stranded DNA.

The process by which RNA is synthesised by reading a DNA template is called transcription, and is carried out by RNA polymerase. The process by which a polypeptide is synthesised by reading an RNA template is called translation, and is carried out by the ribosome. These two processes, along with DNA replication, form the foundation of the central dogma of molecular biology. Most RNA viruses replicate RNA from RNA using an RNA-dependent RNA polymerase; retroviruses instead copy their RNA genome into DNA by reverse transcription, carried out by reverse transcriptase.

Figure 16: The central dogma states that information in nucleic acid can be perpetuated or transferred, but the transfer of information into a polypeptide is irreversible.

However, one process which does not occur in nature is the transfer of information from the polypeptide level and back to the nucleotide level.

Nucleic Acids Hybridise by Base Pairing

The concept of base pairing is central to all processes involving nucleic acids. Disruption of the hydrogen bonds which form base pairs is crucial to the function of a double-stranded nucleic acid, whereas the ability to form base pairs is essential for the activity of a single-stranded nucleic acid.

Figure 17: Base pairing occurs in duplex DNA, and also in intra-and intermolecular interactions in single-stranded RNA (or DNA).

Base pairing is generally intercompatible between complementary strands of the same nucleic acid (Figure 17), but single stranded DNA and single stranded RNA can also base pair to form double stranded structures known as DNA-RNA heteroduplexes

Mutations Change the Sequence of DNA

Mutations provide decisive evidence that DNA is the genetic material of life. When a change in the sequence of DNA causes an alteration in the sequence of a gene, the resulting gene expression product, for example a protein, could be altered. This could result in a change in the phenotype of the organism.

All organisms experience a certain number of mutations as the result of normal cellular operations, or random interactions with the environment, these are called spontaneous mutations. The rate at which these occur is different among species, and can be different among tissue types within the same species.

Mutations are rare events, and, of course, those that have deleterious effects are selected against during evolution. It is therefore difficult to observe large numbers of spontaneous mutants from natural populations.

Mutation rates can be increased by exposure to radiation, chemicals, and other agents, these are called mutagens, and the changes they cause are called induced mutations. Most mutagens work by modifying either a particular base of the DNA, or by becoming incorporated into the nucleic acid.

TipEukaryotic mutation rates

Parent-offspring genome sequencing directly estimates human germline mutation rates. One large trio study found about 63 autosomal de novo single-nucleotide variants per child on average, with the count varying especially with parental age.

In bacteria, the rate of mutations is approximately 10−6^{-6} per locus per generation. This corresponds an average of approximately 10−910^{-9}-10−1010^{-10} mutations per cell per base pair per generation.

Figure 18: Mutation rates at different scales in bacteria.

Mutations can affect single base pairs, or long sequences

Any base pair on the DNA can be mutated. A point mutation (Figure 19) is one that changes only a single base pair, and can be caused by two types of events:

  • Chemical modification of the DNA directly converting one base into another.
  • An error during the replication of DNA causes the wrong base to be inserted into a polynucleotide.
Figure 19: Mutations can be induced by chemical modification of a base.

Point mutations can be further divided into two types (Figure 20), depending on the nature of the base substitution:

  • Transition: the conversion of one pyrimidine into another pyrimidine, or one purine into another purine. This would result in a G-C pair being replaced by an A-T pair, or vice versa. This is the most common type of point mutation.
  • Transversion: the conversion of one pyrimidine into a purine, or vice versa. This would result in an A-T pair being replaced with a T-A pair OR a C-G pair. This is less common.
Figure 20: The two types of point mutation: transitions and transversions.

Point mutations were thought for a long time to be the principal means of change in individual genes. We now know, though, that insertions of short sequences are quite frequent. Often, the insertions are the result of transposable elements, which are sequences of DNA with the ability to move from one site to another.

An insertion within a coding region usually abolishes the activity of the gene because it can alter the reading frame; such an insertion is a frameshift mutation. (Similarly, a deletion within a coding region is usually a frameshift mutation.) Insertions of transposable elements can subsequently result in deletion of part or all of the inserted material, and sometimes of the adjacent regions.

Some mutations can be reversed

Figure 21 shows that the possibility of reversion mutations, or revertants, is an important characteristic that distinguishes point mutations and insertions from deletions:

  • A point mutation can revert either by restoring the original sequence (a true reversion) or by gaining a compensatory mutation elsewhere in the gene (a second-site reversion).
  • An insertion can revert by deletion of the inserted sequence.
  • A deletion of a sequence cannot revert in the absence of some mechanism to restore the lost sequence.
Figure 21: Point mutations and insertions can revert, but deletions cannot revert.