CONTENTS · FULL LECTURE

2: Prokaryotic transcription

Author

Georgeos Hardo

Introduction to transcription

Transcription is the process by which one DNA strand is read to produce an RNA chain complementary to the DNA template strand. The DNA which is read is called the template strand. The DNA template is read by a molecule called RNA polymerase in the 3’ –> 5’ direction. Therefore, the growing RNA chain is produced in the 5’ –> 3’ direction (Figure 1).

Figure 1: The function of RNA polymerase is to copy one strand of duplex DNA into RNA.

Transcription starts when an RNA polymerase binds to a special region called the promoter at the start of the gene. The promoter includes the first base pair that is transcribed into RNA, known as the start point. RNA polymerase moves along the template strand, synthesising RNA until it reaches a terminator sequence, where transcription ends. Thus, we call a transcription unit the region of the gene which extends from the promoter to the terminator, and codes for a single RNA molecule (Figure 2).

Figure 2: A transcription unit is a sequence of DNA transcribed into a single RNA, starting at the promoter and ending at the terminator.

Sequences prior to the start point, not transcribed, are called upstream, and those after the start point (which are transcribed) are called downstream. By convention, we write down the sequences so that transcription proceeds from the left (upstream) to the right (downstream), which corresponds to the 5’ –> 3’ direction of the RNA chain.

We also additionally often write the DNA sequence of a gene to only show the nontemplate strand, which has the same sequence as the RNA (except for the uracil instead of thymine). We also number the position of bases relative to the start point, which is the +1 position. The base before the start point is the -1 position, and so on (there is no 0 position).

The lifetime of prokaryotic transcripts

The initial product of transcription, which contains the original 5’ end is known as the primary transcript. RNA products, such as rRNA and tRNA primary transcripts undergo a maturation process by which an endonuclease cleaves off the ends, this increases the stability of the RNA molecule, allowing them to have lifetimes approaching that of the bacterium’s division time (approx 20 minutes for E. coli). mRNA primary transcripts, on the other hand, are quite unstable, and are subject to immediate attack by the cell’s endonucleases and exonucleases, and so have lifetimes of only 1-3 minutes. (In eukaryotes, mRNA is much more stable, and will be discussed later in the course.)

Transcription’s role in gene expression

Transcription is the first stage in gene expression, and is one of the key points of gene regulation. One of the most common ways in which genes are switched on and off is through regulation at the transcriptional level (as opposed to the translational level).

Two important questions which we will learn the answers to in this course are:

  • How does RNA polymerase (RNAP) find the promoter of a gene on the DNA?
  • How do regulator molecules interact with the RNAP to activate or inhibit transcription?

Transcription Occurs by Base Pairing in a “Bubble” of Unpaired DNA

RNA synthesis takes place within a transcription bubble, in which DNA is transiently separated into its single strands, revealing the template and allowing it to be used to direct the synthesis of RNA (Figure 3).

Figure 3: DNA strands separate to form a transcription bubble. RNA is synthesised by complementary base pairing with one of the DNA strands.
TipTranscription rate

It’s interesting to note that the rate of transcription in bacteria (40-50 nucleotides per second) is much slower than the rate of DNA replication (800 bp per second).

The RNA chain is synthesised starting at its 5’ end and growing in the 3’ direction. It’s important to note here that we are referring to the direction of the RNA chain itself, not the direction of the template, in other words, the 3’-OH group of the last nucleotide added to the chain reacts with the 5’-triphosphate of the incoming nucleotide.

RNA polymerase creates the transcription bubble when it binds to the promoter of a gene. RNA polymerase then moves along the DNA, with the bubble moving along with it, and the RNA chain growing in length. RNA polymerase is able to perform the base pairing and nucleotide addition by itself (Figure 4).

Figure 4: Transcription takes place in a bubble, in which RNA is synthesised by base pairing with one strand of DNA in the transiently unwound region. As the bubble progresses, the DNA duplex reforms behind it, displacing the RNA in the form of a single polynucleotide chain.

As RNA polymerase moves along the DNA template, it unwinds the duplex at the front of the bubble (the unwinding point), and the DNA automatically reforms the double helix at the back (the rewinding point). The length of the transcription bubble is about 12 to 14 bp, but the length of the RNA–DNA hybrid within the bubble is only 8 to 9 bp.

As the enzyme moves along the template, the DNA duplex reforms, and the RNA is displaced as a free polynucleotide chain. The last 14 ribonucleotides in the growing RNA are complexed with the DNA and/or the enzyme at any given moment.

The three stages of transcription

Transcription in prokaryotes can generally be divided into three stages:

Figure 5: Transcription has three stages: The enzyme binds to the promoter and melts DNA and remains stationary during initiation; moves along the template during elongation; and dissociates at termination.

Initiation

Initiation can be divided into multiple steps.

  • Template recognition: RNAP binds to the double stranded DNA at the promoter site. The enzyme forms a closed complex in which the DNA remains double stranded (no bubble is formed).
  • RNAP then locally unwinds the promoter region, including the transcription start point, to form the open complex.
  • Multiple rounds of abortive initiation occur, in which the RNAP enters into cycles of synthesis of short mRNA transcripts which are released before successful initiation occurs, and the RNAP clears the promoter region.

Elongation

Elongation involves the processive movement of the enzyme by disruption of base pairing in double stranded DNA. RNAP exposes the template strand for nucleotide addition and moves the transcription bubble with it as it moves downstream.

In recent years it has been found that during elongation, RNAP pauses, and even arrests at certain sequences. RNAP can even “backtrack” along the DNA template and remove a few nucleotides from the RNA chain in the case of errors or displacement of the 3’ end of the growing RNA chain.

Elongation factors: NusA and NusG

In addition to the core functions of RNAP, the elongation process is regulated by accessory proteins known as elongation factors. Two of the most important in bacteria are NusA and NusG. These factors bind to the elongating RNAP and influence its speed and pausing behaviour.

  • NusA primarily enhances transcriptional pausing, especially at sequences that form hairpin structures. By encouraging RNAP to pause at these sites, NusA plays a crucial role in facilitating intrinsic termination. It is also thought that NusA can compete with the sigma factor for binding to the core enzyme and help to displace it. Sigma release is common, but is not always required for elongation.
  • NusG, in contrast, is often considered an anti-pausing factor that increases the overall rate and processivity of transcription. It also plays a vital dual role by physically coupling the ribosome to the RNAP complex, directly linking transcription and translation. Furthermore, it is a key component in Rho-dependent termination, where it helps the Rho factor interact with the paused polymerase.

The Mechanism of Nucleotide Addition

The process of adding new nucleotides to the growing RNA chain is catalysed and monitored by the RNA polymerase enzyme itself. The synthesis proceeds in the 5’ to 3’ direction, meaning new nucleotides are always added to the 3’ end of the chain. The core chemical reaction involves the 3’–OH group of the last nucleotide in the chain attacking an incoming nucleoside 5’–triphosphate. During this reaction, the incoming nucleotide loses its two terminal phosphate groups (γ\gamma and β\beta). The remaining α\alpha phosphate group is then used to form a phosphodiester bond, linking the new nucleotide to the chain.

TipATP and nucleotide addition

The molecule we are all familiar with as the energy currency of the cell, ATP, is also the same substrate for the nucleotide addition of adenine to RNA, with two of its phosphate groups being cleaved, providing the energy for the formation of the phosphodiester bond. (In DNA synthesis, it is dATP which is added to the growing chain.)

Figure 6: The role of the phosphate groups in the nucleotide addition reaction. (Shown here for DNA nucleobases, but also valid for RNA nucleobases)

As RNA polymerase moves along the DNA template, it continuously unwinds the DNA at the front of the bubble and rewinds it at the back, maintaining the bubble’s structure and displacing the newly synthesised RNA strand.

Inside the enzyme, the DNA template makes a sharp 90° turn at the active site, which is facilitated by a “wall” of protein. The incoming nucleotides are thought to enter the active site through a secondary channel, sometimes called a pore. The rudder contacts the nascent RNA at the upstream edge of the RNA–DNA hybrid and helps stabilise the elongation complex; it does not melt the incoming DNA duplex.

Figure 7: DNA is forced to make a turn at the active site by a wall of protein. Nucleotides may enter the active site through a pore in the protein.

This process presents a challenge: the polymerase must maintain tight contact with the nucleic acids but must also be able to break and remake these contacts with each cycle of nucleotide addition. As shown in Figure 8, the specific bases that occupy the contact points within the enzyme change every time the enzyme moves forward one position. This is achieved through conformational changes in flexible parts of the enzyme, which folds around the incoming nucleotide to facilitate catalysis and then unfolds to allow the enzyme to move to the next position.

Figure 8: Movement of a nucleic acid polymerase requires breaking and remaking bonds to the nucleotides at fixed positions relative to the enzyme structure. The nucleotides in these positions change each time the enzyme moves a base along the template.

Termination

Termination involves the recognition of the sequences that signal RNAP to halt additional nucleotide addition. This sequence is known as the terminator.

Termination can occur prematurely, for example due to long pauses in transcription, which can cause the transcription bubble to collapse and disrupt the RNA-DNA heteroduplex.

Bacterial RNA Polymerase consists of Multiple Subunits

The best genetically and biochemically characterised RNA polymerases are from bacteria, namely from E. coli, Thermus aquaticus, and Thermus thermophilus.

However, in all bacteria, a single type of RNAP is responsible for the synthesis of rRNA, mRNA, and tRNA (unlike in eukaryotes). Bacteria are small, and so each cell only has approximately 13,000 RNAP molecules at a time.

Figure 9: Bacterial RNAPs have 5 types of subunits

We call the complete RNAP enzyme the holoenzyme. In E. coli it has a weight of approximately 460kD, and is made up of 5 subunits (Figure 9):

RNAP can be split into two parts, the core enzyme which is composed of the α2ββ′ω\alpha_2 \beta \beta' \omega subunits, and the holoenzyme, which is the complete enzyme including the sigma factor (α2ββ′ωσ\alpha_2 \beta \beta' \omega \sigma).

Figure 10: The upstream face of the core RNA polymerase, illustrating the “crab claw” shape of the enzyme.

The β\beta and β′\beta' subunits

The β\beta and β′\beta' subunits account for most of the mass of the core enzyme. Interestingly, their amino acid sequences and 3D structures are conserved with the largest subunits of the RNAPs from all other domains of life, bacteria, archaea, and eukaryotes. This indicates that the basic features of transcription are shared amongst all organisms on earth.

β\beta and β′\beta' form:

  • the active centre of the enzyme,
  • the main channel through which the DNA passes during the transcription cycle,
  • the secondary channel through which the substrate ribonucleotides enter the enzyme (on their way to the active site)
  • and the exit channel through which the RNA leaves the enzyme.

Mutations to the genes which encode these subunits, rpoB and rpoC have been shown to affect all stages of transcription.

The α\alpha subunits

The dimer formed by the α\alpha subunits serve as the scaffold for the assembly of the core enzyme. The C-terminus of the subunits also directly contact the promoter, and thereby contribute somewhat to promoter recognition.

σ\sigma factors

The core enzyme has a general affinity for DNA, primarily through electrostatic interactions between the protein (basic) and the DNA (acidic). When bound to DNA in this fashion, the DNA remains in duplex form. This brings us to a key point:

  • The core enzyme can synthesise RNA on a DNA template, but alone it cannot recognise promoters.

The form of the enzyme responsible for actually initiating transcription from promoters is called the holoenzyme (α2ββ′ωσ\alpha_2 \beta \beta' \omega \sigma).

Figure 11: Core enzyme binds indiscriminately to any DNA. Sigma factor reduces the affinity for sequence-independent binding and confers specificity for promoters.

The addition of the σ\sigma factor to the core enzyme is what allows the enzyme to both recognise the promoter sequence, and initiate transcription. The σ\sigma also helps to reduce non-specific binding of RNAP to the DNA, so much so that the association constant to random DNA is reduced by a factor of 1,000, but the association constant to the promoter sequence is increased by a factor of 1,000.

σ\sigma factors are an example of a transcription factor. Transcription factors are proteins which modulate the activity of RNAP for specific promoters.

RNAP-promoter binding rates

The rate at which the holoenzyme binds to different promoters varies widely, and thus this is an important parameter in determining promoter strength; that is, the efficiency of an individual promoter in initiating transcription.

How does RNAP find the promoter?

How does RNAP find promoters amongst the approximately 4.6 million base pairs of DNA in the E. coli genome? It was thought that RNAP would spend most of its time non-specifically interacting with random portions of DNA, in the hopes that it finds a promoter to bind to. Kinetic measurements of RNAP’s binding and unbinding rates have shown that this mechanism of finding a promoter is far too slow.

In actual fact, the process is probably sped up by the fact that RNAP can bind strongly anywhere on the genome, and not just to specific promoter regions. An alternative model has been shown to be much more likely:

These proposed mechanisms may allow RNAP to search locally along DNA or move between nearby DNA segments.

Sigma Factors Control Binding to DNA by Recognising Promoter Sequences

Recall that promoters themselves are not transcribed, but are simply sequences of DNA which RNAP recognises and binds to. The sequence of the promoter itself defines its function. This is an example of a cis-acting site.

RNAP typically can span approximately 75bp of DNA when bound, however it has been found that across that 75bp there is very little consensus across promoters.

Only a few short sequences are significantly conserved across promoters. There are two 6-bp elements referred to as the:

Their numbers refer to their positions relative to the transcription start site (they are upstream).

These two elements are usually the most important for promoter recognition since they have been shown to interact with the σ\sigma factor, but many others are important, and are shown in Figure 13

Figure 13: DNA elements and RNA polymerase modules that contribute to promoter recognition by sigma factor.

It is also important to note that mutations in the promoter sequence can have a significant effect on the strength of the promoter:

Mutations in a promoter which bring the sequence closer to the consensus are likely to increase the strength of the promoter, and vice versa. Additionally, mutations in the spacer which bring the length closer to 17bp are likely to increase the strength of the promoter, and vice versa. In other words, modifying the consensus sequence of the promoter can reduce the strength of the promoter.

Different sigma factors bind to different consensus sequences

E. coli has seven sigma factors, each of which causes RNA polymerase to initiate at a set of promoters defined by specific sequences. Sigma-70-family factors generally recognise –35 and –10 elements, whereas sigma 54 recognises distinct –24 and –12 elements. These sigma factors control the expression of sets of genes associated with different cellular functions under different environmental conditions. These specialised sigma factors bind the promoters of genes appropriate to the environmental conditions, increasing the transcription of those genes.

Below is a table of the different sigma factors and their associated promoters.

Subunit (Gene) Size (Number of Amino Acids) Approximate Number of Promoters Promoter Sequence Recognised
Sigma 70 (rpoD) 613 1,000 TTGACA—16 to 18 bp—TATAAT
Sigma 54 (rpoN) 477 5 TGGCACG—N4—TTGC (–24/–12 elements)
Sigma S (rpoS) 330 100 TTGACA—16 to 18 bp—TATAAT
Sigma 32 (rpoH) 284 30 CCCTTGAA—13 to 15 bp—CCCGATNT
Sigma F (rpoF) 239 40 CTAAA—15 bp—GCCGATAA
Sigma E (rpoE) 202 20 GAA—16 bp—YCTGA
Sigma FecI (fecI) 173 1–2 ?

Two most important sigma factors to note are:

  • Sigma 70: the “housekeeping” sigma factor or also called as primary sigma factor (Group 1), transcribes most genes in growing cells. Every cell has a “housekeeping” sigma factor that keeps essential genes and pathways operating.
  • Sigma 38: the starvation/stress response sigma factor. When the cell is starved for nutrients, sigma 38 RNAP holoenzymes bind to their consensus sequence, and transcribe the genes which are required to respond to environmental stress.

Differences in the protein structure of the sigma factors are what allow them to bind to different promoter sequences (Figure 14).

Figure 14: The structure of sigma factor in the context of the holoenzyme: −10 and −35 interactions. Sigma factor is extended and its domains are connected by flexible linkers.

Transcription Termination

Transcription termination typically occurs after a long pause in transcription, or at the site of a terminator sequence. Termination requires that all hydrogen bonds holding together the RNA-DNA heteroduplex be broken, allowing the DNA duplex to reform.

There are two types of terminators in prokaryotes, intrinsic terminators and terminators that require a rho factor protein. We will start with intrinsic terminators.

Intrinsic terminators

Intrinsic terminators are so named because they are not dependent on any other protein to stop transcription. Termination depends only on the RNA product, and are typically found in G+C rich regions of the transcript, which fold to form a hairpin structure (Figure 15).

Figure 15: Intrinsic terminators include palindromic regions that form hairpins varying in length from 7 to 20 bp. The stem-loop structure includes a G-C–rich region and is followed by a run of U residues.
TipTranscripts vs Genes

You may recall previously that E. coli has over 4200 genes, so how can it be that 1,100 unique transcripts account for over half of them? The answer is simply that many genes are transcribed as a single transcript known as an operon, this will be covered later in the course.

These hairpins typically have a second feature: a series of up to seven uracil residues (thymine on the DNA nontemplate strand and adenine on the template strand). It has been found that in E. coli, approximately 1,100 sequences in the genome fit this pattern, meaning that approximately half of the terminators are intrinsic.

Figure 16: The DNA sequences required for termination are located upstream of the terminator sequence. Formation of a hairpin in the RNA may be necessary.

Rho-dependent terminators

The other type of terminator is the rho-dependent terminator. These terminators are dependent on the presence of a protein called the rho factor to stop transcription.

Rho factor first binds to a sequence within the transcript upstream of the termination site. This is known as a rho utilisation site (rut site for short). The rho factor then tracks along the RNA until it catches up to the RNAP. When RNAP reaches the termination site, rho freezes the structure of the polymerase, invades the exit channel, and destabilises the enzyme, which causes the transcript to be released (Figure 17).

Figure 17: Rho factor binds to RNA at a rut site and translocates along RNA until it reaches the RNA–DNA hybrid in RNA polymerase, where it releases the RNA from the DNA.

The fact that rho needs to translocate from the rut site to the point of termination suggests that it needs to somehow catch up to RNAP during transcription.

Rho has a C-terminal ATPase domain, allowing it to translocate along the RNA. Under normal conditions, rho can translocate faster than RNAP. RNA-polymerase pauses nevertheless to provide favourable sites and time for efficient rho-dependent termination (Figure 17).

The Sigma Cycle

The process by which the sigma (σ\sigma) factor associates with the core RNA polymerase (RNAP) to initiate transcription and is often released after promoter escape is known as the sigma cycle. Promoter escape weakens sigma–core interactions, but sigma release is not always required for elongation and some complexes retain sigma far downstream. The cycle can be broken down into several steps:

  1. Holoenzyme Formation: A free σ\sigma factor associates with a core enzyme (α2ββ′ω\alpha_2 \beta \beta' \omega) to form the RNAP holoenzyme. This complex is competent to bind specifically to promoter sequences.
  2. Promoter Binding and Initiation: The holoenzyme binds to a promoter, forming a closed and then an open complex. It then begins synthesising RNA, often going through several rounds of abortive initiation where short transcripts are made and released.
  3. Promoter Clearance and Possible σ\sigma Release: Once the transcript reaches a length of approximately 10-12 nucleotides, the polymerase undergoes promoter clearance. Sigma–core interactions weaken and sigma is often released, but some elongating complexes retain sigma.
  4. Elongation: RNA polymerase moves down the DNA template and continues to synthesise the RNA transcript.
  5. Recycling: When sigma is released, it becomes available to bind to another core enzyme and start a new cycle of transcription initiation at another promoter.
Figure 18: A pool of σ\sigma factors compete for binding to core RNAP to form a holoenzyme, which binds promoter DNA to form an open complex (OC). Sigma is often released from the elongating complex (EC) and can be reused to direct transcription initiation by other core RNAP molecules, although release is not required for elongation. Competitive binding to the EC by NusA may help displace σ\sigma from the EC.

Detecting DNA-Protein Interactions with DNA Footprinting

One can learn about the specific binding and ability of proteins, in our case RNAP, to recognise DNA sequences using a technique known as DNA footprinting.

Say you have a stretch of DNA to which RNAP is bound. This DNA can be digested with an endonuclease, and the resulting fragments can be separated by gel electrophoresis. The key point to note is that DNA which is obscured by the protein will not be digested, and will therefore remain intact.

By comparing the pattern of the digested DNA to a control stretch of DNA, to which no protein was bound, one can learn about the specific sequence the protein was bound to.

In practice, the endonuclease will digest the DNA into fragments of varying length. By only doing a partial digestion, one will have a mixture of all possible fragment lengths.

By running these fragments through polyacrylamide gel electrophoresis, one can separate the fragments by size with single nucleotide resolution. Figure 19 shows 31 bands. In the protected fragments, bonds cannot be broken when they are obscured by the protein. The absence of bands 9 through 18 in the figure identifies a protein-binding site covering the region located 9 to 18 bases from the labelled end of the DNA.

Figure 19: Footprinting identifies DNA-binding sites for proteins by their protection against digestion.

One can sequence the original stretch of DNA and identify the region located 9 to 18 bases from the labelled end of the DNA, and thus know the specific sequence the protein was bound to.