2: Prokaryotic transcription
Introduction to transcription
Transcription is the process by which one DNA strand is read to produce an RNA chain complementary to the DNA template strand. The DNA which is read is called the template strand. The DNA template is read by a molecule called RNA polymerase in the 3’ –> 5’ direction. Therefore, the growing RNA chain is produced in the 5’ –> 3’ direction (Figure 1).
Transcription starts when an RNA polymerase binds to a special region called the promoter at the start of the gene. The promoter includes the first base pair that is transcribed into RNA, known as the start point. RNA polymerase moves along the template strand, synthesising RNA until it reaches a terminator sequence, where transcription ends. Thus, we call a transcription unit the region of the gene which extends from the promoter to the terminator, and codes for a single RNA molecule (Figure 2).
Sequences prior to the start point, not transcribed, are called upstream, and those after the start point (which are transcribed) are called downstream. By convention, we write down the sequences so that transcription proceeds from the left (upstream) to the right (downstream), which corresponds to the 5’ –> 3’ direction of the RNA chain.
We also additionally often write the DNA sequence of a gene to only show the nontemplate strand, which has the same sequence as the RNA (except for the uracil instead of thymine). We also number the position of bases relative to the start point, which is the +1 position. The base before the start point is the -1 position, and so on (there is no 0 position).
The lifetime of prokaryotic transcripts
The initial product of transcription, which contains the original 5’ end is known as the primary transcript. RNA products, such as rRNA and tRNA primary transcripts undergo a maturation process by which an endonuclease cleaves off the ends, this increases the stability of the RNA molecule, allowing them to have lifetimes approaching that of the bacterium’s division time (approx 20 minutes for E. coli). mRNA primary transcripts, on the other hand, are quite unstable, and are subject to immediate attack by the cell’s endonucleases and exonucleases, and so have lifetimes of only 1-3 minutes. (In eukaryotes, mRNA is much more stable, and will be discussed later in the course.)
Transcription’s role in gene expression
Transcription is the first stage in gene expression, and is one of the key points of gene regulation. One of the most common ways in which genes are switched on and off is through regulation at the transcriptional level (as opposed to the translational level).
Two important questions which we will learn the answers to in this course are:
- How does RNA polymerase (RNAP) find the promoter of a gene on the DNA?
- How do regulator molecules interact with the RNAP to activate or inhibit transcription?
Transcription Occurs by Base Pairing in a “Bubble” of Unpaired DNA
RNA synthesis takes place within a transcription bubble, in which DNA is transiently separated into its single strands, revealing the template and allowing it to be used to direct the synthesis of RNA (Figure 3).
It’s interesting to note that the rate of transcription in bacteria (40-50 nucleotides per second) is much slower than the rate of DNA replication (800 bp per second).
The RNA chain is synthesised starting at its 5’ end and growing in the 3’ direction. It’s important to note here that we are referring to the direction of the RNA chain itself, not the direction of the template, in other words, the 3’-OH group of the last nucleotide added to the chain reacts with the 5’-triphosphate of the incoming nucleotide.
RNA polymerase creates the transcription bubble when it binds to the promoter of a gene. RNA polymerase then moves along the DNA, with the bubble moving along with it, and the RNA chain growing in length. RNA polymerase is able to perform the base pairing and nucleotide addition by itself (Figure 4).
As RNA polymerase moves along the DNA template, it unwinds the duplex at the front of the bubble (the unwinding point), and the DNA automatically reforms the double helix at the back (the rewinding point). The length of the transcription bubble is about 12 to 14 bp, but the length of the RNA–DNA hybrid within the bubble is only 8 to 9 bp.
As the enzyme moves along the template, the DNA duplex reforms, and the RNA is displaced as a free polynucleotide chain. The last 14 ribonucleotides in the growing RNA are complexed with the DNA and/or the enzyme at any given moment.
The three stages of transcription
Transcription in prokaryotes can generally be divided into three stages:
- Initiation: RNAP recognises the promoter and forms a transcription bubble.
- Elongation: the bubble moves along the DNA as the RNA transcript is synthesised.
- Termination: the RNA transcript is released, the bubble closes, and RNAP detaches from the DNA.
Initiation
Initiation can be divided into multiple steps.
- Template recognition: RNAP binds to the double stranded DNA at the promoter site. The enzyme forms a closed complex in which the DNA remains double stranded (no bubble is formed).
- RNAP then locally unwinds the promoter region, including the transcription start point, to form the open complex.
- Multiple rounds of abortive initiation occur, in which the RNAP enters into cycles of synthesis of short mRNA transcripts which are released before successful initiation occurs, and the RNAP clears the promoter region.
Elongation
Elongation involves the processive movement of the enzyme by disruption of base pairing in double stranded DNA. RNAP exposes the template strand for nucleotide addition and moves the transcription bubble with it as it moves downstream.
In recent years it has been found that during elongation, RNAP pauses, and even arrests at certain sequences. RNAP can even “backtrack” along the DNA template and remove a few nucleotides from the RNA chain in the case of errors or displacement of the 3’ end of the growing RNA chain.
Elongation factors: NusA and NusG
In addition to the core functions of RNAP, the elongation process is regulated by accessory proteins known as elongation factors. Two of the most important in bacteria are NusA and NusG. These factors bind to the elongating RNAP and influence its speed and pausing behaviour.
- NusA primarily enhances transcriptional pausing, especially at sequences that form hairpin structures. By encouraging RNAP to pause at these sites, NusA plays a crucial role in facilitating intrinsic termination. It is also thought that NusA can compete with the sigma factor for binding to the core enzyme and help to displace it. Sigma release is common, but is not always required for elongation.
- NusG, in contrast, is often considered an anti-pausing factor that increases the overall rate and processivity of transcription. It also plays a vital dual role by physically coupling the ribosome to the RNAP complex, directly linking transcription and translation. Furthermore, it is a key component in Rho-dependent termination, where it helps the Rho factor interact with the paused polymerase.
The Mechanism of Nucleotide Addition
The process of adding new nucleotides to the growing RNA chain is catalysed and monitored by the RNA polymerase enzyme itself. The synthesis proceeds in the 5’ to 3’ direction, meaning new nucleotides are always added to the 3’ end of the chain. The core chemical reaction involves the 3’–OH group of the last nucleotide in the chain attacking an incoming nucleoside 5’–triphosphate. During this reaction, the incoming nucleotide loses its two terminal phosphate groups ( and ). The remaining phosphate group is then used to form a phosphodiester bond, linking the new nucleotide to the chain.
The molecule we are all familiar with as the energy currency of the cell, ATP, is also the same substrate for the nucleotide addition of adenine to RNA, with two of its phosphate groups being cleaved, providing the energy for the formation of the phosphodiester bond. (In DNA synthesis, it is dATP which is added to the growing chain.)
As RNA polymerase moves along the DNA template, it continuously unwinds the DNA at the front of the bubble and rewinds it at the back, maintaining the bubble’s structure and displacing the newly synthesised RNA strand.
Inside the enzyme, the DNA template makes a sharp 90° turn at the active site, which is facilitated by a “wall” of protein. The incoming nucleotides are thought to enter the active site through a secondary channel, sometimes called a pore. The rudder contacts the nascent RNA at the upstream edge of the RNA–DNA hybrid and helps stabilise the elongation complex; it does not melt the incoming DNA duplex.
This process presents a challenge: the polymerase must maintain tight contact with the nucleic acids but must also be able to break and remake these contacts with each cycle of nucleotide addition. As shown in Figure 8, the specific bases that occupy the contact points within the enzyme change every time the enzyme moves forward one position. This is achieved through conformational changes in flexible parts of the enzyme, which folds around the incoming nucleotide to facilitate catalysis and then unfolds to allow the enzyme to move to the next position.
Termination
Termination involves the recognition of the sequences that signal RNAP to halt additional nucleotide addition. This sequence is known as the terminator.
Termination can occur prematurely, for example due to long pauses in transcription, which can cause the transcription bubble to collapse and disrupt the RNA-DNA heteroduplex.
Bacterial RNA Polymerase consists of Multiple Subunits
The best genetically and biochemically characterised RNA polymerases are from bacteria, namely from E. coli, Thermus aquaticus, and Thermus thermophilus.
However, in all bacteria, a single type of RNAP is responsible for the synthesis of rRNA, mRNA, and tRNA (unlike in eukaryotes). Bacteria are small, and so each cell only has approximately 13,000 RNAP molecules at a time.
We call the complete RNAP enzyme the holoenzyme. In E. coli it has a weight of approximately 460kD, and is made up of 5 subunits (Figure 9):
- and subunits: the catalytic subunits, which are responsible for the addition of nucleotides to the RNA chain, and make up most of the holoenzyme by mass.
- Two subunits: these are responsible for enzyme assembly and promoter recognition. They form a dimer which serves as a scaffold for the core enzyme.
- One subunit: the function of this subunit is not yet fully understood, and is not even fully essential for the function of the enzyme. It is beyond the scope of this course.
- One factor: this subunit is responsible for the recognition of the promoter sequence.
RNAP can be split into two parts, the core enzyme which is composed of the subunits, and the holoenzyme, which is the complete enzyme including the sigma factor ().
The and subunits
The and subunits account for most of the mass of the core enzyme. Interestingly, their amino acid sequences and 3D structures are conserved with the largest subunits of the RNAPs from all other domains of life, bacteria, archaea, and eukaryotes. This indicates that the basic features of transcription are shared amongst all organisms on earth.
and form:
- the active centre of the enzyme,
- the main channel through which the DNA passes during the transcription cycle,
- the secondary channel through which the substrate ribonucleotides enter the enzyme (on their way to the active site)
- and the exit channel through which the RNA leaves the enzyme.
Mutations to the genes which encode these subunits, rpoB and rpoC have been shown to affect all stages of transcription.
The subunits
The dimer formed by the subunits serve as the scaffold for the assembly of the core enzyme. The C-terminus of the subunits also directly contact the promoter, and thereby contribute somewhat to promoter recognition.
factors
The core enzyme has a general affinity for DNA, primarily through electrostatic interactions between the protein (basic) and the DNA (acidic). When bound to DNA in this fashion, the DNA remains in duplex form. This brings us to a key point:
- The core enzyme can synthesise RNA on a DNA template, but alone it cannot recognise promoters.
The form of the enzyme responsible for actually initiating transcription from promoters is called the holoenzyme ().
The addition of the factor to the core enzyme is what allows the enzyme to both recognise the promoter sequence, and initiate transcription. The also helps to reduce non-specific binding of RNAP to the DNA, so much so that the association constant to random DNA is reduced by a factor of 1,000, but the association constant to the promoter sequence is increased by a factor of 1,000.
factors are an example of a transcription factor. Transcription factors are proteins which modulate the activity of RNAP for specific promoters.
RNAP-promoter binding rates
The rate at which the holoenzyme binds to different promoters varies widely, and thus this is an important parameter in determining promoter strength; that is, the efficiency of an individual promoter in initiating transcription.
How does RNAP find the promoter?
How does RNAP find promoters amongst the approximately 4.6 million base pairs of DNA in the E. coli genome? It was thought that RNAP would spend most of its time non-specifically interacting with random portions of DNA, in the hopes that it finds a promoter to bind to. Kinetic measurements of RNAP’s binding and unbinding rates have shown that this mechanism of finding a promoter is far too slow.
In actual fact, the process is probably sped up by the fact that RNAP can bind strongly anywhere on the genome, and not just to specific promoter regions. An alternative model has been shown to be much more likely:
- RNAP diffuses through the cell, and binds to the DNA non-specifically.
- The enzyme engages in a sliding motion randomly along the DNA.
- Because the DNA in bacteria is intricately and tightly packed into the nucleoid, the rate at which the enzyme can find another piece of DNA to bind to is very high. This is called intersegment transfer.
- This can repeat until the RNAP finds a promoter.
These proposed mechanisms may allow RNAP to search locally along DNA or move between nearby DNA segments.
Sigma Factors Control Binding to DNA by Recognising Promoter Sequences
Recall that promoters themselves are not transcribed, but are simply sequences of DNA which RNAP recognises and binds to. The sequence of the promoter itself defines its function. This is an example of a cis-acting site.
RNAP typically can span approximately 75bp of DNA when bound, however it has been found that across that 75bp there is very little consensus across promoters.
Only a few short sequences are significantly conserved across promoters. There are two 6-bp elements referred to as the:
- -10 element (sometimes known as the TATA box, or Pribnow box) which has the consensus sequence TATAAT.
- -35 element, which has consensus sequence TTGACA.
- A 17bp spacer between the -10 and -35 elements. There is no specific consensus sequence, but 17bp is the consensus length.
Their numbers refer to their positions relative to the transcription start site (they are upstream).
These two elements are usually the most important for promoter recognition since they have been shown to interact with the factor, but many others are important, and are shown in Figure 13
It is also important to note that mutations in the promoter sequence can have a significant effect on the strength of the promoter:
- Up mutations are mutations which increase the strength of the promoter, and thus result in increased transcription, and subsequent gene expression.
- Down mutations are mutations which decrease the strength of the promoter, and thus result in decreased transcription, and subsequent gene expression.
Mutations in a promoter which bring the sequence closer to the consensus are likely to increase the strength of the promoter, and vice versa. Additionally, mutations in the spacer which bring the length closer to 17bp are likely to increase the strength of the promoter, and vice versa. In other words, modifying the consensus sequence of the promoter can reduce the strength of the promoter.
Different sigma factors bind to different consensus sequences
E. coli has seven sigma factors, each of which causes RNA polymerase to initiate at a set of promoters defined by specific sequences. Sigma-70-family factors generally recognise –35 and –10 elements, whereas sigma 54 recognises distinct –24 and –12 elements. These sigma factors control the expression of sets of genes associated with different cellular functions under different environmental conditions. These specialised sigma factors bind the promoters of genes appropriate to the environmental conditions, increasing the transcription of those genes.
Below is a table of the different sigma factors and their associated promoters.
| Subunit (Gene) | Size (Number of Amino Acids) | Approximate Number of Promoters | Promoter Sequence Recognised |
|---|---|---|---|
| Sigma 70 (rpoD) | 613 | 1,000 | TTGACA—16 to 18 bp—TATAAT |
| Sigma 54 (rpoN) | 477 | 5 | TGGCACG—N4—TTGC (–24/–12 elements) |
| Sigma S (rpoS) | 330 | 100 | TTGACA—16 to 18 bp—TATAAT |
| Sigma 32 (rpoH) | 284 | 30 | CCCTTGAA—13 to 15 bp—CCCGATNT |
| Sigma F (rpoF) | 239 | 40 | CTAAA—15 bp—GCCGATAA |
| Sigma E (rpoE) | 202 | 20 | GAA—16 bp—YCTGA |
| Sigma FecI (fecI) | 173 | 1–2 | ? |
Two most important sigma factors to note are:
- Sigma 70: the “housekeeping” sigma factor or also called as primary sigma factor (Group 1), transcribes most genes in growing cells. Every cell has a “housekeeping” sigma factor that keeps essential genes and pathways operating.
- Sigma 38: the starvation/stress response sigma factor. When the cell is starved for nutrients, sigma 38 RNAP holoenzymes bind to their consensus sequence, and transcribe the genes which are required to respond to environmental stress.
Differences in the protein structure of the sigma factors are what allow them to bind to different promoter sequences (Figure 14).
Transcription Termination
Transcription termination typically occurs after a long pause in transcription, or at the site of a terminator sequence. Termination requires that all hydrogen bonds holding together the RNA-DNA heteroduplex be broken, allowing the DNA duplex to reform.
There are two types of terminators in prokaryotes, intrinsic terminators and terminators that require a rho factor protein. We will start with intrinsic terminators.
Intrinsic terminators
Intrinsic terminators are so named because they are not dependent on any other protein to stop transcription. Termination depends only on the RNA product, and are typically found in G+C rich regions of the transcript, which fold to form a hairpin structure (Figure 15).
You may recall previously that E. coli has over 4200 genes, so how can it be that 1,100 unique transcripts account for over half of them? The answer is simply that many genes are transcribed as a single transcript known as an operon, this will be covered later in the course.
These hairpins typically have a second feature: a series of up to seven uracil residues (thymine on the DNA nontemplate strand and adenine on the template strand). It has been found that in E. coli, approximately 1,100 sequences in the genome fit this pattern, meaning that approximately half of the terminators are intrinsic.
Rho-dependent terminators
The other type of terminator is the rho-dependent terminator. These terminators are dependent on the presence of a protein called the rho factor to stop transcription.
Rho factor first binds to a sequence within the transcript upstream of the termination site. This is known as a rho utilisation site (rut site for short). The rho factor then tracks along the RNA until it catches up to the RNAP. When RNAP reaches the termination site, rho freezes the structure of the polymerase, invades the exit channel, and destabilises the enzyme, which causes the transcript to be released (Figure 17).
The fact that rho needs to translocate from the rut site to the point of termination suggests that it needs to somehow catch up to RNAP during transcription.
Rho has a C-terminal ATPase domain, allowing it to translocate along the RNA. Under normal conditions, rho can translocate faster than RNAP. RNA-polymerase pauses nevertheless to provide favourable sites and time for efficient rho-dependent termination (Figure 17).
The Sigma Cycle
The process by which the sigma () factor associates with the core RNA polymerase (RNAP) to initiate transcription and is often released after promoter escape is known as the sigma cycle. Promoter escape weakens sigma–core interactions, but sigma release is not always required for elongation and some complexes retain sigma far downstream. The cycle can be broken down into several steps:
- Holoenzyme Formation: A free factor associates with a core enzyme () to form the RNAP holoenzyme. This complex is competent to bind specifically to promoter sequences.
- Promoter Binding and Initiation: The holoenzyme binds to a promoter, forming a closed and then an open complex. It then begins synthesising RNA, often going through several rounds of abortive initiation where short transcripts are made and released.
- Promoter Clearance and Possible Release: Once the transcript reaches a length of approximately 10-12 nucleotides, the polymerase undergoes promoter clearance. Sigma–core interactions weaken and sigma is often released, but some elongating complexes retain sigma.
- Elongation: RNA polymerase moves down the DNA template and continues to synthesise the RNA transcript.
- Recycling: When sigma is released, it becomes available to bind to another core enzyme and start a new cycle of transcription initiation at another promoter.
Detecting DNA-Protein Interactions with DNA Footprinting
One can learn about the specific binding and ability of proteins, in our case RNAP, to recognise DNA sequences using a technique known as DNA footprinting.
Say you have a stretch of DNA to which RNAP is bound. This DNA can be digested with an endonuclease, and the resulting fragments can be separated by gel electrophoresis. The key point to note is that DNA which is obscured by the protein will not be digested, and will therefore remain intact.
By comparing the pattern of the digested DNA to a control stretch of DNA, to which no protein was bound, one can learn about the specific sequence the protein was bound to.
In practice, the endonuclease will digest the DNA into fragments of varying length. By only doing a partial digestion, one will have a mixture of all possible fragment lengths.
By running these fragments through polyacrylamide gel electrophoresis, one can separate the fragments by size with single nucleotide resolution. Figure 19 shows 31 bands. In the protected fragments, bonds cannot be broken when they are obscured by the protein. The absence of bands 9 through 18 in the figure identifies a protein-binding site covering the region located 9 to 18 bases from the labelled end of the DNA.
One can sequence the original stretch of DNA and identify the region located 9 to 18 bases from the labelled end of the DNA, and thus know the specific sequence the protein was bound to.