4: RNA splicing and processing
Introduction, and the importance of RNA processing
RNA is a central player in gene expression. It was first characterised as the intermediate in the process of protein synthesis (i.e solely as an information carrier, like DNA), but since then, many RNAs that play structural and functional roles have been discovered (think tRNA, rRNA, etc).
Recall that an interrupted gene is one where the sequences retained in the mature RNA (exons) are separated by intervening sequences called introns. These introns are transcribed into RNA along with the exons but are then removed by splicing. The exons are joined in the mature RNA, and only their coding-sequence portions are translated into protein.
In Eukaryotes, RNAs transcribed from their respective genes require further processing to become mature and functional. Interrupted genes are found in all groups of eukaryotes:
- Most genes in multicellular eukaryotes are interrupted.
- A few genes in unicellular eukaryotes (such as yeast) are interrupted.
Genes vary widely according to the numbers and lengths of introns, but a typical mammalian gene has 7-8 exons spread out over 16kb. The exons are relatively short, only around 100-200bp, and the intron are long, approximately 1kb. Eukaryotes have introns, intervening sequences removed from precursor RNA, for several reasons, primarily related to gene expression regulation and the potential for generating protein diversity.
The discrepancy between the interrupted organisation of the gene on the DNA, and the uninterrupted organisation of its mature RNA (mRNA) requires a mechanism to remove the introns and join the exons together in a process called splicing. An pre-sliced mRNA is known as a pre-mRNA.
Splicing occurs in the nucleus (Figure 1), along with other modifications to the new RNA. In brief:
- The transcript is capped at the 5’ end
- has the introns removed
- is polyadenylated at the 3’ end (adding a tail of adenine nucleotides)
- is then transported through nuclear pores to the cytoplasm, where it is translated into protein.
The 5’ end of Eukaryotic mRNA is capped
Transcription starts with a nucleoside triphosphate, usually an A or a G. When a mature mRNA is examined, however, the 5’ end does not have the expected nucleoside triphosphate. Instead, it contains a 5’ cap, which is formed by adding a guanine (G) to the terminal base of the transcript via an unusual 5’–5’ triphosphate linkage.
The capping process happens during transcription. Shortly after initiation, RNA Polymerase II pauses about 30 nucleotides downstream. This pause allows capping enzymes to be recruited to add the cap to the 5’ end of the new RNA molecule. This capping is crucial because, without this protection, the nascent RNA is vulnerable to degradation by exonucleases. Thus, capping is an important checkpoint for the polymerase to continue elongation and transcribe the rest of the gene.
The addition of the terminal G is catalyzed by an enzyme called guanylyl-transferase (GT). The new G residue is added in the reverse orientation to all the other nucleotides. This structure is then a substrate for methylation. The most critical modification is the addition of a methyl group at the 7th position of the terminal guanine, creating a monomethylated cap, which is found on most mRNAs (Figure 2). Some small noncoding RNAs can be further methylated to form a trimethylated cap.
The 5’ cap has several important functions:
- It protects the mRNA from being degraded by 5’–3’ exonucleases.
- In the nucleus, it is recognized by the cap-binding complex (CBP20/80), which stimulates the splicing of the first intron and helps with mRNA export from the nucleus.
- Once in the cytoplasm, a different set of proteins (eIF4F) binds to the cap to initiate translation.
Nuclear Splice Sites are Short Sequences
By comparing the nucleotide sequence of a mature mRNA with that of the original gene, the junctions between exons and introns can be determined. These exon-intron boundaries contain short, well-conserved consensus sequences known as splice sites.
The splice sites define the ends of the intron directionally. They are named from left to right (5’ to 3’) along the intron:
- The 5’ splice site is at the 5’ end of the intron (sometimes called the left, or donor, site).
- The 3’ splice site is at the 3’ end of the intron (also known as the right, or acceptor, site).
As shown in Figure 3, the sequence of a generic intron is GU...AG. Because the intron starts with the dinucleotide GU and ends with the dinucleotide AG, the junctions are described as conforming to the GU-AG rule. In the coding strand of the DNA, this corresponds to a GT-AG sequence. The importance of these consensus sequences is confirmed by point mutations, which can prevent splicing from occurring.
While the vast majority of introns belong to the major class, a small fraction are minor introns. Introns are classified based on their consensus signals and the distinct splicing machinery that processes them (Figure 3):
- U2-type introns: The major class of introns that generally follow the GU-AG rule.
- U12-type introns: These were first identified through AU-AC termini, but most have GU-AG termini. They are defined by their consensus signals and use of the minor spliceosome.
Splice Sites Are Read in Pairs
A major challenge in pre-mRNA splicing is ensuring the correct 5’ and 3’ splice sites are joined together, especially in genes with many long introns. The sequences at the ends of an intron have no complementarity, so the pairing cannot rely on base-pairing.
Experiments have shown that the splicing machinery treats splice sites as generic and interchangeable. All 5’ splice sites are functionally equivalent, as are all 3’ splice sites. For example, if an exon from one gene (like the SV40 virus) is experimentally linked to an exon from a completely different gene (like mouse β-globin), the hybrid intron is correctly removed. However, splice sites are not functionally equivalent: their sequence strength and surrounding context determine how efficiently they are recognised. This tells us two things:
- The splicing machinery does not depend on specific secondary structures within a particular intron to identify its ends.
- The splicing apparatus is not tissue-specific and can process most RNA precursors, regardless of which cell they come from.
When several 5’ and 3’ splice sites are available, what rules prevent incorrect pairing? The process is guided by two main principles:
- Coupling with Transcription: Splicing often occurs as the pre-mRNA is still being transcribed. This temporal link imposes a direction, suggesting that splice sites are recognized in a 5’ to 3’ order, much like a “first-come, first-served” system.
- Sequence Context: A functional splice site is not just the core consensus sequence. It is defined by surrounding sequences in both exons and introns that act as splicing enhancers or splicing silencers (another example of cis-regulatory elements). For a splice site to be efficiently recognized, it must have a dominant set of enhancing elements over suppressing ones.
Together, these mechanisms ensure that splice sites are read in pairs in a linear order, correctly joining the exons of a gene.
Splicing Proceeds Through a Lariat Intermediate
The process of splicing requires three key sequences on the pre-mRNA: the 5’ splice site, the 3’ splice site, and a branch site located 18 to 40 nucleotides upstream of the 3’ splice site. The branch site contains a critical adenine (A) nucleotide that is essential for the chemical reaction. While this site is highly conserved in yeast (UACUAAC), it is less defined in multicellular eukaryotes.
Splicing occurs through two sequential transesterification reactions, where one phosphodiester bond is exchanged for another.
The mechanism is as follows (Figure 6):
- First Transesterification: The 2’-OH group of the adenine nucleotide at the branch site performs a nucleophilic attack on the 5’ splice site. This reaction cleaves the bond at the 5’ end of the intron and, in the same step, attaches this free 5’ end to the branch site adenine. This creates a looped, branched structure called a lariat. A key feature of the lariat is the unusual 2’-5’ phosphodiester bond that forms the branch point.
- Second Transesterification: The newly freed 3’-OH group of the 5’ exon now attacks the 3’ splice site. This reaction joins the two exons together (ligation) and simultaneously releases the intron, still in its lariat form.
After being released, the intron lariat is “debranched” by an enzyme, converting it back to a linear RNA molecule, which is then rapidly degraded in the nucleus.
The main role of the branch site is to identify the nearest 3’ splice site to be used in the second reaction. If the authentic branch site is mutated, especially in multicellular eukaryotes, the splicing machinery may use nearby, related sequences called cryptic sites. This demonstrates that proximity to the 3’ splice site is a critical factor for branch site function.
snRNAs are required for Splicing
The splicing apparatus is a complex machinery that assembles on the pre-mRNA to ensure the precise removal of introns. This process involves both proteins and specialized RNA molecules.
In eukaryotic cells, there are many types of discrete, small RNA molecules. Those confined to the nucleus are known as small nuclear RNAs (snRNAs), while those in the cytoplasm are called small cytoplasmic RNAs (scRNAs). A specific class found in the nucleolus, the small nucleolar RNAs (snoRNAs), is involved in processing ribosomal RNA. In their natural state, these RNAs exist as ribonucleoprotein particles, referred to as snRNPs and scRNPs (sometimes colloquially called “snurps” and “scyrps”).
The splicing reaction occurs within a large complex called the spliceosome. This massive structure, which is larger than a ribosome, is formed by the sequential assembly of snRNPs and other proteins onto the pre-mRNA. Its primary role is to bring the 5’ and 3’ splice sites together before any cleavage occurs.
The Spliceosome and its Components
The spliceosome is a 50S-60S ribonucleoprotein particle with a mass of about 12 MDa. Its composition is a mix of RNA and protein, with five snRNAs and their associated proteins making up almost half of its mass.
The key components are:
- Five snRNPs: The spliceosome contains five essential snRNPs, named after the snRNA they contain: U1, U2, U5, U4, and U6. The U4 and U6 snRNPs typically exist together as a single U4/U6 di-snRNP particle. The importance of these snRNAs is demonstrated by experiments where inactivating any one of them prevents splicing.
- Associated Proteins: Each snRNP contains a single snRNA and several proteins.
- A core group of eight conserved proteins, known as Sm proteins, is found in the U1, U2, U4, and U5 snRNPs. These proteins are recognized by anti-Sm antibodies, which are often found in patients with autoimmune diseases like systemic lupus erythematosus.
- The U6 snRNP lacks Sm proteins but contains a set of Sm-like (Lsm) proteins.
- Splicing Factors: In addition to the snRNP proteins, about 70 other proteins are considered splicing factors. These are required for the assembly of the spliceosome, binding to the pre-mRNA, and creating the catalytic center for the transesterification reactions.
Like a ribosome, the function of the spliceosome relies heavily on RNA-RNA interactions, both between the snRNAs and the pre-mRNA transcript and among the snRNAs themselves. This suggests that the RNA components play a direct, possibly catalytic, role in the splicing reaction.
Committing Pre-mRNA to the Splicing Pathway
For splicing to begin, the machinery must first recognize and commit to the correct splice sites on the pre-mRNA transcript. This initial stage involves a series of binding events that form a stable foundation for the rest of the spliceosome to assemble.
Initial Splice Site Recognition and the E Complex
The very first step in splicing is the binding of the U1 snRNP to the 5’ splice site. The U1 snRNA has a single-stranded 5’ end that is complementary to the consensus sequence at the 5’ splice site, allowing it to bind through a direct RNA-RNA base pairing interaction.
The necessity of this base pairing has been confirmed experimentally. Mutations that disrupt the pairing between the 5’ splice site and U1 snRNA abolish splicing, but a second, “suppressor” mutation in the U1 snRNA that restores this complementarity can rescue the splicing process.
This initial binding of U1 is part of a larger assembly called the commitment complex (also known as the E complex in mammals). This stable complex forms when:
- U1 snRNP binds to the 5’ splice site.
- The protein U2AF binds to the polypyrimidine tract located between the branch site and the 3’ splice site.
- The branch point binding protein (BBP/SF1) interacts with the branch site sequence.
These interactions are cooperative and lock the pre-mRNA into the splicing pathway.
The Role of SR Proteins
In multicellular eukaryotes, where splice site consensus sequences can be weak or divergent, an additional class of factors is essential for forming the commitment complex: SR proteins.
These proteins are characterised by two key domains:
- An N-terminal RNA-recognition motif for binding to specific RNA sequences.
- A C-terminal RS domain (rich in Arginine/Serine dipeptides) that mediates protein-protein interactions.
SR proteins act as a molecular “glue,” stabilizing the E complex by forming a network of interactions. They can bind to the U1 snRNP at the 5’ site and to U2AF at the 3’ site, bridging the two ends of the intron and ensuring the splice sites are properly recognised.
Interestingly, SR proteins are not found in organisms like yeast, where splice sites are highly conserved and don’t require this extra layer of recognition.
Intron vs. Exon Definition
The way the splicing machinery initially identifies a pair of splice sites can occur via two different strategies, largely dependent on the size of the introns and exons.
Intron Definition: In this model, the machinery recognizes the 5’ and 3’ splice sites across the intron simultaneously. This works well for organisms like yeast, where genes typically contain a single, short intron. It is the conceptually simpler method.
Exon Definition: This is the predominant method in multicellular eukaryotes, where introns are often extremely long and exons are relatively short (100-300 nucleotides). Instead of trying to span a massive, variable intron, the machinery identifies an exon first. U2AF binds to the 3’ splice site at the beginning of the exon, and U1 snRNP binds to the 5’ splice site at the end of the same exon. SR proteins often help define the exon by binding to specific sequences within it. The machinery then “rearranges” these interactions to pair the 3’ splice site with the upstream 5’ splice site, effectively defining the correct intron-exon boundaries and ensuring adjacent exons are joined correctly.
Alternative Splicing is a Rule, Not an Exception
When we think of a gene, we often imagine a single pre-mRNA transcript being processed into one specific mature mRNA. However, for the majority of genes in multicellular eukaryotes, this is not the case. Instead, a process called alternative splicing is incredibly common. It’s estimated that over 90% of mammalian genes undergo alternative splicing. This isn’t a mistake by the cellular machinery; it’s a regulated and essential part of the gene expression program that allows a single gene to produce multiple distinct mRNA sequences, and thus, multiple protein products.
Modes of Alternative Splicing
Alternative splicing can occur through several different mechanisms, which can be used individually or in combination for a single pre-mRNA transcript (Figure 12). The main modes include:
- Constitutive splicing: This is the process of removing introns and joining exons of a pre-mRNA in a consistent, non-variable way, producing a single type of mature mRNA. This is the most common form of splicing and is not considered a type of alternative splicing.
- Exon inclusion or skipping (cassette exons): An entire exon is either included or excluded from the final mRNA.
- Intron retention: An entire intron is kept in the mature mRNA.
- Mutually exclusive exons: One of two or more exons is included, but not both. This is often regulated in a tissue-specific manner.
- Alternative 5’ splice-site selection: An alternative 5’ splice site is used, changing the 3’ end of the upstream exon.
- Alternative 3’ splice-site selection: An alternative 3’ splice site is used, changing the 5’ end of the downstream exon.
Functional Consequences of Alternative Splicing
Alternative splicing dramatically expands the coding capacity of the genome and has profound effects on cellular function in two primary ways:
- Creating Structural and Functional Diversity of Proteins: By including or excluding specific exons, alternative splicing can add or remove protein domains. This can change a protein’s function, its location within the cell, or its interaction with other molecules. A classic example is the CaMKIIδ gene (Figure 13).
- Skipping all three alternative exons produces a cytoplasmic kinase.
- Including exon 14, which contains a nuclear localization signal, sends the kinase to the nucleus.
- Including exons 15 and 16 targets the kinase to the cell membrane, which is common in neurons.
In other cases, alternatively spliced products can have opposite functions. For instance, many genes involved in apoptosis (programmed cell death) produce one isoform that promotes cell death and another that protects the cell. The balance between these isoforms can determine the cell’s fate.
For example: Bcl-x, a key member of the Bcl-2 family of apoptosis regulators. The BCL2L1 gene, which codes for Bcl-x, undergoes alternative splicing to produce two major isoforms with opposing functions:
- Bcl-xL (Long): This isoform is anti-apoptotic (pro-survival). It is the larger protein and works by preventing the release of mitochondrial factors like cytochrome c, which would otherwise trigger cell death.
- Bcl-xS (Short): This isoform is pro-apoptotic (pro-death). It is a smaller protein that lacks a key domain found in Bcl-xL. It actively promotes apoptosis, in part by inhibiting the protective effect of Bcl-xL.
Therefore, the ratio of Bcl-xL to Bcl-xS in a cell is a critical switch that determines its susceptibility to apoptotic signals. A high Bcl-xL/Bcl-xS ratio favors survival, while a low ratio sensitizes the cell to death.
- Regulating Gene Expression: Alternative splicing can also control the abundance of an mRNA. Sometimes, splicing includes an exon that contains a premature stop codon. This targets the mRNA for rapid degradation, effectively reducing the amount of protein produced. This mechanism is a way to fine-tune gene expression levels, similar to transcriptional regulation.