What this is
- This review discusses advanced methods for mapping RNA modifications in mammalian cells.
- Chemical modifications, such as and pseudouridination, play critical roles in gene expression.
- The authors summarize recent developments in sequencing technologies that allow for quantitative and analysis of these modifications.
Essence
- New sequencing methods enable detailed mapping of RNA modifications at , enhancing understanding of their biological roles.
Key takeaways
- mA-SAC-seq allows mapping of internal mA methylomes at , revealing ∼30,000–130,000 mA sites in RNA with high reproducibility.
- BID-seq quantitatively maps (Ψ) modifications, uncovering thousands of Ψ sites with stoichiometric information, crucial for understanding their function.
- UBS-seq effectively quantifies 5-methylcytosine (mC) modifications in structured RNAs, identifying 2,723 mC sites in HeLa cells with a modification fraction ≥5%.
Caveats
- Current methods may require optimization for sensitivity and accuracy, particularly for low-abundance RNA species.
- Sequence context preferences of certain enzymes can affect the accuracy of modification stoichiometry measurements.
Definitions
- methylation: A chemical modification involving the addition of a methyl group to RNA, affecting its stability and function.
- pseudouridine (Ψ): A modified nucleoside in RNA that enhances stability and translation efficiency.
- base resolution: The ability to identify modifications at the level of individual nucleotides in RNA sequences.
Simplified
Key References
Introduction
Chemical modifications exist in almost all RNA species of mammalian cells, playing critical roles in regulating gene expression. These RNA modifications mainly include base methylation (e.g., N6-methyladenosine (m6A), N1-methyladenosine (m1A), N7-methylguanosine (m7G), 5-methylcytidine (m5C), N3-methylcytidine (m3C), N1-methylguanosine (m1G), N2,N2-dimethylguanosine (m22G)), backbone methylation (2′-O-methylation (Nm)), base acetylation (e.g., N4-acetylcytidine (ac4C)), base isomerization (e.g., pseudouridine (Ψ)), base editing (e.g., inosine (I)), and base oxidation/reduction (e.g., dihydrouridine (D), 8-oxoguanosine (o8G); Figure). Abundant noncoding RNA species such as rRNA and tRNA are known to possess dense chemical modifications5 that impact their biogenesis, local structure, and stability and modulate the subsequent translation. High modification stoichiometry (modification fraction) is observed at most modified sites in mammalian rRNA and tRNA. The high abundances of rRNA and tRNA allow functional study of their RNA modifications using diverse approaches including quantification of the modification site and stoichiometry with mass spectrometry.6
Besides rRNA and tRNA, RNA modifications are also present in low-abundance RNA species, such as mRNA, regulatory noncoding RNA, and chromatin-associated RNA (caRNA), which could not be investigated at each modified site through conventional biochemical approaches. The next-generation sequencing technology offers an opportunity to map the global distributions of these mRNA modifications.7,8 Using mRNA purified from human cells as an example, mass spectrometry has revealed the presence of different mRNA modifications, including m6A, Ψ, 2′-O-methylation, m5C, m1A, internal m7G, ac4C, etc. To unveil the biological functions of these mRNA modifications, an ideal sequencing method would not only uncover the location of individual RNA modification transcriptome-wide but also the stoichiometric information at each modified site. In the past several years, our laboratory has been working on developing such quantitative base-resolution methods that allow the evaluation of modification stoichiometry at individual modified sites for almost all major mRNA modifications. Here, we summarize m6A-SAC-seq1 and eTAM-seq9 for m6A; BID-seq for Ψ;2 UBS-seq for m5C;3 DAMM-seq4 for one-pot sequencing of m1A, m3C, m1G, and m22G; m1A-quant-seq10 for m1A; Nm-Mut-seq11 for Nm; and m7G-seq and m7G-quant-seq for internal m7G12,13 in this Account.
Major RNA modifications in mammalian cells.
-Methyladenosine (mA) N 6 6
Among all internal mRNA modifications in mammals, m6A is the most abundant, exhibiting ∼0.4–0.6% m6A/A abundance in mRNA.14,15 m6A affects almost every aspect of mRNA processing and metabolism, impacting a variety of biological processes in a living cell.5,7 m6A methylation on mammalian mRNA is mediated by the METTL3/METTL14 methyltransferase complex as the main m6A “writer” protein.16 FTO and ALKBH5 were identified by our laboratory as the m6A “eraser” proteins in 2010 and 2013,17,18 respectively, which reverse m6A methylation. Proteins that preferentially bind m6A, or “reader” proteins, were also discovered, with YTHDF1–319 −21 and YTHDC1–222 −24 possessing an evolutionally conserved YTH domain that binds selectively to m6A. These proteins bind m6A-methylated RNA and regulate mRNA stability, translation, pre-mRNA splicing, nuclear export, and other biological processes. The discovery of m6A “writer,” “eraser,” and “reader” proteins largely promote the recent explosion of m6A epitranscriptome research; the application of next-generation sequencing technology to map mRNA m6A distribution has been an indispensable part in nearly every study related to m6A biology.
For over a decade, the antibody-based m6A-seq or m6A-MeRIP-seq strategy14,15 has been a prevalent method to study the distribution of m6A methylations at 100–200 nucleotide (nt) resolution. Relying on the anti-m6A antibody, miCLIP25 and m6A-LAIC-seq26 provided updated versions. However, these antibody-based methods were unable to provide single-base resolution, m6A stoichiometry, nor the sensitivity to compare m6A methylation dynamics. MAZTER-seq27 and m6A-REF-seq,28 exploring MazF RNase, selectively cleave RNA at ACA motif without m6A methylation. Although this approach could detect m6A methylation within the ACA motif, it only accounts for around 15% of the overall m6A sites in mRNA and lacks sensitivity. DART-seq,29 m6A-SEAL,30 and m6A-label-seq31 have also been reported, but they either lack stoichiometric information at m6A-modified sites or cannot be applied transcriptome-wide.
We have recently reported m6A-selective allyl chemical labeling and sequencing (m6A-SAC-seq) that reads out m6A as mutated signals.1 Our approach is based on the unique activity of the Methanocaldococcus jannaschii homologue MjDim1, which is known to convert A to m6A and then m6A to m62A.32 We found that this enzyme can employ a chemically modified allylic-SAM to install an allyl group at the N6 position of m6A. With the allylic-SAM as the cofactor, MjDim1 exhibited a ∼10-fold preference for m6A over A in the allyl group transfer reaction, converting m6A into allyl-modified m6A (N6-allyl,N6-methyladenosine, a6m6A; Figurea).1 To further induce misincorporation signatures at the generated a6m6A sites, the subsequent I2 treatment converts a6m6A and a6A into the corresponding N1,N6-ethanoadenine and N1,N6-propanoadenine derivatives, respectively, which can be read out as misincorporation signatures by human immunodeficiency virus 1 (HIV-1) reverse transcriptase (HIV RT; Figurea). HIV RT generated ∼10-fold higher mutation rates at the cyclized a6m6A sites (m6A sites) than the cyclized a6A sites (unmodified A sites) in almost all sequence contexts in NNm6 ANN versus NNANN. Meanwhile, HIV RT could also read through cyclized NNa6m6 ANN without noticeable RT stops. Therefore, the MjDim1-catalyzed allyl transfer using allylic-SAM shows a ∼10-fold preference for m6A over A, and the cyclized a6m6A adduct (generated from m6A) induces another ∼10-fold higher misincorporation rate than the cyclized a6A adduct formed from unmodified A. Overall, m6A-SAC-seq offers a ∼100-fold higher selectivity at m6A over unmodified A sites1. Before MjDim1 labeling, RNA fragments could also be split into two sections, “untreated” and “FTO treated”, in which the FTO treatment erases a large portion of mRNA m6A as a background control to further eliminate false positives.
m6A-SAC-seq maps internal m6A methylomes at base resolution, and the misincorporation rates readout at each modified site could be used for assessing the m6A modification stoichiometry. By applying m6A-SAC-seq to the mixtures of synthetic NNm6ANN and NNANN oligo probes in different ratios, we established a spike-in calibration system for each sequencing library and monitored m6A modification fractions transcriptome-wide through computation. The initial m6A-SAC-seq protocol starts with ∼30 ng of poly(A) RNA or rRNA-depleted RNA. After further optimization of library preparation and bioinformatic pipelines, the latest m6A-SAC-seq starts with ∼2 ng of poly(A) RNA with high reproducibility,33 revealing ∼30 000–130 000 m6A sites that overlap well with m6A profiles obtained by antibody-based approaches. For the first time, m6A-SAC-seq demonstrated the quantitative base-resolution maps of internal m6A modifications at diverse motif contexts unbiasedly, and set the technological basis for sensitively monitoring m6A dynamics in diverse biological processes.
As a complement sequencing tool to m6A-SAC-seq, which involves enzymic reactions targeting m6A bases over unmodified adenosines, through collaboration we have also helped develop evolved TadA-assisted N6-methyladenosine sequencing (eTAM-seq).9 eTAM-seq employs highly efficient global adenosine deamination of unmethylated A sites mediated by TadA8.20 and reads out these unmodified sites as guanosines (G) in next-generation sequencing (NGS) data, while m6A is resistant to this enzymic deamination and thus is still read as A (Figureb). In this way, eTAM-seq achieves transcriptome-wide, base-resolution detection and quantification of m6A. It uncovered ∼35 000 m6A-modified sites in HeLa mRNA,9 which overlaps very well with m6A-SAC-seq results. Note that eTAM-seq enables site-specific and quantitative sequencing of m6A with as few as 10 cells,9 which requires much lower input RNA than other existing quantitative sequencing methods. Similar to the deamination principle in eTAM-seq but employing chemical reactions instead, Liu et al. developed glyoxal and nitrite-mediated deamination of unmethylated adenosines (GLORI) to quantitatively map m6A at base precision.34 GLORI maps m6A methylomes transcriptome-wide in mouse and human cells, revealing clustered m6A dynamics with information on site distribution and modification stoichiometry. The current version requires several hundred nanograms of RNA as the starting material, but further improvements should be able to lower the input amount substantially.
m6A-SAC-seq has two advantages over deamination-based methods such as eTAM-seq9 and GLORI:34 (i) only the m6A sites show mutation signatures, which preserves the sequence complexity and allows more accurate sequence mapping; (ii) the positive readout of m6A significantly reduces sequencing costs. A main limitation is the sequence context preference of the MjDim1 enzyme. Calibration probes are highly recommended for each sequencing library to reveal modification stoichiometry at different sequence contexts, since MjDim1 does exhibit sequence preference toward the GA motif. Note that while m6A-SAC-seq requires much less sequencing depth compared with eTAM-seq and GLORI, sequence bias needs to be calibrated. Taken together, m6A-SAC-seq, eTAM-seq, and GLORI currently serve as the three quantitative mapping tools targeting m6A (Table S1↗).
Base-resolution quantitative mA-SAC-seq and eTAM-seq for mapping mA modification in mammalian mRNA. (a) Quantitative mA-SAC-seq maps internal mA sites as misincorporation signatures. (b) Quantitative eTAM-seq induces A to G conversion at unmethylated A sites instead of mA sites. 6 6 6 6 6
Pseudouridine (Ψ)
Following m6A, pseudouridine (Ψ) is the second most abundant mRNA modification in mammalian mRNA, with ∼0.2% Ψ/U levels measured by mass spectrometry. Ψ modifications are known to distribute broadly in abundant noncoding RNAs, such as rRNA, tRNA, and small nuclear RNA (snRNA). Ψ is also known to exist in mammalian mRNA; however, sequencing Ψ within low-abundance mRNA was challenging using traditional experimental methods. Based on the chemical reaction with N-cyclohexyl-N′-(2-morpholinoethyl) carbodiimide methyl-p-toluenesulfonate (CMC) to yield CMC-modified Ψ and subsequent RT truncation signatures induced during reverse transcription,35 −37 several next-generation sequencing methods were developed to map the transcriptome-wide distribution of cellular Ψ modifications, including Pseudo-seq,35 Ψ-seq,36 and PSI-seq.37 The RT truncation signatures are difficult to detect, especially at low- and medium-modified sites. Only modest numbers of Ψ sites were identified previously in mammalian mRNA with a low overlap among different data sets. An azide-modified CMC was applied to enrich Ψ-containing mRNA fragments in CeU-seq,38 with many more Ψ sites detected, but this method could not reveal Ψ stoichiometry transcriptome-wide. HydraPsi-seq,39 a method that relies on the resistance of pseudouridine toward hydrazine/aniline cleavage, was also reported to map Ψ modifications in yeast mRNA, but it cannot reveal Ψ stoichiometry either. The lack of a reliable method to map Ψ transcriptome-wide has been a bottleneck in functional investigations of Ψ in mRNA and other low-abundance RNA species.
We reported bisulfite-induced deletion sequencing (BID-seq)2 to quantitatively map Ψ at single-base resolution in 2022. Inspired by Khoddami et al., who reported RBS-seq (a modification of RNA bisulfite sequencing),40 a modified version of RNA bisulfite sequencing that enables the simultaneous detection of m5C, Ψ, and m1A at single-base resolution transcriptome-wide. A key discovery in RBS-seq was the observation of Ψ-dependent deletion signatures generated by a Ψ-bisulfite adduct during RT.40,41 Fleming et al. further investigated the chemical products41 generated at Ψ-modified sites after the bisulfite reaction in RBS-seq and identified two main adducts that were shown to induce the opening of the ribose ring and cause deletion signatures during reverse transcription. N1-methylpseudouridine (m1Ψ) in mRNA vaccines was shown to react similarly with bisulfite to yield ribose-opening products as well.42
RBS-seq still only uncovered very limited numbers of Ψ sites with weak signatures, mainly because of the conversion ratio of Ψ into Ψ-bisulfite adducts under conventional bisulfite reaction conditions. In the conventional bisulfite reaction, the protonation at the N3 position of cytosine typically requires an acidic pH, which facilitates the attack of BS to the C6 position to generate the C-BS adduct and subsequent deamination. We reasoned that the acidic condition is not optimal for the formation of Ψ-bisulfite adducts critical to the deletion signature during RT. We hypothesized that a neutral pH could not only enhance the production of Ψ-BS adducts but also inhibit C-to-U conversion to avoid reduced sequence complexity, achieving high Ψ detection sensitivity and accuracy2 (Figurea).
Indeed, when testing the synthetic RNA oligo probes of AGΨGA versus AGUGA and AGCGA, the matrix-assisted laser desorption/ionization-time-of-flight (MALDI-TOF) MS measurement clearly showed that bisulfite offered an almost quantitative conversion of Ψ to Ψ-BS adducts under neutral conditions,2 but the unmodified uridines were not affected after bisulfite treatment and desulphonation. No detectable C-to-U conversion was observed. Furthermore, we systematically screened all commercially available RT enzymes and self-made evolved RTs and observed that SuperScript IV reads out Ψ-BS adducts as deletion signatures with the highest deletion ratios among all known RTs.2 We then constructed BID-seq libraries using the NGS platform (Figureb). With the fully modified Ψ site within a synthetic oligo containing NNΨNN, 232 out of 256 motifs in NNΨNN gave deletion ratios over 50% at the Ψ sites after bisulfite treatment, with nearly all motifs showing >25% deletion ratios. We further validated 42, 53, and two known Ψ sites in HeLa 18S, 28S, and 5.8 rRNAs, respectively, without any false positives observed. BID-seq was applied to mRNA samples from human cell lines and mouse tissues, uncovering thousands of Ψ sites with stoichiometric information at each modified site.2
Thirteen pseudouridine synthase (PUS) enzymes are encoded in the human genome, with several specific PUS enzymes reported to install Ψ modifications in human mRNA.43 The quantitative feature of BID-seq enabled us to monitor Ψ stoichiometry change in PUS-depleted cells versus the control, and we found that mRNA Ψ sites could be either installed by a single specific PUS enzyme or by multiple PUS proteins.2 BID-seq revealed more than 100 Ψ-modified stop codons in 12 mouse tissues and confirmed the role of Ψ in promoting stop codon readthrough in vivo.2 Collectively, BID-seq set the stage for functional and mechanistic investigation of Ψ in low-abundance mRNA. In 2023, a similar approach, PRAISE, was also reported.44 In a further optimized BID-seq protocol, we achieved even lower background deletions at unmodified uridines after bisulfite treatment and uncovered around 8500 Ψ sites starting from as little as 10 ng poly(A) RNA from mouse embryonic stem cells (mESC;45Table S2↗).
Base-resolution quantitative BID-seq for mapping pseudouridines in mammalian mRNA. (a) BID-seq bisulfite selectively reacts with Ψ sites but does not affect other RNA modifications or unmodified bases. (b) The brief BID-seq pipeline built on the NGS platform induces deletion signatures at internal Ψ sites.
5-Methylcytosine (mC) 5
m5C exists in diverse RNA species including rRNA, tRNA, mRNA, and various noncoding RNAs. The antibody-based mapping method for RNA m5C, such as m5C-RIP-seq46 and 5-azacytidine-mediated RNA immunoprecipitation (Aza-IP),47 could provide neither single-base resolution nor m5C stoichiometry information, while miCLIP48 requires overexpression of the mutant enzyme. In recent years, BS-seq has been increasingly utilized to examine m5C modifications, with several commercial RNA BS conversion kits available, including the EZ RNA Methylation Kit from Zymo Research and the Methylamp RNA BS Conversion Kit from Epigentek. Using BS-seq, recent studies49 −52 have revealed that m5C modification in mRNA and its regulator proteins impacts diverse cellular functions and plays crucial roles in development and cancer. However, the exact level and stoichiometry of m5C on mRNA have been a subject of debate due to the absence of a sensitive, robust, and quantitative sequencing method. Discrepancies were observed when conventional BS-seq was applied to low-abundance mRNA, with some studies detecting thousands of m5C sites in mRNAs53 while other studies discovered only a few sites.54 More recent studies have reported only a few hundred m5C sites in human and mouse transcriptomes using an improved bisulfite sequencing method and a more stringent computational approach.50,51 These inconsistent findings have raised the need to develop more sensitive and robust methods for identifying and quantifying real m5C sites in mRNA.
A major challenge for RNA m5C BS sequencing has been the high false positive rates or high background caused by incomplete C-to-U conversion due to reduced reaction temperature and reduced reaction time to avoid severe RNA degradation. The reduced temperature is also ineffective in denaturing local secondary structures of highly structured RNAs, leading to further reduced C-to-U conversion at structured regions. Mechanistically, two competing pathways exist in bisulfite conversion of RNA, with one giving the desired C-to-U conversion and the other leading to the undesired RNA degradation. The protonated N3 nitrogen under acid conditions facilitates cytosine’s reaction with BS to give the C-BS adduct, which is converted to U-BS adduct by deamination. Subsequent desulphonation of the U-BS adduct under basic conditions generates U, completing C-to-U conversion. Alternatively, the U-BS adduct may undergo spontaneous depyrimidination to cause DNA degradation. To improve C-to-U conversion efficiency and reduce RNA damage, we developed ultrafast BS sequencing (UBS-seq)3 using a new recipe with a high BS concentration (∼10 M) and a high reaction temperature (98 °C; Figure). Applying UBS-seq to highly structured rRNA as a model, we showed that the average detected fraction for the two known m5C sites was >95%, while no false positive site was detected when a 5% cutoff for the unconverted rate of C was used. This new approach allowed detection of m5C stoichiometry in highly structured RNA species and outperformed all the reported BS conditions in terms of lower background and higher sensitivity in detecting real m5C sites in structured rRNA.3
When UBS-seq was applied to polyA+-enriched RNA from HeLa and HEK293T cell lines, input mRNA as low as 10–20 ng could yield transcriptome-wide m5C quantification, identifying 2723 and 2404 m5C sites with a modification fraction ≥5%, respectively.3 The quantitative nature of UBS-seq allowed us to reveal sequence motifs of m5C sites in mRNA and assign NSUN2 as the main m5C methyltransferase that installs ∼90% m5C sites to HeLa mRNA. To further validate the detected m5C sites with low modification fractions, with NSUN2 and NSUN6 as potential mRNA m5C “writer” proteins, we conducted rescue experiments by transfecting the corresponding methyltransferase plasmids back to the HeLa cells with NSUN2 or NSUN6 depletion and sequenced the isolated poly(A)-tailed RNA. Indeed, we observed that the decreased m5C fractions in the depleted strains were mostly rescued, further confirming that these lowly modified m5C sites are real. In addition, our results showed that m5C sites deposited by NSUN2 but not by NSUN6 are enriched in 5′-UTR regions in both HeLa and HEK293T mRNA, suggesting that m5C modification or its binding proteins may be involved in regulating mRNA translation.3 This new method and the data sets will aid future functional investigations on RNA m5C (Table S3↗).
Chemistry principle of UBS-seq, a base-resolution approach for mC quantification in mammalian mRNA. 5
DAMM-seq to Map Base Methylations at Watson–Crick Base Pairing Interface
We and others previously detected N1-methyladenosine (m1A) at ∼0.02% m1A/A abundance in mammalian poly(A)-tailed RNA.55,56 Different from m6A and Ψ, which cannot naturally induce misincorporation signatures during RT, the methyl group at the N1 position of the m1A base can disrupt base pairing and induce misincorporation in the presence of many commercially available RT enzymes, such as HIV RT, AMV RT, SuperScript II RT, and SuperScript IV RT. The mutation signals tend to be weak with notable RT stops observed. The TGIRT-based m1A-MAP displays excellent performance in m1A site detection in tRNA, but the low turnover of the TGIRT enzyme impedes the actual application of m1A-MAP for longer RNAs.57 We and our collaborators set up an evolution platform to evolve engineered RT enzymes for efficient readthrough of m1A with high rates of mutation signatures. We developed a fluorescence-based RT evolution platform for the direct evolution of RT enzymes, to select enzymes that give high RT misincorporation ratios of any given type of RNA modification.10 We started with HIV RT and identified RT1306, an evolved HIV RT to give robust readthrough and high misincorporation rates at m1A sites. Taking advantage of the evolved RT1306 and the AlkB-mediated demethylation at m1A sites as controls, we developed m1A-quant-seq10 (Figurea, which uncovered several hundred m1A sites in human polyA-tailed RNA, with stoichiometric information (Table S3↗). m1A-quant-seq also uncovered an array of mRNAs and lncRNAs containing highly modified m1A sites (above 30% modification stoichiometry),10 such as MALAT1, mt-ND5, PRUNE, etc., which is consistent with the previous publications and suggested functional roles.58
During our studies of m1A in mammalian cytosolic tRNA and mitochondrial tRNAs, we noticed m1A methylation on nascent mitochondrial RNA (mt-RNA). Mitochondrial transcription is unique in human cells, with bidirectional transcription producing a long precursor mt-RNA that contains two mt-rRNA, 13 mt-mRNA, and 22 mt-tRNA serving as junctions within the mitochondrial polycistronic RNA. In addition to m1A methylations, other base methylations appear to also play roles in nascent mt-RNA processing. We therefore developed demethylation-assisted multiple methylation sequencing (DAMM-seq),4 which starts with ∼10 ng of input RNA and quantitatively maps not only m1A but also N3-methylcytidine (m3C), N1-methylguanosine (m1G), and N2,N2-dimethylguanosine (m22G) in mitochondrial polycistronic RNA. These modifications all block the Watson–Crick base pairing interface and exist in limited sequence motifs in tRNAs; HIV RT already reads through these known motifs well and induces high misincorporation rates for quantitative stoichiometry determination.4
DAMM-seq utilizes the misincorporation ratios obtained at each methylated site for estimating the methylation stoichiometry, enabling the quantitative characterization of m1A, m1G, m3C, and m22G in nascent mt-RNAs.4 As an application example of quantitative DAMM-seq, we applied DAMM-seq to sequence mitochondrial polycistronic RNA under different cellular treatments, particularly ALKBH7 depletion and overexpression. DAMM-seq sensitively monitored the methylation level changes at methylated m1A, m1G, m3C, and m22G sites within mitochondrial polycistronic RNA and uncovered an ALKBH7 demethylation effect at m1A and m22G sites within pre-tRNA regions of mt-Leu1 and mt-Ile, respectively. This study identified ALKBH7 as an RNA demethylase regulating mitochondrial RNA processing and mitochondrial activity.4 DAMM-seq serves as an effective approach for the quantitative investigation of multiple RNA methylations within nascent RNA species of an input RNA amount as low as 10 ng (Table S3↗). A similar method named PANDORA-seq was also reported around the same time.59
Quantitative base-resolution sequencing technology for mA and 2′--methylation modifications in mammalian mRNA. (a) The brief mA-quant-seq pipeline built on NGS platform induces a major A → T misincorporation signature at internal mA sites. (b) The brief N-Mut-seq pipeline built on the NGS platform induces A → T or C → T or G → T misincorporation signatures at internal A, C, and Gsites. 1 1 1 O m m m m
2′--methylation (N) O m
Among the expanding list of functionally relevant post-transcriptional RNA modifications, methylation of the 2′-OH of an RNA ribose (2′-O-methylation, Nm) is both abundant and unique. 2′-O-methylation is one of the most common RNA modifications present in many cellular RNAs, such as tRNA, rRNA, and small nuclear RNA (snRNA). It also exists in mammalian mRNA and is about 10-fold less abundant compared to m6A. Unlike base modifications, which are specific for one of the four RNA bases, Nm methylation occurs at the 2′-OH position of each ribonucleoside. Nm modifications can modulate the secondary structure of rRNA and tRNA and affect mRNA stabilization and translation.60,61 While functional roles of Nm modifications on abundant tRNAs and rRNAs have been well documented, studies of their impacts on lower-abundance RNAs, such as mRNAs and long noncoding RNAs (lncRNAs), have been hampered by the lack of effective antibodies and robust sequencing methods that map Nm at base resolution.
Previous base-resolution approaches using NGS have been successful at identifying 2′-O-methylation in highly abundant RNA.62 −66 However, most of these methods cannot provide stoichiometric information at the modified sites. RiboMethSeq65 is the only quantitative method developed thus far, but very deep sequencing depth is required for confident Nm detection in low-abundance RNAs and at low-stoichiometry Nm sites. In addition, most current Nm mapping methods mainly rely on RT truncation signatures induced by chemical-assisted cleavage, which may include substantial false positives arising from RT stops at structured regions in RNA or unmodified RNA ends from fragmentation biases or RNA processing. Compared with RNA truncation signatures, modification-dependent misincorporations as readouts are generally more reliable in modification detection and quantification. However, since Nm modification occurs on the ribose and does not directly modulate the Watson–Crick base pairing interface during reverse transcription, all known reverse transcriptases (RT) are unable to induce misincorporations, which has precluded the use of mutation-based analysis methods for Nm mapping. Leveraging the fluorescence-based RT evolution platform to generate RTs that are mutagenic at specific RNA modification of interest,8 we speculated that evolved RTs might be selected to generate misincorporations opposite of Nm-modified bases.
Indeed, a new HIV RT variant, RT-41B4, was selected to yield notable misincorporation signatures at Nm-modified sites. In the presence of adjusted dATP/dNTP ratios during RT, we developed Nm-Mut-seq (an RT-mutation-based Nm mapping method)11 that maps Cm, Gm, and Am methylations at single-base resolution with stoichiometry information (Figureb). To validate Nm-Mut-seq, we mapped almost all known Cm, Gm, and Am sites on human rRNAs with high misincorporation rates and without observing any false positives at unmodified sites or other rRNA modifications.11 Applying Nm-Mut-seq to HepG2 cellular mRNA, we uncovered around 1000 Cm, Gm, and Am sites, a subset of which were validated by multiple orthogonal methods,11 further confirming the fidelity of the method. Fibrillarin (FBL) was shown to be a main Nm “writer” protein for HepG2 mRNA, and the methylation is guided by snoRNAs that interact with the target mRNAs. Note that the current Nm-Mut-seq still requires ∼200–800 ng of input polyA+ RNA (Table S3↗); otherwise, the high PCR cycle numbers lead to dramatic PCR duplicates in the obtained NGS data. Nm-Mut-seq, therefore, provides an effective method to not only detect Nm at base resolution in low abundant RNAs such as mRNA and lncRNA but also measure the stoichiometry of the modified sites transcriptome-wide.
Internal-Methylguanosine (mG) N 7 7
m7G methylation is a well-known modification in the cap structure of the 5′ end of mammalian mRNA, which impacts diverse biological processes such as mRNA stability, splicing, nuclear export, and translation. Highly modified internal m7G sites were also found at cytoplasmic tRNA and 18S rRNA.2 However, whether m7G methylation could be deposited internally to mammalian mRNA remained unknown for a long time. In 2019, we and another group demonstrated the existence of internal m7G modification in mammalian mRNA and miRNA.12,67 We reported the antibody-based m7G-MeRIP-seq and a single-base resolution m7G sequencing method m7G-seq12 (Table S3↗).
Using m7G-MeRIP-seq, we identified METTL1 as the main “writer” protein that installs mRNA internal m7G. m7G is positively charged and could be converted into a reduced state in the presence of sodium borohydride; the subsequent heating under acidic conditions leads to depurination at the reduced m7G site, yielding an abasic site (AP site).12 After screening different commercially available RT enzymes, we found that the abasic site generated at the m7G site could be read as RT misincorporation signals using HIV RT; the abasic sites could be further captured and enriched by biotin-tagged hydrazine to yield a higher misincorporation rate. Overall, m7G-seq includes three libraries for each sequencing library: zero mutations at internal m7G sites in the “Input” library, a moderate mutation rate induced at abasic sites in “Before pulldown” libraries, and a high mutation rate induced at biotin-enriched abasic sites in “Pulldown” libraries.12 The base-resolution m7G-seq confirmed the presence of mRNA internal m7G methylations in cancer cell lines. Note that a recent study revealed that quaking proteins (QKIs) bind to internal m7G-modified mRNAs with GA-rich motifs, as the first reported internal m7G “reader” proteins,68 suggesting functional roles of mRNA internal m7G methylations in mammals.
Besides the role of METTL1 as the 'writer' protein for mRNA internal m7G installation, METTL1 is well-known to mediate m7G methylation at position 46 of cytoplasmic tRNA, and this methylation has been shown to link with cancer progression and tumorigenesis,69 −71 requiring a quantitative method for monitoring tRNA m7G dynamics in pathological processes. Based on the chemical principle in m7G-seq, we further optimized it to be more quantitative (Figure), or m7G-quant-seq,13 without the need for the biotin pulldown enrichment. In m7G-quant-seq, we performed the depurination at a more acidic pH (∼2.9) under mild heating. The internal m7G sites could be almost completely converted to abasic sites when extending the incubation time of both KBH4 reduction and depurination.13 Note that the current m7G-quant-seq requires ∼200 ng of cellular small RNA as input, to avoid high PCR cycle number and the consequent PCR duplicates in NGS13 (Table S3↗).
Chemistry principle for base-resolution quantitative mapping of internal mG methylations. 7
CONCLUSION AND PERSPECTIVES
In summary, the lack of reliable sequencing methods to map modifications on mammalian mRNA has hampered functional investigations of these modifications. Applying chemical and biochemical knowledge we have developed a set of methods to map mRNA modifications at base resolution with modification stoichiometry information. These sequencing methods allow base resolution mapping of most major mRNA modifications often with limited input sample requirements, particularly m6A-SAC-seq, BID-seq, UBS-seq, and DAMM-seq, which could start with ∼2 ng, ∼10 ng, ∼10–20 ng, and ∼10 ng input RNA, respectively.1 −4 They could be used to uncover dynamic changes in modification stoichiometry during biological processes. The use of calibration spike-in probes could further enhance the accuracy of stoichiometry determination. For m6A-SAC-seq, the use of calibration probes is highly recommended to ensure accurate modification stoichiometry quantification at different sequence contexts because of the sequence context preference of the MjDim1 enzyme. Spike-in probes are not necessary for UBS-seq, since the chemical treatment displays an extremely high conversion ratio with no noticeable preference in sequence context around unmethylated cytidines. For other methods mentioned in this Account, such as BID-seq, Nm-Mut-seq, and m1A-quant-seq,2,10,11 calibration probes are recommended for reproducing observations on modification stoichiometry when using different batches of engineered or commercially available RT enzymes. These RT enzymes may have sequence preferences. The application scope of these methods goes beyond mRNA and could also be readily applied to other RNA species, such as nuclear nascent RNA, caRNA, cellular small RNA, cell-free RNA, etc.
The application of quantitative m6A-SAC-seq has revealed dynamics of m6A methylation stoichiometry during cell differentiation and characterized a number of cell-state-specific m6A sites.1 Quantitative BID-seq confirmed Ψ-modified stop codons within mammalian mRNAs, as in vivo on/off switches for converting nonsense codons into sense codons.2 In this Account, we emphasized the quantitative feature of our recently developed sequencing methods, and these developments may lead to large-scale high-dimensional maps on quantifying the dynamics of RNA modification stoichiometry in diverse biological processes, instead of only depicting the occurrence of an RNA modification type. Furthermore, the high conversion ratios of Ψ to Ψ-BS in BID-seq and C to U in UBS-seq and the ease of use, respectively, may make both BID-seq and UBS-seq common methods for future detection of Ψ and m5C modifications in mRNA and other RNA species. There are advantages and disadvantages for the base-resolution methods to map m6A, namely, m6A-SAC-seq, eTAM-seq, and GLORI.1,9,34 While m6A-SAC-seq exhibits sequence bias, it is a method that reads out m6A as a mutant, thereby maintaining the sequence complexity and notably reducing sequencing costs.
There is still plenty of space to optimize these methods to be more sensitive and more accurate and use less starting material. Single-cell level sequencing could reveal heterogeneity information and help further classify cell types. A future challenge is to sequence multiple modifications on the same RNA molecule. How compatible these methods are with each other and how practical to run multiple reactions before sequencing are questions that will need to be addressed. Chemical or biochemical labeling of specific modifications, coupled with long-read sequencing platforms such as nanopore or PacBio, may offer an attractive solution moving forward to simultaneously map multiple modifications on a single RNA molecule.
Acknowledgments
This project was supported by National Institutes of Health (NIH) grant RM1 HG008935 (C.H.). C.H. is an investigator at Howard Hughes Medical Institute.
Biographies
Li-Sheng Zhang earned his B.S. in Materials Chemistry from Peking University and received his Ph.D. in Chemistry at The University of Chicago where he worked with Professor Chuan He. He worked as a postdoctoral scholar at Professor He’s laboratory before starting his independent research in Hong Kong. Currently, he is a tenure-track Assistant Professor of chemistry and life science at The Hong Kong University of Science and Technology (HKUST). His research interests focus on quantitative sequencing technology and biological functions of RNA modifications.
Qing Dai earned his B.S. degree in chemistry from Central China Normal University and completed his Ph.D. in organic chemistry at Nankai University. He then carried out postdoctoral work in nucleic acid chemistry at the University of Chicago. Currently, he is a Research Professor at the University of Chicago. His research interests include developing new sequencing methods for DNA and RNA modifications and their applications.
Chuan He received his Ph.D. from Massachusetts Institute of Technology (MIT), where he worked with Professor Stephen J. Lippard. He did his postdoctoral work with Professor Gregory L. Verdine at Harvard University. Currently, He is a John T. Wilson Distinguished Service Professor at The University of Chicago and an investigator at Howard Hughes Medical Institute.
Supporting Information Available
The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.accounts.3c00532↗.
Author Contributions
C.H. conceived the original ideas and supervised the projects mentioned in this Account. L.-S.Z., Q.D., and C.H. wrote the manuscript. L.-S.Z. and Q.D. contributed equally to this work.
Author Contributions
⊥ These authors contributed equally.
The authors declare the following competing financial interest(s): C.H. is a scientific founder, member of the scientific advisory board, and equity holder of Aferna Bio, Inc. and AccuaDX Inc., a scientific cofounder and equity holder of Accent Therapeutics, Inc., and a member of the scientific advisory board of Rona Therapeutics.
Special Issue
Published as part of the Accounts of Chemical Research special issue “RNA Modifications↗ ”.