PLoS computational biology

Using deep learning and similarity-based transfer to better predict CRISPR-Cas9 off-target effects

Updated

Abstract

Essence

Similarity-guided improved CRISPR-Cas9 off-target prediction, with cosine distance and several deep neural networks performing best in these simulations.

Evidence

This computational benchmarking study compared similarity metrics plus four deep learning and two traditional machine-learning models across CRISPR-Cas9 source and target datasets for transfer-learning-based off-target prediction.

Caveat

The findings come from simulation and dataset-comparison experiments, so they show better predictive performance rather than experimental proof in biological editing systems.

Simplified

Key numbers

0.9585
Cosine Similarity for CD33_BS
Similarity score between CD33 source dataset and CD33_BS target dataset.
0.9863
ROC for MLP1 on CD33_BS
Performance metric for MLP1 model using CD33 dataset as source.
0.012
Class Imbalance Ratio for CIRCLE_BS
Class imbalance ratio for the bootstrapped CIRCLE_BS target dataset.

Key figures

Fig 1
Similarity analysis vs : selecting and applying source datasets for CRISPR
Highlights how similarity-based source selection enhances transfer learning accuracy in CRISPR off-target prediction
pcbi.1013606.g001
  • Panel A
    Shows use of three distance metrics (cosine, Euclidean, Manhattan) to compare encoded source datasets (CD33, CIRCLE, SITE) with encoded target datasets to identify the optimal source dataset for transfer learning
  • Panel B
    Illustrates transfer learning process where learned model knowledge from the selected source dataset is transferred to the target dataset to improve off-target prediction accuracy
Fig 2
Encoding of -DNA sequence pairs as fixed-length matrices for insertion, , and deletion cases
Frames how sequence differences like insertions and mismatches are systematically encoded for CRISPR
pcbi.1013606.g002
  • Panel Insertion
    A 7 × L with a five-bit character channel and two-bit direction channel representing an insertion event with DNA/RNA bulges indicated by underscores
  • Panel Mismatch
    A 7 × L encoded matrix showing nucleotide mismatches between on-target and off-target sequences with corresponding direction channel encoding
  • Panel Deletion
    A 7 × L encoded matrix illustrating a deletion event with bulges and direction channel marking mismatches and
Fig 3
architectures for , , and in CRISPR-Cas9
Highlights consistent transfer learning architectures enabling improved target predictions across different neural network types
pcbi.1013606.g003
  • Panels FNN
    Source and target input matrices feed into a multi-layer FNN with dense layers, , and dropout layers, producing target predictions
  • Panels CNN
    Source and target input matrices feed into a CNN with three convolutional layers, flattening, dense, and dropout layers, producing target predictions
  • Panels RNN
    Source and target input matrices feed into an RNN with two RNN layers, batch normalization, dense, and dropout layers, producing target predictions
Fig 4
Similarity scores between three source and seven target CRISPR-Cas9 datasets using three distance metrics
Highlights higher similarity scores within source datasets, anchoring dataset selection for
pcbi.1013606.g004
  • Panels (a-c)
    Similarity scores for CD33 dataset using cosine, Euclidean, and Manhattan distances; CD33 shows highest similarity to itself with cosine (0.9585), Euclidean (0.7843), and Manhattan (0.8933) distances
  • Panels (d-f)
    Similarity scores for CIRCLE dataset using cosine, Euclidean, and Manhattan distances; CIRCLE shows highest similarity to itself with cosine (0.865), Euclidean (0.445), and Manhattan (0.6809) distances
  • Panels (g-i)
    Similarity scores for SITE dataset using cosine, Euclidean, and Manhattan distances; SITE shows highest similarity to itself with cosine (0.8841), Euclidean (0.4555), and Manhattan (0.7052) distances
Fig 5
ROC curves for models trained on CD33, CIRCLE, and SITE datasets evaluated on bootstrapped targets
Highlights higher values for models in CIRCLE dataset versus lower AUCs for CNN10 in SITE dataset
pcbi.1013606.g005
  • Panel A
    ROC curves for models trained on the CD33 dataset and evaluated on the CD33_BS dataset; MLP1 has the highest AUC (0.986 ± 0.003), and has the lowest AUC (0.886 ± 0.011)
  • Panel B
    ROC curves for models trained on the CIRCLE dataset and evaluated on the CIRCLE_BS dataset; MLP1 and MLP2 have the highest AUCs (0.996 ± 0.007), CNN10 has the lowest AUC (0.737 ± 0.063)
  • Panel C
    ROC curves for models trained on the SITE dataset and evaluated on the SITE_BS dataset; FFN3 has the highest AUC (0.994 ± 0.01), CNN10 has the lowest AUC (0.777 ± 0.031)
1 / 5

Full Text

What this is

  • This research investigates the use of to enhance CRISPR-Cas9 off-target prediction accuracy.
  • It emphasizes the importance of selecting appropriate source datasets based on similarity metrics.
  • The study compares various deep learning architectures and traditional machine learning models to identify the best performers.

Essence

  • Similarity-based significantly improves CRISPR-Cas9 off-target predictions by optimizing source dataset selection. Cosine distance outperforms other metrics in identifying suitable datasets, leading to better predictive accuracy.

Key takeaways

  • Cosine distance is the most effective metric for assessing dataset similarity in for CRISPR-Cas9. It provides higher overall similarity values compared to Euclidean and Manhattan distances.
  • The study identifies that models like MLP variants and RNN-GRU achieve superior performance in off-target predictions. These models consistently outperform traditional machine learning methods across various evaluation metrics.
  • The proposed framework streamlines the process by focusing on similarity-based source dataset selection, which mitigates the risk of negative transfer and enhances predictive accuracy.

Caveats

  • The study relies on the quality and representativeness of the datasets used. If the datasets are not representative of the target task, prediction accuracy may suffer.
  • While cosine distance proved effective, its performance may vary with different types of datasets or applications beyond CRISPR-Cas9, requiring further validation in other contexts.

Definitions

  • Transfer Learning: A machine learning technique that leverages knowledge from a large source dataset to improve model performance on a smaller target dataset.
  • Off-target effect: Unintended modifications in the genome caused by CRISPR-Cas9 targeting regions other than the intended DNA sequence.

Simplified

Funding

Competing interests

0 of 3
authors report competing interests
3 report none
PubMed

What Lands in Your Inbox Each Week:

  • 📚7 fresh studies
  • 📝plain-language summaries
  • direct links to original studies
  • 🏅top journal indicators
  • 📅weekly delivery
  • 🧘‍♂️always free