Reads2Vec: Efficient Embedding of Raw High-Throughput Sequencing Reads Data

1 April 2023

journal article
research article
Published by Mary Ann Liebert Inc in Journal of Computational Biology

Vol. 30 (4), 469-491
https://doi.org/10.1089/cmb.2022.0424

Abstract

The massive amount of genomic data appearing for SARS-CoV-2 since the beginning of the COVID-19 pandemic has challenged traditional methods for studying its dynamics. As a result, new methods such as Pangolin, which can scale to the millions of samples of SARS-CoV-2 currently available, have appeared. Such a tool is tailored to take as input assembled, aligned, and curated full-length sequences, such as those found in the GISAID database. As high-throughput sequencing technologies continue to advance, such assembly, alignment, and curation may become a bottleneck, creating a need for methods that can process raw sequencing reads directly. In this article, we propose Reads2Vec, an alignment-free embedding approach that can generate a fixed-length feature vector representation directly from the raw sequencing reads without requiring assembly. Furthermore, since such an embedding is a numerical representation, it may be applied to highly optimized classification and clustering algorithms. Experiments on simulated data show that our proposed embedding obtains better classification results and better clustering properties contrary to existing alignment-free baselines. In a study on real data, we show that alignment-free embeddings have better clustering properties than the Pangolin tool and that the spike region of the SARS-CoV-2 genome heavily informs the alignment-free clusterings, which is consistent with current biological knowledge of SARS-CoV-2.

Keywords

This publication has 44 references indexed in Scilit:

CoMeta: Classification of Metagenomes Using k-mers
PLOS ONE, 2015
Kraken: ultrafast metagenomic sequence classification using exact alignments
Genome Biology, 2014
The real cost of sequencing: higher than you think!
Genome Biology, 2011
Mind the Gaps: Evidence of Bias in Estimates of Multiple Sequence Alignments
Molecular Biology and Evolution, 2007
Reducing storage requirements for biological sequence comparison
Bioinformatics, 2004
Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
Journal of Computational and Applied Mathematics, 1987
A measure of the similarity of sets of sequences not requiring sequence alignment.
Proceedings of the National Academy of Sciences of the United States of America, 1986
Comparing partitions
Journal of Classification, 1985
Least squares quantization in PCM
IEEE Transactions on Information Theory, 1982
A dendrite method for cluster analysis
Communications in Statistics - Theory and Methods, 1974

Cited by 7 articles