Cell Identity Codes: Understanding Cell Identity from Gene Expression Profiles using Deep Neural Networks

Open Access

20 February 2019

journal article
research article
Published by Springer Science and Business Media LLC in Scientific Reports

Vol. 9 (1), 1-14
https://doi.org/10.1038/s41598-019-38798-y

Abstract

Understanding cell identity is an important task in many biomedical areas. Expression patterns of specific marker genes have been used to characterize some limited cell types, but exclusive markers are not available for many cell types. A second approach is to use machine learning to discriminate cell types based on the whole gene expression profiles (GEPs). The accuracies of simple classification algorithms such as linear discriminators or support vector machines are limited due to the complexity of biological systems. We used deep neural networks to analyze 1040 GEPs from 16 different human tissues and cell types. After comparing different architectures, we identified a specific structure of deep autoencoders that can encode a GEP into a vector of 30 numeric values, which we call the cell identity code (CIC). The original GEP can be reproduced from the CIC with an accuracy comparable to technical replicates of the same experiment. Although we use an unsupervised approach to train the autoencoder, we show different values of the CIC are connected to different biological aspects of the cell, such as different pathways or biological processes. This network can use CIC to reproduce the GEP of the cell types it has never seen during the training. It also can resist some noise in the measurement of the GEP. Furthermore, we introduce classifier autoencoder, an architecture that can accurately identify cell type based on the GEP or the CIC.

This publication has 37 references indexed in Scilit:

A Self-Directed Method for Cell-Type Identification and Separation of Gene Expression Microarrays
PLoS Computational Biology, 2013
Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool
BMC Bioinformatics, 2013
Naive pluripotency is associated with global DNA hypomethylation
Nature Structural & Molecular Biology, 2013
NCBI GEO: archive for functional genomics data sets—update
Nucleic Acids Research, 2012
Human stem cell research and regenerative medicine--present and future
British Medical Bulletin, 2011
EEG signal classification using PCA, ICA, LDA and support vector machines
Expert Systems with Applications, 2010
ToppCluster: a multiple gene list feature analyzer for comparative enrichment clustering and network-based dissection of biological systems
Nucleic Acids Research, 2010
Oct4 expression revisited: potential pitfalls for data misinterpretation in stem cell research
Biological Chemistry, 2008
Reducing the Dimensionality of Data with Neural Networks
Science, 2006
Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles
Proceedings of the National Academy of Sciences of the United States of America, 2005

Cited by 14 articles