NNAlign: A Web-Based Prediction Method Allowing Non-Expert End-User Discovery of Sequence Motifs in Quantitative Peptide Data
Open Access
- 2 November 2011
- journal article
- research article
- Published by Public Library of Science (PLoS) in PLOS ONE
- Vol. 6 (11), e26781
- https://doi.org/10.1371/journal.pone.0026781
Abstract
Recent advances in high-throughput technologies have made it possible to generate both gene and protein sequence data at an unprecedented rate and scale thereby enabling entirely new “omics”-based approaches towards the analysis of complex biological processes. However, the amount and complexity of data that even a single experiment can produce seriously challenges researchers with limited bioinformatics expertise, who need to handle, analyze and interpret the data before it can be understood in a biological context. Thus, there is an unmet need for tools allowing non-bioinformatics users to interpret large data sets. We have recently developed a method, NNAlign, which is generally applicable to any biological problem where quantitative peptide data is available. This method efficiently identifies underlying sequence patterns by simultaneously aligning peptide sequences and identifying motifs associated with quantitative readouts. Here, we provide a web-based implementation of NNAlign allowing non-expert end-users to submit their data (optionally adjusting method parameters), and in return receive a trained method (including a visual representation of the identified motif) that subsequently can be used as prediction method and applied to unknown proteins/peptides. We have successfully applied this method to several different data sets including peptide microarray-derived sets containing more than 100,000 data points. NNAlign is available online at http://www.cbs.dtu.dk/services/NNAlign.Keywords
This publication has 48 references indexed in Scilit:
- Exploring Antibody Recognition of Sequence Space through Random-Sequence Peptide MicroarraysMolecular & Cellular Proteomics, 2011
- MHC Class II epitope predictive algorithmsImmunology, 2010
- Development of a novel peptide microarray for large-scale epitope mapping of food allergensJournal of Allergy and Clinical Immunology, 2009
- The Universal Protein Resource (UniProt)Nucleic Acids Research, 2007
- MEME: discovering and analyzing DNA and protein sequence motifsNucleic Acids Research, 2006
- DILIMOT: discovery of linear motifs in proteinsNucleic Acids Research, 2006
- WebLogo: A Sequence Logo Generator: Figure 1Genome Research, 2004
- Prediction of lipoprotein signal peptides in Gram‐negative bacteriaProtein Science, 2003
- Selection of representative protein data setsProtein Science, 1992
- Sequence logos: a new way to display consensus sequencesNucleic Acids Research, 1990