NNAlign: A Web-Based Prediction Method Allowing Non-Expert End-User Discovery of Sequence Motifs in Quantitative Peptide Data

Open Access

2 November 2011

journal article
research article
Published by Public Library of Science (PLoS) in PLOS ONE

Vol. 6 (11), e26781
https://doi.org/10.1371/journal.pone.0026781

Abstract

Recent advances in high-throughput technologies have made it possible to generate both gene and protein sequence data at an unprecedented rate and scale thereby enabling entirely new “omics”-based approaches towards the analysis of complex biological processes. However, the amount and complexity of data that even a single experiment can produce seriously challenges researchers with limited bioinformatics expertise, who need to handle, analyze and interpret the data before it can be understood in a biological context. Thus, there is an unmet need for tools allowing non-bioinformatics users to interpret large data sets. We have recently developed a method, NNAlign, which is generally applicable to any biological problem where quantitative peptide data is available. This method efficiently identifies underlying sequence patterns by simultaneously aligning peptide sequences and identifying motifs associated with quantitative readouts. Here, we provide a web-based implementation of NNAlign allowing non-expert end-users to submit their data (optionally adjusting method parameters), and in return receive a trained method (including a visual representation of the identified motif) that subsequently can be used as prediction method and applied to unknown proteins/peptides. We have successfully applied this method to several different data sets including peptide microarray-derived sets containing more than 100,000 data points. NNAlign is available online at http://www.cbs.dtu.dk/services/NNAlign.

Keywords

This publication has 48 references indexed in Scilit:

Exploring Antibody Recognition of Sequence Space through Random-Sequence Peptide Microarrays
Molecular & Cellular Proteomics, 2011
MHC Class II epitope predictive algorithms
Immunology, 2010
Development of a novel peptide microarray for large-scale epitope mapping of food allergens
Journal of Allergy and Clinical Immunology, 2009
The Universal Protein Resource (UniProt)
Nucleic Acids Research, 2007
MEME: discovering and analyzing DNA and protein sequence motifs
Nucleic Acids Research, 2006
DILIMOT: discovery of linear motifs in proteins
Nucleic Acids Research, 2006
WebLogo: A Sequence Logo Generator: Figure 1
Genome Research, 2004
Prediction of lipoprotein signal peptides in Gram‐negative bacteria
Protein Science, 2003
Selection of representative protein data sets
Protein Science, 1992
Sequence logos: a new way to display consensus sequences
Nucleic Acids Research, 1990

Cited by 53 articles