Transcriptome sequencing in an ecologically important tree species: assembly, annotation, and marker discovery
Open Access
- 1 January 2010
- journal article
- Published by Springer Science and Business Media LLC in BMC Genomics
- Vol. 11 (1), 1-16
- https://doi.org/10.1186/1471-2164-11-180
Abstract
Massively parallel sequencing of cDNA is now an efficient route for generating enormous sequence collections that represent expressed genes. This approach provides a valuable starting point for characterizing functional genetic variation in non-model organisms, especially where whole genome sequencing efforts are currently cost and time prohibitive. The large and complex genomes of pines (Pinus spp.) have hindered the development of genomic resources, despite the ecological and economical importance of the group. While most genomic studies have focused on a single species (P. taeda), genomic level resources for other pines are insufficiently developed to facilitate ecological genomic research. Lodgepole pine (P. contorta) is an ecologically important foundation species of montane forest ecosystems and exhibits substantial adaptive variation across its range in western North America. Here we describe a sequencing study of expressed genes from P. contorta, including their assembly and annotation, and their potential for molecular marker development to support population and association genetic studies. We obtained 586,732 sequencing reads from a 454 GS XLR70 Titanium pyrosequencer (mean length: 306 base pairs). A combination of reference-based and de novo assemblies yielded 63,657 contigs, with 239,793 reads remaining as singletons. Based on sequence similarity with known proteins, these sequences represent approximately 17,000 unique genes, many of which are well covered by contig sequences. This sequence collection also included a surprisingly large number of retrotransposon sequences, suggesting that they are highly transcriptionally active in the tissues we sampled. We located and characterized thousands of simple sequence repeats and single nucleotide polymorphisms as potential molecular markers in our assembled and annotated sequences. High quality PCR primers were designed for a substantial number of the SSR loci, and a large number of these were amplified successfully in initial screening. This sequence collection represents a major genomic resource for P. contorta, and the large number of genetic markers characterized should contribute to future research in this and other pines. Our results illustrate the utility of next generation sequencing as a basis for marker development and population genomics in non-model species.Keywords
This publication has 60 references indexed in Scilit:
- Multilocus Patterns of Nucleotide Diversity and Divergence Reveal Positive Selection at Candidate Genes Related to Cold Hardiness in Coastal Douglas Fir (Pseudotsuga menziesii var. menziesii)Genetics, 2009
- Sequencing and de novo analysis of a coral larval transcriptome using 454 GSFlxBMC Genomics, 2009
- Rapidly developing functional genomics in ecological model systems via 454 transcriptome sequencingGenetica, 2008
- Combining population genomics and quantitative genetics: finding the genes underlying ecologically important traitsHeredity, 2007
- Association Genetics in Pinus taeda L. I. Wood Property TraitsGenetics, 2007
- The molecular ecologist's guide to expressed sequence tagsMolecular Ecology, 2006
- PpRT1: the first complete gypsy-like retrotransposon isolated in Pinus pinasterPlanta, 2006
- A wing expressed sequence tag resource for Bicyclus anynana butterflies, an evo-devo modelBMC Genomics, 2006
- Nucleotide Diversity and Linkage Disequilibrium in Cold-Hardiness- and Wood Quality-Related Candidate Genes in Douglas FirGenetics, 2005
- Expressed Sequence Tag-Linked Microsatellites as a Source of Gene-Associated Polymorphisms for Detecting Signatures of Divergent Selection in Atlantic Salmon (Salmo salar L.)Molecular Biology and Evolution, 2005