Skip to main content

Research Repository

Advanced Search

ProteoAnnotator--open source proteogenomics annotation software supporting PSI standards.

ProteoAnnotator--open source proteogenomics annotation software supporting PSI standards. Thumbnail


Abstract

The recent massive increase in capability for sequencing genomes is producing enormous advances in our understanding of biological systems. However, there is a bottleneck in genome annotation--determining the structure of all transcribed genes. Experimental data from MS studies can play a major role in confirming and correcting gene structure--proteogenomics. However, there are some technical and practical challenges to overcome, since proteogenomics requires pipelines comprising a complex set of interconnected modules as well as bespoke routines, for example in protein inference and statistics. We are introducing a complete, open source pipeline for proteogenomics, called ProteoAnnotator, which incorporates a graphical user interface and implements the Proteomics Standards Initiative mzIdentML standard for each analysis stage. All steps are included as standalone modules with the mzIdentML library, allowing other groups to re-use the whole pipeline or constituent parts within other tools. We have developed new modules for pre-processing and combining multiple search databases, for performing peptide-level statistics on mzIdentML files, for scoring grouped protein identifications matched to a given genomic locus to validate that updates to the official gene models are statistically sound and for mapping end results back onto the genome. ProteoAnnotator is available from http://www.proteoannotator.org/. All MS data have been deposited in the ProteomeXchange with identifiers PXD001042 and PXD001390 (http://proteomecentral.proteomexchange.org/dataset/PXD001042; http://proteomecentral.proteomexchange.org/dataset/PXD001390).

Acceptance Date Oct 2, 2014
Publication Date Dec 1, 2014
Publicly Available Date Mar 29, 2024
Journal Proteomics
Print ISSN 1615-9853
Publisher Wiley
Pages 2731 - 2741
DOI https://doi.org/10.1002/pmic.201400265
Keywords Open source, ProteoAnnotator, Proteogenomics, Proteomics Standards Initiative, mzIdentML, Genomics, Proteins, Proteomics, Software
Publisher URL http://onlinelibrary.wiley.com/doi/10.1002/pmic.201400265/abstract

Files




Downloadable Citations