PGP: parallel prokaryotic proteogenomics pipeline for MPI clusters, high-throughput batch clusters and multicore workstations.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 24470574.
- Also identified by DOI 10.1093/bioinformatics/btu051 and PMC identifier 4016709.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
We present the first public release of our proteogenomic annotation pipeline. We have previously used our original unreleased implementation to improve the annotation of 46 diverse prokaryotic genomes by discovering novel genes, post-translational modifications and correcting the erroneous annotations by analyzing proteomic mass-spectrometry data. This public version has been redesigned to run in a wide range of parallel Linux computing environments and provided with the automated configuration, build and testing facilities for easy deployment and portability. Source code is freely available from https://bitbucket.org/andreyto/proteogenomics under GPL license. It is implemented in Python and C++. It bundles the Makeflow engine to execute the workflows. atovtchi@jcvi.org.
Medical subject headings
- Cluster Analysis
- Genomics
- High-Throughput Nucleotide Sequencing
- Prokaryotic Cells
- Proteomics