NLR-parser: rapid annotation of plant NLR complements.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 25586514.
- Also identified by DOI 10.1093/bioinformatics/btv005 and PMC identifier 4426836.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The repetitive nature of plant disease resistance genes encoding for nucleotide-binding leucine-rich repeat (NLR) proteins hampers their prediction with standard gene annotation software. Motif alignment and search tool (MAST) has previously been reported as a tool to support annotation of NLR-encoding genes. However, the decision if a motif combination represents an NLR protein was entirely manual. The NLR-parser pipeline is designed to use the MAST output from six-frame translated amino acid sequences and filters for predefined biologically curated motif compositions. Input reads can be derived from, for example, raw long-read sequencing data or contigs and scaffolds coming from plant genome projects. The output is a tab-separated file with information on start and frame of the first NLR specific motif, whether the identified sequence is a TNL or CNL, potentially full or fragmented. In addition, the output of the NB-ARC domain sequence can directly be used for phylogenetic analyses. In comparison to other prediction software, the highly complex NB-ARC domain is described in detail using several individual motifs.
Medical subject headings
- Arabidopsis
- Arabidopsis Proteins
- Immunity, Innate
- Molecular Sequence Annotation
- Plant Diseases
- Proteins
- Sequence Analysis, DNA