Clustering of highly homologous sequences to reduce the size of large protein databases.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 11294794.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
We present a fast and flexible program for clustering large protein databases at different sequence identity levels. It takes less than 2 h for the all-against-all sequence comparison and clustering of the non-redundant protein database of over 560,000 sequences on a high-end PC. The output database, including only the representative sequences, can be used for more efficient and sensitive database searches.
Medical subject headings
- Databases, Factual
- Proteins
- Software