The shortest common supersequence problem in a microarray production setting.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 14534185.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
During microarray production, several thousands of oligonucleotides (short DNA sequences) are synthesized in parallel, one nucleotide at a time. We are interested in finding the shortest possible nucleotide deposition sequence to synthesize all oligos in order to reduce production time and increase oligo quality. Thus we study the shortest common super-sequence problem of several thousand short strings over a four-letter alphabet. We present a statistical analysis of the basic ALPHABET-LEFTMOST approximation algorithm, and propose several practical heuristics to reduce the length of the super-sequence. Our results show that it is hard to beat ALPHABET-LEFTMOST in the microarray production setting by more than 2 characters, but these savings can improve overall oligo quality by more than four percent. Source code in C may be obtained by contacting the author, or from http://oligos.molgen.mpg.de.
Medical subject headings
- Algorithms
- DNA
- DNA Probes
- Oligonucleotide Array Sequence Analysis
- Sequence Alignment
- Sequence Analysis, DNA