A local sequence alignment approach to recognizing fixed poetic forms across languages.
Where this comes from
- Record sourced from PubMed, PMID 42447153.
- Also identified by DOI 10.1371/journal.pone.0340514 and PMC identifier 13367689.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Fixed poetic forms such as the sonnet, ottava rima, or terza rima are an important feature of European literary traditions, yet large-scale empirical research on their cross-lingual distribution and evolution has been limited so far. This paper introduces a fully language-independent, unsupervised method for identifying recurrent rhyme-based forms using local sequence alignment. Drawing on 187,719 poems from six European traditions (Czech, English, French, German, Italian, Russian) in the PoeTree collection, we encode rhyme schemes in a compact eight-symbol alphabet and apply the Smith-Waterman algorithm via the Metronome package to compute pairwise distances. Dimensionality reduction (UMAP) and density-based clustering (HDBSCAN) yield 61 clusters, many of which align with known fixed forms. Evaluation against existing Czech and Russian annotations shows strong recall, while supervised classification experiments-both within and across languages-demonstrate that form categories are robustly learnable in the induced vector space. We illustrate the potential of such data for literary research in three showcases: cross-tradition influence in 19th-century Czech poetry, topical affinities of selected forms using multilingual topic modeling, and geographic associations revealed through geonym analysis.
Medical subject headings
- Language
- Sequence Alignment
- Poetry as Topic