CSpace: a concept embedding space for biomedical applications.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40576218.
- Also identified by DOI 10.1093/bioinformatics/btaf376 and PMC identifier 12275461.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The rise of transformer-based architectures has dramatically improved our ability to analyze natural language. However, the power and flexibility of these general-purpose models come at the cost of highly complex model architectures with billions of parameters that are not always needed. In this work, we present CSpace: a concise word embedding of biomedical concepts that outperforms all alternatives in terms of out-of-vocabulary ratio and semantic textual similarity task, and has comparable performance with respect to transformer-based alternatives in the sentence similarity task. This ability can serve as the foundation for semantic search by enabling efficient retrieval of conceptually related terms. Additionally, CSpace incorporates ontological identifiers (MeSH, NCBI gene and taxonomy IDs), enabling computationally efficient disease, gene or condition relatedness measurement, potentially unlocking previously unknown disease-condition associations. Full and compressed models are available on Zenodo at https://doi.org/10.5281/zenodo.14781672, while training code, examples, interactive visualizations and experiments are available at https://doi.org/10.5281/zenodo.15125706 and on the GitHub repository.
Medical subject headings
- Natural Language Processing
- Software
- Computational Biology