Literature-informed gene extraction and ranking for multimodal data fusion.
Where this comes from
- Record sourced from PubMed, PMID 42402042.
- Also identified by DOI 10.1093/bib/bbag348.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Published biomedical experiments provide an increasingly rich collection of results and identify genes potentially involved in diverse biological mechanisms. However, individual studies are often confined to narrow experimental contexts and are restricted to single omics layers. Cross-study knowledge aggregation can broaden this perspective and enable the construction of global, context-aware gene rankings. Recent developments in natural language processing have made large-scale literature mining increasingly feasible. This enables the systematic extraction and fusion of symbolic knowledge from published experiments. We present pathXcite, a software that extracts genes associated with specific contexts, such as diseases or biological mechanisms from the literature, and ranks them by contextual relevance. These relevance-based gene rankings can compress a scientific context into a symbolic representation. This representation enables diverse downstream analyses, including cross-context comparisons, network-based analysis, enrichment analysis, and integration with experimental omics data. In multiple use cases, we show how our extraction and fusion strategy can be applied to uncover hidden aspects in biological data.
Medical subject headings
- Data Mining
- Software
- Computational Biology