Knowledge-based citation reasoning for biomedical domain.
Where this comes from
- Record sourced from PubMed, PMID 41734273.
- Also identified by DOI 10.1093/bioinformatics/btag061 and PMC identifier 12987763.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Citation is central to scholarly communication, enabling researchers to navigate rapidly expanding literature and identify relevant prior work. Yet the 'reasoning' behind why a particular paper is cited is often implicit or opaque. Although academic search engines and literature tools rank candidate papers for a query, the motivations underlying these rankings are rarely transparent, making it difficult for scholars to interpret and act on retrieved results-especially in biomedical research where domain knowledge is essential. We propose an encoder-decoder framework that leverages curated biomedical knowledge to generate 'explanations of citation motivation' in a structured bio-triplet format. We evaluate the approach against recent families of pre-trained language models for text generation, including BERT-style (and variants) and GPT-style (and variants) models. In cancer-focused experiments using PubMed Central, we annotate over 10 000 citation relations with bio-triplets grounded in curated knowledge from multiple biomedical databases. Trained on these annotations, our model outperforms strong sequence-generation baselines, improving precision, recall, and F1 for citation-motivation generation. Code and data are available at Zenodo (archival DOI: 10.281/zenodo.14893445) and GitHub: https://github.com/zhongxiangboy/Knowledge-based-Citation-Reasoning-for-Biomedical-Domain.
Medical subject headings
- Biomedical Research
- Computational Biology
- Data Mining