PMC text mining subset in BioC: about three million full-text articles and growing.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 30715220.
- Also identified by DOI 10.1093/bioinformatics/btz070 and PMC identifier 6748740.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Interest in text mining full-text biomedical research articles is growing. To facilitate automated processing of nearly 3 million full-text articles (in PubMed Central® Open Access and Author Manuscript subsets) and to improve interoperability, we convert these articles to BioC, a community-driven simple data structure in either XML or JavaScript Object Notation format for conveniently sharing text and annotations. The resultant articles can be downloaded via both File Transfer Protocol for bulk access and a Web API for updates or a more focused collection. Since the availability of the Web API in 2017, our BioC collection has been widely used by the research community. https://www.ncbi.nlm.nih.gov/research/bionlp/APIs/BioC-PMC/.
Medical subject headings
- Data Mining