Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach.
Where this comes from
- Record sourced from PubMed, PMID 28541493.
- Also identified by DOI 10.1093/jamia/ocx053 and PMC identifier 5975623.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Identifying all published reports of randomized controlled trials (RCTs) is an important aim, but it requires extensive manual effort to separate RCTs from non-RCTs, even using current machine learning (ML) approaches. We aimed to make this process more efficient via a hybrid approach using both crowdsourcing and ML. We trained a classifier to discriminate between citations that describe RCTs and those that do not. We then adopted a simple strategy of automatically excluding citations deemed very unlikely to be RCTs by the classifier and deferring to crowdworkers otherwise. Combining ML and crowdsourcing provides a highly sensitive RCT identification strategy (our estimates suggest 95%-99% recall) with substantially less effort (we observed a reduction of around 60%-80%) than relying on manual screening alone. Hybrid crowd-ML strategies warrant further exploration for biomedical curation/annotation tasks.
Medical subject headings
- Crowdsourcing
- Information Storage and Retrieval
- Machine Learning
- Randomized Controlled Trials as Topic