Benchmarking large language models for genomic knowledge with GeneTuring.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 40984701.
- Also identified by DOI 10.1093/bib/bbaf492 and PMC identifier 12454257.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1600 curated questions, and manually evaluated 48 000 answers from 10 LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom Generative Pretrained Transformer (GPT) setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with National Center for Biotechnology Information (NCBI) Application Programming Interfaces (APIs), developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.
Medical subject headings
- Genomics
- Benchmarking
- Software
- Programming Languages
- Computational Biology