The inadequacy of offline large language model evaluations: A need to account for personalization in model behavior.
other · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41472831.
- Also identified by DOI 10.1016/j.patter.2025.101397 and PMC identifier 12745978.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Standard offline evaluations for language models fail to capture how these models actually behave in practice, where personalization fundamentally alters model behavior. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other questions to their chat interfaces.