Evaluating the performance of general purpose large language models in identifying human facial emotions.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41102392.
- Also identified by DOI 10.1038/s41746-025-01985-5 and PMC identifier 12533101.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
We evaluated the ability of three leading LLMs (GPT-4o, Gemini 2.0 Experimental, and Claude 3.5 Sonnet) to recognize human facial expression using the NimStim dataset. GPT and Gemini matched or exceeded human performance, especially for calm/neutral and surprise. All models showed strong agreement with ground truth, though fear was often misclassified. Findings underscore the growing socioemotional competence of LLMs and their potential for healthcare applications.