Evaluating the performance of general purpose large language models in identifying human facial emotions.

Nelson, Benjamin W; Winbush, Ari; Siddals, Steven; Flathers, Matthew; Allen, Nicholas B; Torous, John · NPJ Digit Med · 2025

basic_science · Level V

Where this comes from

Abstract

We evaluated the ability of three leading LLMs (GPT-4o, Gemini 2.0 Experimental, and Claude 3.5 Sonnet) to recognize human facial expression using the NimStim dataset. GPT and Gemini matched or exceeded human performance, especially for calm/neutral and surprise. All models showed strong agreement with ground truth, though fear was often misclassified. Findings underscore the growing socioemotional competence of LLMs and their potential for healthcare applications.