SelfCheck-Eval: A multi-module framework for zero-resource hallucination detection in large language models.

Muhammed, Diyana; Tuccari, Giusy Giulia; Rabby, Gollam; Auer, Sören; Vahdati, Sahar · Patterns (N Y) · 2026

Where this comes from

Abstract

Large language models (LLMs) have achieved considerable progress across diverse applications, yet their tendency to generate incorrect or fabricated content, commonly termed hallucinations, remains a fundamental obstacle to reliable deployment in high-stakes domains. Existing detection benchmarks are confined to general-knowledge settings, leaving specialized fields, where accuracy is important, underexplored. To address this gap, we introduce the American Invitational Mathematics Examination (AIME) Math Hallucination dataset, a benchmark for evaluating mathematical reasoning hallucinations, and propose SelfCheck-Eval, an LLM-agnostic, black-box detection framework compatible with open- and closed-source LLMs. The framework integrates three independent modules, semantic, specialized detection, and contextual consistency, into a suitable architecture. Systematic evaluation reveals a noticeable performance gap: existing methods perform well on biographical content but struggle with mathematical reasoning, a deficit that continues across natural language inference (NLI) fine-tuning, preference learning, and process supervision paradigms. These findings expose fundamental limitations of current approaches and motivate the development of specialized, black-box-compatible methods for trustworthy LLM deployment.