How to Train Your Chatbot: Information-Theoretic Foundations of Diagnostic Questioning in Inborn Errors of Immunity.

Lugo Reyes, Saul O; Vásquez Echeverri, Estefanía; Bustamante Ogando, Juan Carlos; Castano-Jaramillo, Lina M; Vélez Tirado, Natalia; Venegas Montoya, Edna; Tarango García, Alejandro; Gómez Tello, Héctor et al. · Allergy · 2026

other · Level V

Where this comes from

Abstract

Navigating the more than 550 inborn errors of immunity (IEI) requires efficient diagnostic reasoning. Information theory suggests questions should be prioritized by their capacity to reduce diagnostic uncertainty (entropy); yet whether experts or large language models (LLMs) optimize for information gain remains unquantified. We compared expert clinician and LLM diagnostic prioritization strategies using an information-theoretic framework. Fifteen immunologists and six LLMs (ChatGPT, Claude, Gemini, Grok, DeepSeek, and Llama) ranked 35 diagnostic questions by efficiency. Shannon's entropy was used to estimate expected information gain (EIG) for each question. Agreement was assessed via Spearman correlations, consensus ranking, and principal components analysis (PCA). Clinician consensus rankings strongly correlated with estimated information gain (Spearman ρ = -0.71, p < 0.001). "Age at onset?" ranked first by clinicians, provided the highest information gain (2.29 bits), reducing diagnostic uncertainty by 80%. Clinicians and LLMs showed strong agreement on top-tier discriminators (Spearman ρ = 0.73, p < 0.001). However, PCA revealed a distinct LLM cluster; clinicians prioritized bedside/history questions, whereas LLMs favored syndromic and laboratory features. Optimal questioning reached diagnostic confidence in 4-5 steps, approaching the theoretical minimum. Expert clinicians implicitly approximate information-theoretic optimization in IEI diagnostics. While LLMs share a core heuristic for high-yield questions, divergence in mid-sequence reasoning suggests a shift from experiential heuristics to probabilistic data-matching. This framework provides a principled basis for training AI-assisted tools that mirror expert diagnostic logic.