Health artificial intelligence is here, but are we measuring what matters?

Payne, Philip R O; Kannampallil, Thomas; Lozovatsky, Margaret · J Am Med Inform Assoc · 2026

other · Level V

Where this comes from

Abstract

Artificial intelligence (AI) is increasingly being used in healthcare settings, yet evidence of its real-world value remains inconsistent. Current evaluation paradigms often emphasize methodological rigor and technical validity over measurable improvements in patient outcomes or system performance. To examine limitations in prevailing approaches to health AI evaluation and propose a framework prioritizing outcomes-based, systems-level assessment aligned with healthcare delivery goals. This perspective analyzes current evaluation practices through conceptual and ethical lenses, contrasting a deontological focus on methodological standards with a consequentialist framework emphasizing real-world impact. A persistent gap exists between how AI systems are evaluated and how their value is realized. Technical metrics are necessary but insufficient; meaningful evaluation requires measuring clinical and operational outcomes. Strategies include standardized outcome frameworks, evaluation infrastructure, multistakeholder governance, and aligned incentives. Advancing health AI requires shifting from process-focused evaluation toward outcome-based assessment embedded within healthcare systems.