AI-based meibography grading shows lower accuracy than human evaluators in diagnosing MGD
Artificial intelligence (AI)–based analysis of meibography images may be less accurate than human graders for diagnosing meibomian gland dysfunction (MGD), according to a new study.
A systematic review of 14 studies included 5,511 predominantly middle-aged participants and analyzed 18,926 meibography images obtained using noncontact infrared imaging or in vivo confocal microscopy. Investigators followed Cochrane diagnostic test accuracy methods, using standardized tools to assess bias and evidence quality.
Across the included studies, most AI models were internally validated, with only two studies reporting external validation and one reporting both. Nearly all studies had a high risk of bias in at least 1 domain, and most raised concerns about applicability.
Based on 3 externally validated datasets, AI models demonstrated a pooled sensitivity of 97.5% and specificity of 85.5% in distinguishing MGD from normal glands. However, variability among internally validated models was attributed to differences in study populations and case mix.
Overall, the certainty of evidence was rated very low to low due to imprecision, bias, and limited generalizability. The authors concluded that AI-based meibography grading currently underperforms compared with human evaluation and emphasized the need for more rigorous study designs, broader datasets, and external validation in future research.
Reference
Liu SH, Shah M, Leslie L, et al. Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies. Cornea. 2026;doi: 10.1097/ICO.0000000000004151. Epub ahead of print. PMID: 41931504.