Visual steatosis assessment: method-dependent variation and systematically higher estimates compared with AI measurement.

Journal: Histopathology
Published Date:

Abstract

AIMS: Despite its clinical importance, histological steatosis assessment remains poorly standardized and highly variable among pathologists. We examined determinants of this variability by comparing pathologist estimates with quantitative artificial intelligence (AI) based measurements. METHODS AND RESULTS: Ten experienced liver pathologists from four institutions evaluated a limited set of 20 permanent H&E-stained whole-slide images (WSIs) from metabolic dysfunction-associated steatotic liver disease (MASLD) cases (10 core biopsies, 7 wedge biopsies and 3 resections) and 6 additional 500 × 500-μm image fields from separate permanent H&E-stained cases selected to represent varying droplet compositions. The WSIs spanned the full spectrum of steatosis severity. Assessments used four phases: default method, steatosis proportionate area (SPA), percentage of hepatocytes containing fat (%HCF) and Banff large-droplet criteria. Image fields were also assessed using no size cut-off or cut-offs based on hepatocyte, nuclear or 2-3× nuclear size. AI models quantified SPA, %HCF and droplet size for comparison with pathologists' assessments. Pathologists differed substantially in terminology, droplet-size definitions and quantification methods. Method choice (SPA vs. %HCF) was a major determinant of estimate dispersion and discrepancy from AI measurements. Visual estimates exceeded AI quantification by 1.4-3.6×. Pathologists' SPA estimates were increasingly higher than the corresponding AI measurements with increasing burden of droplets with area below 50 μm2, whereas, in exploratory analysis, %HCF estimates remained relatively stable relative to the corresponding AI measurements. Mixed-effects calibration equations were derived to relate visual and digital assessment scales. CONCLUSION: Steatosis assessment lacks standardized terminology and measurement practices. SPA and %HCF are conceptually distinct, and their interchangeable use amplifies discrepancies, particularly in small-droplet-rich cases. These findings support standardized, AI-compatible quantification; the exploratory conversion equations require independent validation before clinical application.

Authors

Keywords

No keywords available for this article.