Multi-model LLM assessment of Quality Control Circle methodological quality: a designed-anchor reliability study
Journal:
medRxiv
Published Date:
Oct 2, 2026
Abstract
Background: Quality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. Objective: To estimate inter-model reliability for QCC quality scoring and to assess alignment with designed anchors and descriptive comparability to a small set of public PMC QCC reports. Methods: Thirty synthetic QCC reports and eight public PMC QCC reports were evaluated across four primary LLM evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because a same-family model generated the synthetic cases. Evaluators received the report text only; the trap manifest and designed-anchor scores were withheld, and their absence was verified programmatically. Each case was scored on eight QCC quality dimensions in three runs per evaluator, summarized by median, and ICC(A,1) was estimated across the primary panel. Anchor calibration, trap detection from the evaluators' structured defect output under a map fixed before scoring, sensitivity analyses, and a descriptive synthetic-versus-PMC check were also examined. Results: Inter-model reliability on the primary k=4 panel was good: ICC(A,1) = 0.843 (95% CI 0.812 to 0.870; case-level bootstrap 0.782 to 0.874) on 239 pooled case-dimension rows from all 30 cases; per-dimension estimates ranged from 0.747 to 0.901. The pre-specified k=5 analysis including Claude gave 0.833, leave-one-out estimates ranged from 0.835 to 0.857, and the originally specified score parser gave 0.854. On the 52 trap-affected case-dimension cells the primary-panel consensus fell within 1 of the designed anchor in 50 (96.2%), with mean absolute error 0.54 and exact agreement 34.6%; 51.5% of individual primary-panel run scores equalled the anchor. A defect of the expected category was emitted in 87% of evaluator runs and by 94% of (trap, evaluator) pairs, versus 55% and 68% for a pre-specified keyword detector; because some categories are emitted at high base rates regardless of case, these are upper bounds. All eight dimensions met the descriptive synthetic-versus-PMC margin; synthetic cases tended to score higher, significantly so only for Action Alignment. Conclusions: Multi-model LLM scoring of QCC methodological quality achieved good inter-model reliability and directionally consistent alignment with synthetic anchor scores, with planted weaknesses identified far more often in structured defect output than by keyword matching. Benchmark feasibility is established; expert-labeled validation in clinical QCC reports is the planned next step.