Can large language models provide high-quality desk review decisions in an orthopaedic surgery journal? A concordance study comparing three AI models to human editorial decisions.

Journal: Orthopaedics & traumatology, surgery & research : OTSR
Published Date:

Abstract

BACKGROUND: Artificial intelligence (AI) is increasingly integrated into scientific publishing workflows, yet no study has formally evaluated the ability of large language models (LLMs) to reproduce human editorial desk-review (R0) decisions in a general orthopaedic surgery journal. We investigated whether three commercially available LLMs could accurately replicate the editorial decisions of the Editorial Board of Orthopaedics & Traumatology: Surgery & Research (OTSR). The study addressed four questions: (1) Is the concordance between LLM and human R0 decisions satisfactory for editorial use? (2) Do LLMs exhibit a severity bias? (3) Do LLMs generate decision letters of acceptable quality, and do they reproduce the specific critiques of human reviewers? (4) Does prompt complexity influence LLM decision-making? HYPOTHESIS: LLMs used without task-specific fine-tuning or prior exposure to the study corpus would demonstrate at least moderate concordance (κ ≥ 0.40) with human editorial decisions. MATERIAL AND METHODS: A corpus of 32 manuscripts randomly selected from submissions to OTSR between 2025 and 2026 (n = 10 outright rejected at R0: 3 out of scope, 3 plagiarism/dual submission, 4 direct desk rejection; n = 11 accepted for peer review; n = 11 rejected after full peer review) was anonymised and independently evaluated, without task-specific fine-tuning or prior exposure to the study corpus, by ChatGPT (GPT-5.5, OpenAI), Gemini (3.1, Google), and Claude (Sonnet 4.6, Anthropic) using a structured prompt incorporating the OTSR guidelines. The primary outcome was assessed using Cohen's kappa between LLM and human binary decisions. Secondary outcomes included accuracy, inter-LLM agreement, domain-specific scoring, ARCADIA quality scoring of 115 eligible decision letters by two blinded raters with ICC, human-performed content concordance analysis, sensitivity analysis (structured vs. minimal prompt), and test-retest reproducibility at 24 hours. RESULTS: Overall accuracy (i.e. the decision was similar for LLM and editorial decision) was 59.4% for ChatGPT (19/32) and Claude (19/32), and 62.5% for Gemini (20/32). Cohen's kappa was near-zero for ChatGPT (κ = -0.05) and Gemini (κ = 0.00), and low for Claude (κ = 0.15). All LLMs showed systematic over-rejection of accepted manuscripts. No LLM identified plagiarism or simultaneous dual submission as a rejection motive. Test-retest concordance was 84.4 - 90.6% across models. ARCADIA quality scoring (n = 115 scorable letters, inter-rater ICC = 0.86, 95% CI 0.80-0.90) showed Claude achieved the highest scores (4.46 ± 0.32 /5), significantly above the human OTSR letters (4.01 ± 0.44, p < 0.001), ChatGPT (3.81 ± 0.49, p = 0.002), and Gemini (3.28 ± 0.49, p < 0.001). LLMs reproduced 30 - 41% of human reviewer-specific critiques, with Claude achieving the highest match (40.5%) without hallucinations. Switching to a minimal prompt markedly increased acceptance rates for Gemini (87.5%) and Claude (75%), while ChatGPT remained largely insensitive to prompt simplification (15.6%). CONCLUSION: The principal finding of this study is the dissociation between formal review quality and true editorial reliability. Although modern LLMs generated persuasive and methodologically structured decision letters, they failed to achieve meaningful concordance with real editorial decisions and displayed stable architecture-specific biases that were highly sensitive to prompt design. These results indicate that current LLMs reproduce the surface features of peer review more successfully than its underlying scientific and contextual reasoning. Consequently, LLMs may represent valuable supervised assistants for editorial workflows, but not reliable autonomous substitutes for human editorial expertise in orthopaedic scientific publishing. LEVEL OF EVIDENCE: IV; Observational pilot study, concordance analysis.

Authors

Keywords

No keywords available for this article.