Comparative Evaluation of a System One Model and a General-Purpose Large Language Model on the Korean Physical Therapist Licensing Examination
Journal:
medRxiv
Published Date:
Sep 28, 2026
Abstract
Background: System One Model is designed for structured decision-making and can provide probabilistic outputs, but their performance in domain-specific physical therapy tasks has not been established. Objective: To evaluate the performance and potential utility of Jev, a System One Model, for professional knowledge-based selection tasks in physical therapy by comparing it with GPT-4o. Methods: A total of 380 multiple-choice questions from the 2024 and 2025 Korean Physical Therapist Licensing Examinations were independently submitted to Jev and GPT-4o. The outcomes included answer accuracy, subgroup performance, output validity, API response latency, token usage, estimated API cost, and Jev's prediction uncertainty. A retrospective probability-based model-cascading simulation was also performed. Results: Jev correctly answered 282 questions (74.2%), whereas GPT-4o answered 326 (85.8%), a difference of -11.6 percentage points. Jev produced valid structured responses for all questions, whereas GPT-4o produced three invalid responses. Jev's selected-answer probabilities discriminated correct from incorrect answers (AUROC = 0.866), and accuracy reached 99.3% when probabilities were greater than or equal to 0.90. The median API response latency was 0.94 s for Jev and 1.06 s for GPT-4o, with estimated costs of US$0.008 and US$0.161, respectively. At a retrospective cascading threshold of 0.70, accuracy was 86.6% with 175 simulated GPT-4o calls. Conclusion: Jev showed lower overall accuracy than GPT-4o but provided reliable structured outputs, informative prediction probabilities, lower latency, and lower estimated cost. Its probabilistic outputs warrant further prospective evaluation for selective model escalation and decision-support workflows in physical therapy.