Artificial intelligence performance in generating colorectal surgery board questions.
Journal:
American journal of surgery
Published Date:
Feb 10, 2026
Abstract
BACKGROUND: Large language models (LLM) can pass medical licensing and specialty board exams, but their ability to generate high-quality board-style exam questions is uncertain. METHODS: Three LLMs each generated 20 colorectal surgery board questions in accordance with American Board of Colon and Rectal Surgery guidelines. Questions from the Colon and Rectal Surgery Educational Program (CARSEP) served as comparators. Board-certified colorectal surgeons, blinded to source, graded each question on clarity, relevance, suitability, distractor quality, and adequacy of rationale, and categorized questions as "Approved for Committee," "Author to Review," or "Not Accepted." RESULTS: CARSEP demonstrated the highest "Approved for Committee" rate (65%), compared with ChatGPT-4o (7%), Copilot Pro (10%), and Gemini Advanced (10%). CARSEP significantly outperformed most LLMs across all evaluation domains (p < 0.001), except for question suitability, where most LLM questions received >70% very good to good ratings. CONCLUSIONS: Although LLMs demonstrate potential, they are currently unable to consistently generate high-quality, colorectal surgery board-style questions.
Authors
Keywords
No keywords available for this article.