Dies ist eine Übersichtsseite mit Metadaten zu dieser wissenschaftlichen Arbeit. Der vollständige Artikel ist beim Verlag verfügbar.

Performance of AI Models vs. Orthopedic Residents in Turkish Specialty Training Development Exams in Orthopedics

2025·2 Zitationen·SiSli Etfal Hastanesi Tip Bulteni / The Medical Bulletin of Sisli HospitalOpen Access

Volltext beim Verlag öffnen

Zitationen

Autoren

2025

Jahr

Abstract

A rtificial intelligence (AI) is a rapidly developing tech- nology in medicine that is used in many areas, from medical decision support systems to diagnosis. [1]Specifically, large language models, such as ChatGPT-4o, Gemini, Bing AI, and DeepSeek, are trained on different types of textual data (publicly available content, licensed materials, academic publications, news, etc.), thereby acquiring a broad range of topics and diverse information.All of these Objectives: As artificial intelligence (AI) continues to advance, its integration into medical education and clinical decision making has attracted considerable attention.Large language models, such as ChatGPT-4o, Gemini, Bing AI, and DeepSeek, have demonstrated potential in supporting healthcare professionals, particularly in specialty training examinations.However, the extent to which these models can independently match or surpass human performance in specialized medical assessments remains uncertain.This study aimed to systematically compare the performance of these AI models with orthopedic residents in the Specialty Training Development Exams (UEGS) conducted between 2010 and 2021, focusing on their accuracy, depth of explanation, and clinical applicability.Methods: This retrospective comparative study involved presenting the UEGS questions to ChatGPT-4o, Gemini, Bing AI, and DeepSeek.Orthopedic residents who took the exams during 2010-2021 served as the control group.The responses were evaluated for accuracy, explanatory details, and clinical applicability.Statistical analysis was conducted using SPSS Version 27, with one-way ANOVA and post-hoc tests for performance comparison.Results: All AI models outperformed orthopedic residents in terms of accuracy.Bing AI demonstrated the highest accuracy rates (64.0% to 93.0%), followed by Gemini (66.0% to 87.0%) and DeepSeek (63.5% to 81.0%).ChatGPT-4o showed the lowest accuracy among AI models (51.0% to 59.5%).Orthopedic residents consistently had the lowest accuracy (43.95% to 53.45%).Bing AI, Gemini, and DeepSeek showed knowledge levels equivalent to over 5 years of medical experience, while ChatGPT-4o ranged from to 2-5 years. Conclusion:This study showed that AI models, especially Bing AI and Gemini, perform at a high level in orthopedic specialty examinations and have potential as educational support tools.However, the lower accuracy of ChatGPT-4o reduced its suitability for assessment.Despite these limitations, AI shows promise in medical education.Future research should focus on improving the reliability, incorporating visual data interpretation, and exploring clinical integration.

Autoren

Enver İpek

Institutionen

Themen

Artificial Intelligence in Healthcare and EducationClinical Reasoning and Diagnostic SkillsRadiomics and Machine Learning in Medical Imaging

Volltext beim Verlag öffnen

Performance of AI Models vs. Orthopedic Residents in Turkish Specialty Training Development Exams in Orthopedics

Abstract

Ähnliche Arbeiten

Autoren

Institutionen

Themen