RESERVED CONTENT
Registration is now closed!
Topic: HISTORY AND EDUCATION - HISTORY AND EDUCATION - Education
Victor Leite, MD1, Leonardo Santana Neiva Filho, AA2, Arthur Moreira Meirelles, Student3, Alice Moreira Meirelles, Student4, Idrys Henrique Leite Guedes, MD5, Aline Moreira Meirelles Santana Neiva, MD6, Abdon Moreira Lustosa, MD, PhD6, Marcia Adriana Dias Meirelles Moreira, MD, PhD6.
1Faculdade de Medicina Nova Esperanca, JOAO PESSOA, Brazil, 2Liberty University, Lynchburg, VA, USA, 3Centro Universitario de Joao Pessoa, Joao Pessoa, Brazil, 4Faculdade de Medicina Nova Esperanca, Joao Pessoa, Brazil, 5Universidade Federal de Campina Grande, Campina Grande, Brazil, 6Universidade Federal da Paraiba, Joao Pessoa, Brazil.
Background:
The increasing application of artificial intelligence (AI) in medical education and assessment has raised questions about its reliability in high-stakes medical examinations. Large language models (LLMs) like ChatGPT have demonstrated remarkable capabilities in processing and generating text, but their effectiveness in specialized medical fields remains uncertain. This study aims to evaluate the performance of ChatGPT in the Brazilian Society of Anesthesiology (SBA) specialist examination, analyzing its accuracy across different question difficulty levels, cognitive domains, and anesthesiology topics.
Methods:
A total of 120 multiple-choice questions (MCQs) from SBA examinations were presented to ChatGPT. These questions were classified based on their difficulty level (very easy, easy, medium, difficult, very difficult) and cognitive domain (remember, understand, apply). The performance of ChatGPT was compared across SBA examination levels and anesthesiology topics. Statistical analyses included chi-square and Fisher’s exact tests, with significance set at 5%. Additionally, ChatGPT’s performance was compared with historical pass rates of human candidates to contextualize its accuracy within real-world exam conditions.
Results:
ChatGPT achieved an overall accuracy rate of 84.2% (101/120), with significant variation across difficulty levels. The model correctly answered 97.8% of very easy questions, 87.5% of easy questions, 78.6% of medium questions, and 70.0% of difficult questions, but it failed all very difficult questions (0%). The decline in accuracy as complexity increased suggests that ChatGPT excels in factual recall and moderate problem-solving but struggles with highly complex scenarios requiring deeper clinical reasoning. In terms of cognitive domains, ChatGPT demonstrated superior accuracy in questions requiring understanding (100%) compared to recall (75%) and application (76.7%) (p < 0.001). This result indicates that the model is highly effective in synthesizing and interpreting knowledge but may face limitations in applying it to clinical case-based decision-making. When analyzed by anesthesiology topics, no significant difference in performance was observed (p = 0.169), suggesting a uniform distribution of accuracy across different subject areas. Similarly, no statistical difference was found across different SBA examination levels (p = 0.368), reinforcing the model’s consistent performance regardless of exam stage. Compared to human candidates, ChatGPT’s accuracy exceeded the historical passing threshold for the SBA specialist exam, which typically ranges from 65% to 75%, highlighting its potential as a supplementary tool for exam preparation.
Conclusion:
ChatGPT demonstrated high accuracy in answering SBA anesthesiology exam questions, excelling in comprehension-based tasks while showing limitations in complex problem-solving. Its performance exceeded the minimum passing criteria for human candidates, suggesting that LLMs can be valuable study aids for anesthesiology training. However, the observed decline in accuracy with increasing question difficulty underscores the need for human oversight when interpreting AI-generated responses. Further research is necessary to assess the practical integration of AI tools in medical education and professional certification processes.
Topic: HISTORY AND EDUCATION - HISTORY AND EDUCATION - Education
Victor Leite, MD1, Leonardo Santana Neiva Filho, AA2, Arthur Moreira Meirelles, Student3, Alice Moreira Meirelles, Student4, Idrys Henrique Leite Guedes, MD5, Aline Moreira Meirelles Santana Neiva, MD6, Abdon Moreira Lustosa, MD, PhD6, Marcia Adriana Dias Meirelles Moreira, MD, PhD6.
1Faculdade de Medicina Nova Esperanca, JOAO PESSOA, Brazil, 2Liberty University, Lynchburg, VA, USA, 3Centro Universitario de Joao Pessoa, Joao Pessoa, Brazil, 4Faculdade de Medicina Nova Esperanca, Joao Pessoa, Brazil, 5Universidade Federal de Campina Grande, Campina Grande, Brazil, 6Universidade Federal da Paraiba, Joao Pessoa, Brazil.
Background:
The increasing application of artificial intelligence (AI) in medical education and assessment has raised questions about its reliability in high-stakes medical examinations. Large language models (LLMs) like ChatGPT have demonstrated remarkable capabilities in processing and generating text, but their effectiveness in specialized medical fields remains uncertain. This study aims to evaluate the performance of ChatGPT in the Brazilian Society of Anesthesiology (SBA) specialist examination, analyzing its accuracy across different question difficulty levels, cognitive domains, and anesthesiology topics.
Methods:
A total of 120 multiple-choice questions (MCQs) from SBA examinations were presented to ChatGPT. These questions were classified based on their difficulty level (very easy, easy, medium, difficult, very difficult) and cognitive domain (remember, understand, apply). The performance of ChatGPT was compared across SBA examination levels and anesthesiology topics. Statistical analyses included chi-square and Fisher’s exact tests, with significance set at 5%. Additionally, ChatGPT’s performance was compared with historical pass rates of human candidates to contextualize its accuracy within real-world exam conditions.
Results:
ChatGPT achieved an overall accuracy rate of 84.2% (101/120), with significant variation across difficulty levels. The model correctly answered 97.8% of very easy questions, 87.5% of easy questions, 78.6% of medium questions, and 70.0% of difficult questions, but it failed all very difficult questions (0%). The decline in accuracy as complexity increased suggests that ChatGPT excels in factual recall and moderate problem-solving but struggles with highly complex scenarios requiring deeper clinical reasoning. In terms of cognitive domains, ChatGPT demonstrated superior accuracy in questions requiring understanding (100%) compared to recall (75%) and application (76.7%) (p < 0.001). This result indicates that the model is highly effective in synthesizing and interpreting knowledge but may face limitations in applying it to clinical case-based decision-making. When analyzed by anesthesiology topics, no significant difference in performance was observed (p = 0.169), suggesting a uniform distribution of accuracy across different subject areas. Similarly, no statistical difference was found across different SBA examination levels (p = 0.368), reinforcing the model’s consistent performance regardless of exam stage. Compared to human candidates, ChatGPT’s accuracy exceeded the historical passing threshold for the SBA specialist exam, which typically ranges from 65% to 75%, highlighting its potential as a supplementary tool for exam preparation.
Conclusion:
ChatGPT demonstrated high accuracy in answering SBA anesthesiology exam questions, excelling in comprehension-based tasks while showing limitations in complex problem-solving. Its performance exceeded the minimum passing criteria for human candidates, suggesting that LLMs can be valuable study aids for anesthesiology training. However, the observed decline in accuracy with increasing question difficulty underscores the need for human oversight when interpreting AI-generated responses. Further research is necessary to assess the practical integration of AI tools in medical education and professional certification processes.

