Pith. sign in

REVIEW 2 cited by

Evaluating multiple large language models in pediatric ophthalmology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04368 v1 pith:HWKRSGLV submitted 2023-11-07 cs.CL

classification cs.CL
keywords studentsmedicalophthalmologypediatricllmsphysicianschatgptgpt-3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

IMPORTANCE The response effectiveness of different large language models (LLMs) and various individuals, including medical students, graduate students, and practicing physicians, in pediatric ophthalmology consultations, has not been clearly established yet. OBJECTIVE Design a 100-question exam based on pediatric ophthalmology to evaluate the performance of LLMs in highly specialized scenarios and compare them with the performance of medical students and physicians at different levels. DESIGN, SETTING, AND PARTICIPANTS This survey study assessed three LLMs, namely ChatGPT (GPT-3.5), GPT-4, and PaLM2, were assessed alongside three human cohorts: medical students, postgraduate students, and attending physicians, in their ability to answer questions related to pediatric ophthalmology. It was conducted by administering questionnaires in the form of test papers through the LLM network interface, with the valuable participation of volunteers. MAIN OUTCOMES AND MEASURES Mean scores of LLM and humans on 100 multiple-choice questions, as well as the answer stability, correlation, and response confidence of each LLM. RESULTS GPT-4 performed comparably to attending physicians, while ChatGPT (GPT-3.5) and PaLM2 outperformed medical students but slightly trailed behind postgraduate students. Furthermore, GPT-4 exhibited greater stability and confidence when responding to inquiries compared to ChatGPT (GPT-3.5) and PaLM2. CONCLUSIONS AND RELEVANCE Our results underscore the potential for LLMs to provide medical assistance in pediatric ophthalmology and suggest significant capacity to guide the education of medical students.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Large language models predict retinopathy of prematurity risk poorly from admission notes alone, over-predict medium and high risk, and positive emotional prompt framing partially corrects this bias.

  2. Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios

    cs.CL 2024-11 conditional novelty 4.0 of 10

    Replacing GPT-4 with o1-preview as the backbone of CoD, MedAgents, and AgentClinic improves mean diagnostic accuracy on several medical benchmarks, with higher runtime and mixed results on simple agent roles.

Pith tools