{"id":"980abca0-8c0c-4bcc-865d-c2ebed946c19","arxiv_id":"2507.19498","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM-based myopia education agent with image grading and retrieval-augmented knowledge matched specialists on exams and improved patient satisfaction in a 70-person randomized trial.","lead":"ChatMyopia, an AI agent that pairs a large language model with an image grader and a curated medical knowledge base, answers myopia-related questions before patients see their eye care practitioner. In a randomized trial with 70 patients, it improved self-reported satisfaction compared with printed leaflets, though the study was single-blinded and the agent's code is not released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RCT satisfaction gain may be caused by interactive-format novelty and extra ECP attention, not by ChatMyopia content; a digital-control arm is required.","rationale":"The reader's weakest assumption correctly identifies the single-blind design and novelty effects as the key threat. I agree and sharpen it by pointing to a specific mechanism in the Methods: ECPs in the intervention arm monitored ChatMyopia's responses and could adapt their explanations, creating a differential-attention confound beyond mere novelty. This threat is load-bearing because the paper's central claim—improved patient satisfaction from ChatMyopia—rests entirely on this one small RCT; the SCQ and QA evaluations support capability but not real-world effectiveness. The authors' own limitations section mentions single-blinding but does not propose or test a control for the digital format or ECP attention. A three-arm trial with a digital-leaflet control and blinded or standardized ECP interaction would settle the question. Since the reader already judged the paper conditional and our concern matches, the verdict should remain unchanged.","tokens_in":11101,"tokens_out":4093,"duration_ms":41740,"concrete_test":"Add a third arm to the RCT: a 'digital leaflet' presented on the same tablet with the same static myopia information but no AI interactivity. Keep ECPs blinded to allocation (e.g., by not showing them the tool interface) or standardize the consultation script, and log consultation duration and the number of ECP clarifications. If the digital-leaflet arm's C-MISS-R scores are statistically indistinguishable from ChatMyopia's, the current effect is attributable to format or novelty; if ChatMyopia still significantly outperforms the digital leaflet, the content-based satisfaction claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the internal validity of the RCT's primary outcome. In the Methods, the intervention group spent 10 minutes interacting with ChatMyopia on a tablet, while the control group read paper leaflets for 10 minutes. Participants were unblinded, and ECPs in the intervention arm 'monitored ChatMyopia's responses and addressed any areas requiring further clarification.' This creates a confound: the intervention differs from control not only in educational content but also in delivery format, the novelty of using an AI tool, and the amount or quality of ECP attention during the consultation. The primary endpoint is the self-reported C-MISS-R satisfaction scale, with no objective outcomes. The observed p=0.018 difference could therefore be fully explained by format effects, novelty effects, or differential ECP behavior rather than by the myopia-specific content of ChatMyopia. The authors acknowledge the single-blind design in their Limitations section but do not address the ECP-monitoring confound or provide a control for the digital format. Thus the central claim that ChatMyopia improves patient satisfaction is conditional on ruling out these non-content mechanisms.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes ChatMyopia, an LLM-based agent that combines a myopic maculopathy image classifier with a retrieval-augmented knowledge base to answer myopia-related text and image queries. The authors evaluate the system on an image grading task, national-exam single-choice questions against eye care practitioners, open-ended patient questions against GPT-4 and ECPs, and a randomized trial (n=70) comparing patient satisfaction (C-MISS-R) with traditional leaflets. They report significantly higher satisfaction in the ChatMyopia arm and comparable or superior accuracy to general ECPs.","tokens_in":11240,"tokens_out":5710,"duration_ms":54296,"significance":"If the effectiveness claim is valid, ChatMyopia would be a useful supplement for pre-consultation education in primary eye care. The work has notable strengths: the RCT is pre-registered, the primary outcome uses a validated instrument with a sample size calculation from a pilot, the SCQ questions come from external national exam materials, and the open-ended evaluation uses blinded specialist raters. The RAG-based architecture with a dynamically invoked image tool is pragmatic and the interpretability argument (heatmaps, traceable tool calls) is reasonable. However, the internal validity of the RCT—the paper's central empirical claim—is undermined by an uncontrolled comparison: the intervention differs from control in delivery format, novelty, and ECP attention. This makes the size of the content-specific effect uncertain.","major_comments":[{"comment":"The intervention group's 10-minute tablet interaction with ChatMyopia is compared to a paper-leaflet control, and ECPs in the intervention arm 'monitored ChatMyopia's responses and addressed any areas requiring further clarification.' This means the comparison conflates educational content with delivery format, novelty of an interactive AI tool, and differential ECP attention. Because the primary outcome is self-reported satisfaction, the observed p=0.018 could be explained by these non-content mechanisms. The Limitations section acknowledges single-blinding but does not address the ECP-monitoring asymmetry or the lack of a digital-control arm. The central claim that ChatMyopia improves satisfaction requires either a control that uses an equally interactive non-AI digital tool or measurement/statistical adjustment for ECP attention and consultation duration.","section":"Methods: Randomized controlled trial for real-world validation; Fig. 4"},{"comment":"The randomization is described only as 'simple random sampling' with no mention of sequence generation, allocation concealment, or who performed randomization. The trial is single-blinded, but the paper does not state whether the outcome assessors who collected the C-MISS-R were blinded to allocation. Given that the primary outcome is a subjective self-report, these omissions leave room for bias and should be reported per CONSORT guidelines. Without this information, the reader cannot fully assess the risk of bias in the primary result.","section":"Methods: Randomized controlled trial for real-world validation; Statistical analysis"}],"minor_comments":[{"comment":"The sentence 'Post hoc comparisons indicated that ChatMyopia outperformed general ECPs (C-MISS-R score = 80 vs. 67.07, p = 0.029)' uses 'C-MISS-R score' for what should be the SCQ total score; this is a labeling error that will confuse readers.","section":"Results: Performance in SCQ examination"},{"comment":"The exact versions of Mistral (e.g., Mistral Large 123B vs. Mixtral) and GPT-4 (e.g., gpt-4-0613) are not specified; the authors should include model identifiers and access dates for reproducibility.","section":"Methods: Architecture of the ChatMyopia AI agent; Performance evaluation"},{"comment":"The description 'simple random sampling' should be elaborated with the random sequence generation method and allocation concealment mechanism, in line with CONSORT reporting standards.","section":"Methods: Randomized controlled trial for real-world validation"},{"comment":"The phrase 'we utilized two public datasets (MMAC, HPMI, and our private dataset)' should read 'two public datasets and one private dataset' to be grammatically and logically correct.","section":"Methods: Establishment of tool modules"},{"comment":"The caption does not indicate what the error bars or whiskers represent in panels B and C; this information should be added.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be an incremental step beyond the group's earlier EyeGPT work, but the RCT, if strengthened with a proper digital-control arm or adequate adjustment for the ECP-attention confound, could make a valuable contribution. The main risk is that the current design does not isolate the content effect, and the manuscript's own limitation statement does not fully acknowledge this."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you care about AI agents in patient education. The real contribution is bundling a fine-tuned ViT myopic maculopathy grader with a myopia-specific RAG knowledge base inside an LLM agent, then taking that system into a pre-registered RCT (n=70) against paper leaflets. That integration is new for myopia, and the external evaluations are solid: the grader gets AUROC 0.967, ChatMyopia matches specialists on 150 national-exam SCQs, and beats GPT-4 on safety in 85 open-ended questions. The RCT's primary endpoint, C-MISS-R satisfaction, is significant (p=0.018) and the statistical analysis is standard. The self-citations for EyeFound weights and the adapted evaluation rubric are legitimate; I don't see circularity.\n\nThe soft spot is exactly where the stress-test points: the RCT cannot separate content from delivery. The intervention arm got a 10-minute tablet interaction with an AI agent; the control arm got 10 minutes of paper leaflets. In the intervention arm, ECPs also monitored ChatMyopia's responses during the consultation. So the satisfaction gain could come from the novelty of the tablet, the extra interactivity, or the extra ECP attention, all separate from the myopia-specific content. The authors acknowledge the single-blind design but never address the ECP-monitoring confound or include a digital-leaflet control. That makes the abstract's phrasing too strong. The trial supports 'an interactive tablet tool with clinician oversight improved satisfaction relative to paper leaflets,' not 'ChatMyopia's educational content caused the improvement.' Minor issues: the text accidentally labels an SCQ score as 'C-MISS-R score' in the results, and no code, demo, or data are released, which limits reproducibility.\n\nIf you read this paper, you get a decent example of an LLM agent with tool use and a transparent evaluation pipeline. It deserved a serious referee, and with a revision that either adds a digital control arm or tones down the causal language, it could be a useful citable case study. I'd send it for peer review, but I would not take the RCT as evidence that content alone moved satisfaction.","headline":"Useful integrated myopia-education agent with an honest RCT, but the satisfaction result is confounded by format and clinician attention; read it as 'tablet tool beats leaflet,' not 'ChatMyopia content beats leaflet.'","tokens_in":11835,"tokens_out":2556,"would_cite":true,"duration_ms":24902,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatMyopia, an LLM agent that answers myopia questions and grades fundus images, improved patient satisfaction in a 70-patient randomized trial and matched specialists on standardized exams.","keywords":["large language model","AI agent","myopia","patient education","randomized controlled trial","retrieval-augmented generation","myopic maculopathy grading","eye care"],"falsifier":"A randomized trial comparing ChatMyopia not with a leaflet but with an equally interactive tablet application that presents the same educational content without AI (for example, a fixed set of answers and images) would directly test the content-attribution assumption; if satisfaction is equal in the two interactive arms, the paper's claim that the AI agent specifically improves satisfaction would be refuted.","tokens_in":10883,"feed_emoji":"👁️","tokens_out":6106,"duration_ms":52229,"temperature":0.7,"pith_summary":"ChatMyopia is an AI agent built for pre-consultation education about myopia. The paper claims that in a randomized trial with 70 patients, using ChatMyopia on a tablet for ten minutes before seeing an eye care practitioner raised patient satisfaction (measured by the C-MISS-R scale) compared with traditional leaflets, with the difference significant at p=0.018. The paper also reports that ChatMyopia scored comparably to myopia specialists on standardized single-choice examinations, outperformed general eye care practitioners on knowledge and scenario questions, and produced more contextually appropriate, safer answers than GPT-4 on 85 common myopia questions. If these results hold, they show that a tool-based LLM agent with an image classifier and a retrieval-augmented knowledge base can serve as a scalable, interpretable supplement to human counseling in primary eye care.","feed_headline":"AI myopia tutor beats leaflets in patient trial","feed_subtitle":"ChatMyopia raised satisfaction scores and matched specialists on myopia exams in a randomized eye-care study.","key_machinery":"The load-bearing mechanism is the agent architecture itself: a large language model (Mistral 123B) acts as the 'brain' that interprets the patient's question, plans a task, and calls one of two specialized tools. The first tool is a ViT-large image classifier, a Vision Transformer with 24 layers, initialized with EyeFound pretrained weights and fine-tuned on 2,769 fundus images, that grades myopic maculopathy into five categories following the META-PM classification system. The second tool is a retrieval-augmented generation (RAG) pipeline that encodes queries with multilingual embeddings, searches a knowledge base built from 12 ophthalmology textbooks and 61 clinical guidelines via FAISS cosine similarity, and inserts the retrieved chunks into the prompt. The agent also appends follow-up questions to foster dialogue. What makes this architecture central is that instead of fine-tuning the LLM, the authors add a deterministic, inspectable tool layer, which they argue is cheaper to update and easier to verify, and which they credit with the accuracy and safety gains over a general-purpose LLM.","core_discovery":"The central discovery is that an LLM-based agent can be assembled for a specific ophthalmic domain and can perform for patient education at a level comparable to, and in some dimensions better than, human practitioners and general-purpose LLMs. The paper demonstrates this in three evaluations: myopic maculopathy grading with an overall AUROC of 0.967 and accuracy of 0.934; a 150-question standardized exam on which ChatMyopia scored 80 vs. 67.07 for general eye care practitioners and 78.67 for specialists; and a blinded human evaluation of 85 open-ended questions on which ChatMyopia scored significantly higher than GPT-4 overall (p<0.001) and comparably to eye care practitioners on utility, relevance, safety, and harmlessness. The randomized controlled trial then shows a significant satisfaction benefit in a real clinic, with the cognitive subscale of the C-MISS-R showing the largest improvement (p=0.013). The authors interpret these results as evidence that a transparent, tool-augmented agent can reduce the information gap before consultation without requiring a new foundation model.","pith_inferences":["Editorial inference: We infer that the satisfaction benefit would shrink if the control group received an equally interactive, non-AI tablet app; the paper's single-blinded design with a leaflet control cannot rule this out.","Editorial inference: We infer that the agent's reported weak spots—red light therapy and refractive-surgery pre-operative questions—are exactly where the static knowledge base lags the current literature, so a version with scheduled knowledge refreshes would likely close that gap.","Editorial inference: We infer that the same tool-scheduling architecture generalizes to other ophthalmic imaging tasks, such as OCT or fluorescein angiography, since the interface between the planner and tools is model-agnostic and the paper already notes this extension.","Editorial inference: We infer that a cost-effectiveness trial measuring consultation time and referral rates is needed to know whether the ten-minute AI interaction is a net time saver for clinics; the paper measured satisfaction only."],"forward_implications":["In primary eye care settings, a 10-minute pre-consultation interaction with ChatMyopia can raise patient satisfaction and perceived understanding of eye conditions relative to printed leaflets, with both the cognitive and affective subscales of the C-MISS-R showing gains.","An AI agent that uses retrieval-augmented knowledge plus an image-grade tool can match specialist-level performance on standardized myopia knowledge questions, suggesting that domain-specific agents can substitute for task-specific fine-tuned models.","For common myopia questions, the agent's answers are rated safer and more contextually appropriate than GPT-4's, indicating that general-purpose LLMs carry avoidable clinical risk in this setting.","The system's interpretable pipeline—each diagnostic or textual answer can be traced to a retrieved source or a heatmap—makes it possible for clinicians to review and correct agent output, which the paper argues is a precondition for clinical deployment.","The trial did not find a significant reduction in decision conflict, so the agent's satisfaction benefit is not automatically a decision-support benefit."],"supporting_citations":[{"why":"Supplies the public MMAC fundus image dataset used to train and test the myopic maculopathy classifier.","marker":"[21]"},{"why":"Defines the META-PM international classification system that the image tool grades.","marker":"[25]"},{"why":"Provides the EyeFound pretrained multimodal ophthalmic weights that initialize the ViT-large image model.","marker":"[26]"},{"why":"Supplies the Mistral 123B LLM that acts as the agent's core planner and responder.","marker":"[28]"},{"why":"Provides a prior retrieval-augmented ophthalmology LLM framework and the human evaluation criteria adapted for patient-centered question-answering.","marker":"[29]"},{"why":"Is the validated Chinese version of the Medical Interview Satisfaction Scale-Revised used as the trial's primary outcome.","marker":"[30]"},{"why":"Is the original Medical Interview Satisfaction Scale on which C-MISS-R is based.","marker":"[31]"},{"why":"Is the Decision Conflict Scale used to measure the secondary outcome of decisional conflict.","marker":"[32]"},{"why":"Prior EyeGPT study that supplies adapted evaluation criteria and the comparison showing RAG equals fine-tuning.","marker":"[14]"},{"why":"Benchmarks general LLMs on myopia care, providing the comparison point for why a domain-specific agent is needed.","marker":"[8]"}],"fun_headline_variants":["AI myopia tutor beats leaflets, tops specialists on exam","ChatMyopia: AI agent scores 80, beats GPT-4 in blind test","Myopia AI tutor improves patient satisfaction in RCT","AI eye tutor: 0.967 AUROC on grading, beats leaflets","LLM agent for myopia: better than leaflets, matches specialists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The observed satisfaction improvement is assumed to come from ChatMyopia's educational content rather than from the novelty of using a tablet, the extra attention patients get when an AI tool is present, or other non-content effects, because the trial was single-blinded and used only a static leaflet as the control.","fun_headline_variants_meta":{"raw":{"variants":["AI myopia tutor beats leaflets, tops specialists on exam","ChatMyopia: AI agent scores 80, beats GPT-4 in blind test","Myopia AI tutor improves patient satisfaction in RCT","AI eye tutor: 0.967 AUROC on grading, beats leaflets","LLM agent for myopia: better than leaflets, matches specialists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2581,"prompt_tokens":957,"completion_tokens":1624,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":1533}},"tokens_in":573,"tokens_out":1624,"duration_ms":15508,"temperature":1.0,"reasoning_tokens":1533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T06:01:02.767789+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized trial comparing ChatMyopia not with a leaflet but with an equally interactive tablet application that presents the same educational content without AI (for example, a fixed set of answers and images) would directly test the content-attribution assumption; if satisfaction is equal in the two interactive arms, the paper's claim that the AI agent specifically improves satisfaction would be refuted.","supporting_citations":[{"cited_title":"A Competition for the Diagnosis of Myopic Maculopathy by Artificial Intelligence Algorithms","cited_arxiv_id":null,"evidence_quote":"Supplies the public MMAC fundus image dataset used to train and test the myopic maculopathy classifier."},{"cited_title":"International Photographic Classification and Grading System for Myopic Maculopathy","cited_arxiv_id":null,"evidence_quote":"Defines the META-PM international classification system that the image tool grades."},{"cited_title":"EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging","cited_arxiv_id":null,"evidence_quote":"Provides the EyeFound pretrained multimodal ophthalmic weights that initialize the ViT-large image model."},{"cited_title":"An Empirical Evaluation of Large Language Models on Consumer Health Questions","cited_arxiv_id":null,"evidence_quote":"Supplies the Mistral 123B LLM that acts as the agent's core planner and responder."},{"cited_title":"Development and evaluation of a retrieval-augmented large language model framework for ophthalmology","cited_arxiv_id":null,"evidence_quote":"Provides a prior retrieval-augmented ophthalmology LLM framework and the human evaluation criteria adapted for patient-centered question-answering."},{"cited_title":"Brief communication: The Chinese medical interview satisfaction scale-revised (C-MISS-R): Development and validation","cited_arxiv_id":null,"evidence_quote":"Is the validated Chinese version of the Medical Interview Satisfaction Scale-Revised used as the trial's primary outcome."},{"cited_title":"The medical interview satisfaction scale: Development of a scale to measure patient perceptions of physician behavior","cited_arxiv_id":null,"evidence_quote":"Is the original Medical Interview Satisfaction Scale on which C-MISS-R is based."},{"cited_title":"O’ Connor, Ottawa Hospital Research Institute","cited_arxiv_id":null,"evidence_quote":"Is the Decision Conflict Scale used to measure the secondary outcome of decisional conflict."},{"cited_title":"EyeGPT for Patient Inquiries and Medical Education: Development and Validation of an Ophthalmology Large Language Model","cited_arxiv_id":null,"evidence_quote":"Prior EyeGPT study that supplies adapted evaluation criteria and the comparison showing RAG equals fine-tuning."},{"cited_title":"Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard","cited_arxiv_id":null,"evidence_quote":"Benchmarks general LLMs on myopia care, providing the comparison point for why a domain-specific agent is needed."}],"review_version":1}