{"id":"ea551854-8b84-4508-88ba-391f80c7cd46","arxiv_id":"2505.08414","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Meta-EyeFM uses a language model to route retinal images to specialized classifiers and achieves ophthalmologist-level accuracy for several eye diseases, though external validation is partly compromised by dataset overlap.","lead":"This paper presents Meta-EyeFM, a system that pairs a large language model with specialized eye-image classifiers to detect signs of common eye diseases from retinal photographs. The authors report that it matches the performance of an ophthalmologist on a small test set and outperforms general-purpose AI chatbots on eye-disease identification.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"External test sets ODIR-5K and SP2 appear in both pre-training (Table S1) and external testing (Table S3), so the claimed external generalizability is not measured on unseen data; the abstract's severity claim (≥89%) is also contradicted by Table S7.","rationale":"The paper's intended contribution is a generalizable, conversation-driven diagnosis and triage tool for primary eye care. The only evidence for generalizability is external testing (Table S6) and the LLM comparison (Table S9). For that evidence to support the claim, the external datasets must be independent of the pre-training corpus. Table S1 and Table S3 show ODIR-5K and SP2 in both, so the assumption fails for a substantial subset of external endpoints: cataract relies entirely on ODIR-5K, and systemic disease endpoints rely mostly on ODIR-5K and SP2. A reviewer could argue that MAE pre-training is unsupervised, but that does not restore independence: the visual encoder has been optimized to reconstruct and represent exactly those images and acquisition styles, so performance on them is not a fair estimate of performance on unseen populations or cameras. The paper itself lists no exclusion of pre-training datasets from external testing, and the overlap is directly visible in its own supplementary tables. I also note the abstract's '≥89% severity accuracy' is contradicted by Table S7 (minimum 82.5%), and the Discussion's 'AUC ≥95%' is contradicted by Table S4 (AMD 91.2%, cataract 93.9%, glaucoma 94.2%). These are internal inconsistencies in the headline numbers, not stylistic issues. The external-overlap problem is the more fundamental concern because it cannot be fixed by rephrasing; it requires redoing or reinterpreting the external validation. The proposed hash-based overlap audit plus re-run on non-overlapping datasets would settle whether the concern lands. Given the central generalizability claim is unsupported and the headline numbers are internally inconsistent, the reader's REJECT verdict remains appropriate.","tokens_in":17782,"tokens_out":8906,"duration_ms":87011,"concrete_test":"Compute perceptual hashes (e.g., pHash) for every image in the Table S1 entries for ODIR-5K and SP2 and every image in the Table S3 external test entries for ODIR-5K and SP2; where available, compare subject/patient IDs as well. Then re-run the external evaluation using only the non-overlapping, truly external datasets (e.g., UKBB for systemic disease, PIONEER/GF for glaucoma, MESSIDOR2/IDRiD/APTOS for DR, CGMH/SESA for MMD). If AUC/accuracy for cataract and systemic endpoints drops materially or cannot be computed from non-overlapping data, the reported external performance is inflated and the generalizable decision-support claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's generalizability claim rests on external testing with datasets that are supposed to be independent of model development. Table S1 lists ODIR-5K (6,377 images) and SP2 (9,139 images) as part of the 249,925-image self-supervised pre-training corpus, and Table S3 uses the same two datasets as external test sets. Because MAE pre-training directly models these images, external metrics for cataract (ODIR-5K only), diabetes and hypertension (ODIR-5K and SP2), and AMD/glaucoma (partly ODIR-5K) are not evaluations on unseen data. This violates the independence assumption of external validation and inflates the Discussion's 'robust external performance' claim. In addition, the abstract and Discussion state severity accuracy ≥89%, but Table S7 shows the minimum severity accuracy is 82.5% (early AMD) and 84.0% (mild NPDR); the Discussion's 'AUC ≥95% for major diseases' is also contradicted by internal AUCs of 91.2% (AMD), 93.9% (cataract), and 94.2% (glaucoma) in Table S4. The central claim is therefore not supported as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Meta-EyeFM, a system that combines a LoRA-fine-tuned LLaVA-based multimodal LLM with eight task-specific vision foundation models (RetFound-initialized MAEs) for fundus image analysis. The LLM is trained to route user queries and images to the appropriate VFM, which performs detection of ocular and systemic diseases, severity differentiation, and sign identification. The authors report 100% routing accuracy on an internal SEED test set, internal AUCs between 79.8% and 98.8%, external AUCs above 59%, performance 11-43% better than Gemini-1.5-flash and ChatGPT-4o on several disease detection benchmarks, and F1-based parity with an ophthalmologist on a 60-image benchmark. The central claims are that the system is an accurate conversational diagnostic and triaging tool for primary eye care and that its external performance demonstrates generalizability.","tokens_in":18018,"tokens_out":4596,"duration_ms":45136,"significance":"If the empirical claims held, the paper would be a valuable contribution to conversational AI in ophthalmology: the router design is sensible, the comparison with general-purpose LMMs addresses a practical deployment question, and the few-shot analysis provides useful evidence on data efficiency. The use of a large multi-ethnic SEED cohort with adjudicated labels is a genuine strength, and the explicit definition of referable disease classes aids clinical interpretation. However, the external generalizability claim is compromised by the use of pre-training datasets in external testing, and several abstract/discussion-level performance numbers do not match the supplementary tables. As a result, the paper's headline claims are not currently supported by the evidence as presented.","major_comments":[{"comment":"The self-supervised pre-training corpus in Table S1 includes ODIR-5K (6,377 images) and SP2 (9,139 images), and the same two datasets appear in Table S3 as external test sets. Because MAE pre-training reconstructs images directly from these datasets, the images are not unseen, so the external AUCs and accuracies for AMD, glaucoma, cataract, diabetes, and hypertension in Table S6 that draw on ODIR-5K and SP2 cannot support the Discussion's claim of 'robust external performance' or generalizability to unseen acquisition protocols. The authors should exclude all pre-trained datasets from external validation and re-run the analysis, or provide evidence such as near-duplicate detection that no image-level overlap exists.","section":"Methods, Table S1 vs Table S3"},{"comment":"The Discussion states that Meta-EyeFM detected major ocular diseases with AUC ≥95% and severity accuracy ≥89%, and the abstract repeats the ≥89% severity claim. However, Table S4 reports internal AUCs of 91.2% for AMD, 93.9% for cataract, and 94.2% for glaucoma, and only referable DR and referable MMD reach 95% or higher. Table S7 reports severity accuracy of 82.5% for early AMD and 84.0% for mild NPDR. These internal inconsistencies must be corrected in the text and abstract, or the tables must be revised to support the stated minima.","section":"Abstract and Discussion vs Tables S4 and S7"},{"comment":"The claim that Meta-EyeFM is 'comparable to an ophthalmologist' is based on a single benchmark of 60 images (68 labels) with no confidence intervals, paired hypothesis tests, or measures of inter-grader variability. The glaucoma row alone shows F1-scores ranging from 0 to 0.706 across graders, so a single F1 value for Meta-EyeFM (0.696) does not establish statistical equivalence. The authors should report a paired analysis with uncertainty or substantially soften the claim.","section":"Clinical benchmarking, Table S10"}],"minor_comments":[{"comment":"The caption describes Gemini-1.5 and GPT-4o as 'open-source LLMs', but both are proprietary API-based models; this should be corrected to 'general-purpose LMMs'.","section":"Table S9 caption"},{"comment":"The quoted prompt (3) is missing a closing quotation mark and contains a grammatical error ('presents in this image'); please fix the quotation and wording.","section":"Figure 5 footnote and Table S9 footnote"},{"comment":"The phrase 'F1-scores ≤0.814' is uninformative because 0.814 is the largest F1 value in Table S4; the authors should report the range (e.g., 0.236-0.814) instead.","section":"Results, Ocular disease detection, first paragraph"},{"comment":"The text states 'AUCs ≥81%' for referable DR and AMD, but Table S6 reports 80.7% for MESSIDOR2 (DR) and 80.5% for PIONEER (AMD); either the text or the table must be corrected.","section":"Results, External testing, first paragraph"}],"recommendation":"reject","confidential_remarks":"The dataset overlap between pre-training and external testing is a data-integrity issue that cannot be fixed with a text revision, because it invalidates a central claim of the paper. Even after correcting the internal numerical inconsistencies, the external validation would need to be redone with disjoint datasets. I therefore recommend rejection, though a future substantially revised version that addresses the leakage and reporting issues could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious engineering effort that deserves a hard look, but the abstract overstates what the tables show, and the external validation is partly contaminated by pre-training overlap. I wouldn't trust the headline numbers as written.\n\nWhat's actually new: the integration of a LLaVA-based router with several RetFound-initialized VFMs for detection, severity, and sign identification across a broad set of ocular and systemic diseases is not something I've seen in one package. The routing idea comes from LISA, and each component is known, but the combination and the breadth of tasks is a legitimate contribution. The internal test results, and the comparison against Gemini/ChatGPT-4o and human graders, give useful evidence that a routed VFM system can beat generic LMMs on fundus tasks.\n\nSoft spots, in order of importance. First, the external test sets ODIR-5K and SP2 appear in both Table S1 (pre-training) and Table S3 (external testing). Since MAE pre-training models those images, the AUCs for cataract, diabetes, hypertension, AMD and glaucoma on those datasets do not measure generalization to unseen data. That guts the 'robust external performance' claim in the Discussion. Some other external datasets (MESSIDOR2, IDRiD, APTOS, PIONEER, CGMH, SESA, UKBB) are independent, so not all external numbers are void, but the paper needs to be re-analyzed without the overlapping sets and the claims rescaled. Second, the abstract and Discussion say severity accuracy ≥89%, but Table S7 shows early AMD at 82.5% and mild NPDR at 84.0%. Similarly, the Discussion's 'AUC ≥95% for major diseases' is contradicted by internal AUCs of 91.2% for AMD and 93.9% for cataract. These are not typos in the margin; they're the central selling points. Third, the external F1-scores for some classes are very low (AMD F1 0.261 on PIONEER, cataract 0.279 on ODIR-5K), which the abstract's '≥82.2% accuracy' glosses over.\n\nNone of this makes the internal results meaningless. The routing accuracy at 100% is plausible. But the paper as written does not support the deployable-tool claim. It needs a major revision that separates truly external from pre-training-overlapping data, corrects the summary numbers, and tones down the generalizability language.\n\nWho's this for? Someone building multimodal medical AI systems will find the architecture useful. A clinician evaluating a ready-to-use triage tool should not rely on the current numbers. I'd send it to peer review because the system is substantive and the flaws are identifiable enough that a good referee could guide a fix. But I would not want it accepted in its current form.","headline":"A serious engineering integration, but the external validation is partly contaminated by pre-training overlap and the abstract's headline numbers contradict the paper's own tables.","tokens_in":18727,"tokens_out":2770,"would_cite":false,"duration_ms":25938,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI chat tool reads eye photos at ophthalmologist level","keywords":["Meta-EyeFM","fundus photography","foundation model","large language model","vision foundation model","conversational triage","primary eye care","diabetic retinopathy"],"falsifier":"Re-run the claimed external experiments after deleting every image from ODIR-5K and SP2, and any other overlapping dataset, from pretraining; if the reported external AUCs drop materially, the generalization claim is refuted. A simpler check is to inspect Table S1 and Table S3 side by side, as both list ODIR-5K and SP2.","tokens_in":17562,"feed_emoji":"👁️","tokens_out":5000,"duration_ms":46515,"temperature":0.7,"pith_summary":"Meta-EyeFM is presented as a decision-support system that lets a user type a question about a fundus photograph and get a diagnostic answer. The paper's claim is that routing each image to one of eight disease-specific vision models, rather than asking a general multimodal AI to diagnose directly, combines the image-reading strength of specialized models with the flexibility of a chatbot. If the reported results hold, primary-care staff without ophthalmology training could screen for diabetic retinopathy, age-related macular degeneration, myopic macular degeneration, cataract, and glaucoma, and also receive alerts about associated systemic conditions. The authors report that the router made the correct assignment in every tested case and that the system matched the best ophthalmologist in a small head-to-head benchmark.","feed_headline":"AI chat tool reads eye photos at ophthalmologist level","feed_subtitle":"A language router sends each fundus image to the right vision model, beating Gemini and GPT-4o by 11-43 percent.","key_machinery":"The mechanism that carries the argument is the embedding-as-router design. The large language model, initialized from LLaVA and fine-tuned with Low-Rank Adaptation, expands its vocabulary with a special <Router> token; when a user supplies a fundus image and a text query, the final-layer embedding of that token is projected through a multilayer perceptron to choose one of eight vision foundation models. Each vision foundation model begins from RetFound pretrained weights and is further pretrained as a masked autoencoder with a 75% mask ratio, then fine-tuned on the SEED cohort. The routing loss and the text-generation loss are trained jointly, which is what lets one system answer questions about disease, severity, and signs from the same image.","core_discovery":"Meta-EyeFM's central claim is that a conversational eye-care assistant can be built by connecting a large language model to eight specialized vision foundation models through a routing layer. On the internal SEED test set, the router sent every fundus image to the correct expert, and the experts reached AUCs of 97.4% for referable diabetic retinopathy, 91.2% for AMD, 98.8% for referable myopic macular degeneration, 94.2% for glaucoma, and 93.9% for visually significant cataract, with disease-detection accuracy at or above 82.2%, severity differentiation at or above 89%, and sign identification at or above 76%. On external datasets the model kept AUCs of at least 81% for retinal disease detection and outperformed Gemini-1.5-flash and ChatGPT-4o by 11 to 43 percent depending on disease. In a 60-image benchmark against clinicians, Meta-EyeFM's F1 score (0.853) was closest to the best ophthalmologist (0.857) and above the other graders. The authors position the system as decision support for primary care and online fundus evaluation, while noting that further diagnostic improvement is needed before screening use.","pith_inferences":["Because ODIR-5K and SP2 appear in both the pretraining list (Table S1) and the external test list (Table S3), the external accuracy figures should be treated as likely optimistic for truly unseen data until the model is retested on datasets excluded from pretraining.","A natural next test is to run the same routing mechanism on OCT or anterior-segment photos, since the design is not eye-specific; if it transfers, the same architecture could become a general medical-image triage shell.","An extension that follows directly from the design is to measure whether the conversational interface changes real referral decisions, not just label accuracy, in a prospective primary-care workflow.","The routing accuracy was measured only on three query types; testing with free-form, out-of-distribution patient language would show whether the 100% routing figure survives real-world phrasing."],"forward_implications":["A single system can handle disease detection, severity grading, and sign identification from one fundus photo through natural-language questions, removing the need for separate apps per task.","The routing design lets a general-purpose LLM handle conversation while specialized vision models do the image diagnosis, sidestepping the known weakness of LLMs on visual medical tasks.","Fine-tuning with only 10% of the training data still kept internal AUCs within a few points of full-data training, supporting data-efficient expansion to rarer diseases and underrepresented groups.","Sign identification (drusen, exudates, microaneurysms) gives clinicians a visible basis for the diagnosis, which the authors argue supports documentation and helps right-site referrals.","Systemic disease prediction from fundus photos opens a non-invasive screening channel for diabetes, hypertension, and chronic kidney disease, although the authors state that performance must improve before real use."],"supporting_citations":[{"why":"Supplies the RetFound pretrained weights and the retinal-image foundation-model approach that the vision experts are built on.","marker":"14"},{"why":"Supplies the Low-Rank Adaptation method used for parameter-efficient fine-tuning of the language model.","marker":"19"},{"why":"Supplies the LLaVA multimodal LLM architecture that is fine-tuned into the router.","marker":"20"},{"why":"Supplies the embedding-as-router paradigm behind the <Router> token routing mechanism.","marker":"21"},{"why":"Supplies the masked autoencoder pretraining objective used for the vision foundation models.","marker":"22"},{"why":"Supplies the SEED cohort profile, the source of the fine-tuning and internal test images with adjudicated labels.","marker":"18"}],"fun_headline_variants":["Conversational AI triages eye scans at ophthalmologist level","Language router sends eye photos to right vision experts","Eye AI beats Gemini and GPT-4o by up to 43 percent","Chat tool matches best ophthalmologist on eye photo diagnosis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The external test results measure generalization only if the external datasets were not part of the self-supervised pretraining corpus; ODIR-5K and SP2 appear in both tables, so that independence is not established.","fun_headline_variants_meta":{"raw":{"variants":["Conversational AI triages eye scans at ophthalmologist level","Language router sends eye photos to right vision experts","Eye AI beats Gemini and GPT-4o by up to 43 percent","Chat tool matches best ophthalmologist on eye photo diagnosis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1492,"prompt_tokens":1002,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":618,"tokens_out":490,"duration_ms":4553,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:55:03.225125+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the claimed external experiments after deleting every image from ODIR-5K and SP2, and any other overlapping dataset, from pretraining; if the reported external AUCs drop materially, the generalization claim is refuted. A simpler check is to inspect Table S1 and Table S3 side by side, as both list ODIR-5K and SP2.","supporting_citations":[{"cited_title":"Levy PI, New York, NY 10029, United States","cited_arxiv_id":null,"evidence_quote":"Supplies the RetFound pretrained weights and the retinal-image foundation-model approach that the vision experts are built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Low-Rank Adaptation method used for parameter-efficient fine-tuning of the language model."},{"cited_title":"The Lancet global health Commission on global eye health: vision beyond 2020","cited_arxiv_id":null,"evidence_quote":"Supplies the LLaVA multimodal LLM architecture that is fine-tuned into the router."},{"cited_title":"Prevalence of undiagnosed age-related macular degeneration in primary eye care","cited_arxiv_id":null,"evidence_quote":"Supplies the embedding-as-router paradigm behind the <Router> token routing mechanism."},{"cited_title":"The global extent of undetected glaucoma in adults: a systematic review and meta-analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the masked autoencoder pretraining objective used for the vision foundation models."},{"cited_title":"Could you identify the specific disease present in this fundus image?","cited_arxiv_id":null,"evidence_quote":"Supplies the SEED cohort profile, the source of the fine-tuning and internal test images with adjudicated labels."}],"review_version":1}