REVIEW 3 major objections 5 minor 1 cited by
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read LLM eye-care answers are systematically worse in LMIC languages, and a new agentic pipeline narrows the gap.
desk verdict A genuinely new paired multilingual ophthalmology benchmark and a sensible debiasing pipeline, but the translation-equivalence assumption is unvalidated, so the exact gaps and CLARA's gains are provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
CLARA (Cross-Lingual Reflective Agentic system) is the load-bearing mechanism, a multi-agent inference-time pipeline. A translation agent converts the query to English; an evaluation agent judges translation quality and the model's certainty; a knowledge agent runs weighted retrieval over biomedical abstracts, medical textbooks, and general encyclopedic articles, reweighting relevance scores by confidence in each part of the question and by expanded ophthalmology jargon; a second evaluation agent critiques retrieved documents and can trigger a web search; and a rewriting agent decomposes complex queries before another pass. The pipeline is the mechanism that turns the paper's diagnosis of language bias into a correction.
What would settle it
Have independent professional translators back-translate all 1,184 questions from each language and have ophthalmologists rate equivalence of meaning, difficulty, and cultural neutrality. If the languages with the largest measured gaps (Filipino, Hindi, Mandarin) also show the largest translation-equivalence problems, the central bias claim would lose support; if equivalence is high, the claim is strengthened.
Extended reading notes
Core claim
The paper's central discovery is that cross-lingual bias in ophthalmological question answering is systematic: every model evaluated scores lower in Filipino, Hindi, and Mandarin than in English, and the worst gaps appear in precisely the languages most relevant to low- and middle-income countries. The authors trace the failures to three causes—limited language proficiency, shallow ophthalmology-specific knowledge, and difficulty with linguistic nuance and cultural context—and document that existing fixes (English chain-of-thought, pre-translation, web search, plain RAG) are inconsistent. CLARA, the proposed system, combines translation, weighted corrective retrieval, web search, iterative relevance verification, and query rewriting; it improves accuracy in every tested language and cuts the largest LMIC gaps roughly in half. The paired question design is what makes these comparisons interpretable, since the same question content is evaluated across languages.
Load-bearing premise
The benchmark's cross-lingual comparisons assume the manually translated questions preserve meaning, difficulty, and cultural neutrality across all seven languages.
Editorial extensions
If this is right
- Deploying LLMs for ophthalmology triage, patient education, or documentation in Filipino, Hindi, or Mandarin without debiasing would systematically under-serve speakers of those languages relative to English speakers.
- CLARA's accuracy gains of roughly 10–15 percentage points in lower-resourced languages can be achieved at inference time, without any fine-tuning, making equitable multilingual deployment more feasible.
- None of the standard single-component fixes—translation, chain-of-thought, web search, or basic RAG—closes the gap on its own; the ablation shows each CLARA component adds a small but consistent gain.
- The released paired benchmark gives other groups a reusable protocol for auditing new models for cross-lingual medical bias rather than relying on English-only evaluations.
Reading between the lines
- I would predict that most of CLARA's benefit comes from the translation step plus English-centric retrieval; isolating that would require an oracle condition in which the model receives perfect English translations with no retrieval.
- The same paired-question methodology could be extended to other specialties with high LMIC burden, such as obstetrics or tropical medicine, to expose analogous language gaps.
- Because the paper documents expert review but no back-translation or inter-annotator equivalence scoring, some portion of the measured gap may reflect translation artifacts; a human equivalence audit of the 1,184 items would settle that.
- It remains untested whether CLARA's gains survive domain shift to new questions, other dialects, or spoken-language triage; running it on an independent multilingual medical exam set would be the next check.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Multi-OphthaLingua, a parallel multilingual ophthalmology multiple-choice benchmark in seven languages (English, Spanish, Portuguese, Filipino, Mandarin, Hindi, French), evaluates six LLMs, and proposes CLARA, an inference-time pipeline that combines translation, weighted RAG, web search, and query rewriting. The main empirical claims are that LLM accuracy is systematically lower in LMIC-representative languages such as Filipino, Hindi, and Mandarin, and that CLARA improves absolute accuracy in all seven languages while narrowing the gap relative to English, with the largest reported effect for GPT-4 in Filipino (51.8% direct to 67.1% with CLARA; gap reduced from 11.6 to 5.1 points, Table 3).
Significance. If the benchmark is valid, it is a useful contribution: it appears to be the first paired multilingual ophthalmology QA dataset, it is expert-curated, a sample is publicly available, and the evaluation covers six models and seven languages. The CLARA system is compared against Translate-COT, Web-ToolCall, and an ablation ladder, so the reported gains are not simply fit to the answer key, and the qualitative observations about ambiguous terms such as 'namamaga' and 'mancha' are informative. However, the central quantitative comparisons currently rest on unvalidated translation equivalence and unreported variance; the significance of the headline claims cannot be assessed until those are addressed.
major comments (3)
- [Benchmark Construction] The cross-lingual comparisons are load-bearing but translation equivalence is asserted, not demonstrated. The text states that questions were 'originally written in Portuguese' and 'manually translated' into six languages, then 'reviewed and curated by board-certified native-speaker ophthalmologists,' but no back-translation, inter-annotator agreement, or per-language difficulty calibration is reported. The paper itself shows that ambiguous terms exist in the benchmark ('namamaga' in Filipino, 'mancha' in Portuguese), so without evidence that all seven versions have equivalent difficulty and meaning, the language accuracy gaps in Tables 2 and 3, and the gap reductions attributed to CLARA, could reflect translation artifacts rather than model bias. Please provide translation-equivalence checks, per-language item statistics, and a protocol for verifying cultural neutrality, or explicitly report these as limitations.
- [Benchmark Results and LLM-Failure Analysis, Tables 2 and 3] The quantitative results are reported only as 8-run averages without standard deviations, confidence intervals, or significance tests. Claims such as 'significant disparities' and 'significantly reduces the multilingual bias gap' are therefore not statistically supported. Additionally, no item counts are given for the per-language or per-subgroup cells; with 1184 questions split across at least ten subgroups and seven languages, some cells are likely small, and accuracy values such as 65.9% versus 70.5% in the Basic-Sciences rows may not be distinguishable. Please report per-cell sample sizes, variance across runs, and appropriate tests or confidence intervals for the gap comparisons.
- [CLARA, Eqs. (1) and (2)] The method description leaves free parameters unspecified: the query-part weights w_j in Eq. (1), the jargon weights w_k in Eq. (2), the maximum iteration limit (given as 5), and the evaluation thresholds for translation certainty and retrieval relevance. It is not stated how these weights are set, whether they were tuned on the benchmark, or whether they are fixed ahead of time. This affects reproducibility and raises the risk of optimistic results if any component was selected on the test set. Please specify the parameter-setting procedure and clarify whether any development data or validation split was used.
minor comments (5)
- [Tables 2 and 3] The Hindi column is labeled 'HI' in Table 2 and 'HIN' in Table 3; please unify the notation.
- [Experimental Setup] The text states that model configuration uses temperature = 0, while Table 2 reports results averaged over 8 runs; please clarify whether the runs varied and report the observed variance, since deterministic decoding would make averaging redundant.
- [Abstract] There is a typo in the abstract: 'we propose CLARA' should be capitalized consistently as 'We propose CLARA'.
- [Benchmark Construction] The sentence 'comprised of 1184 questions across English, Spanish, Filipino, Portuguese, Mandarin, French, and Hindi' is ambiguous: it should state whether 1184 is the number of unique question stems with translations per language or the total number of items across all languages.
- [Benchmark Construction] The paper claims questions were 'carefully crafted to ensure question neutrality across regions,' but no procedure for verifying regional neutrality is described; please provide the protocol or soften the claim.
Circularity Check
No significant circularity: the paper is an empirical evaluation benchmark plus an inference-time pipeline compared against external baselines.
full rationale
The paper's central claims are empirical evaluations of LLM accuracy across languages and of an inference-time debiasing pipeline (CLARA) against direct inference and baseline methods. The benchmark questions are manually curated parallel translations; translation fidelity is an assumption about data quality, not a fitted parameter or derivation step, so any risk it creates is a validity risk rather than circularity. The debiasing results are measured, not derived from the benchmark definitions: CLARA's translation-to-English step is an intervention whose effect is reported empirically (e.g., GPT-4 Filipino rises from 51.8 to 67.1 in Table 3), and the residual gaps could in principle remain large. No equation in the paper defines a reported accuracy in terms of another reported quantity, and no fitted parameter is renamed as a prediction. Self-citations (e.g., Dychiao et al. 2024, Restrepo et al. 2024a/b) appear only as background support for LMIC data scarcity and prior observations, and the benchmark's comparisons rest on external model outputs rather than on those citations. The paper is therefore self-contained with respect to its evaluation claims, and the main caveat—whether translations preserve difficulty and cultural neutrality—is a benchmark-validity concern, not circularity.
Assumptions & free parameters
free parameters (3)
- RAG query-part weights w_j (and jargon weights w_k) =
not reported
- Maximum iteration limit for retrieval/refinement =
5
- Evaluation thresholds for translation certainty and retrieval relevance =
not reported
assumptions (4)
- domain assumption Manual translation preserves meaning, difficulty, and cultural neutrality across the seven languages.
- domain assumption Board-certified native-speaker ophthalmologist review guarantees correct gold answers and cultural appropriateness.
- domain assumption Averaged accuracy over 8 runs with temperature 0 is a stable estimator suitable for cross-lingual comparisons.
- domain assumption The selected chain-of-thought examples are representative of systematic model failures.
Cite this review
Pith. "Pith review of Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs." pith.science (2026). https://pith.science/paper/AFDLQ4DM
@misc{pith2026241214304,
author = {Pith},
title = {Pith review of: Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs},
year = {2026},
howpublished = {\url{https://pith.science/paper/AFDLQ4DM}},
note = {Machine review of arXiv:2412.14304}
}
read the original abstract
Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report summaries. However, LLMs have demonstrated significantly varied performance across different languages in natural language question-answering tasks, potentially exacerbating healthcare disparities in Low and Middle-Income Countries (LMICs). This study introduces the first multilingual ophthalmological question-answering benchmark with manually curated questions parallel across languages, allowing for direct cross-lingual comparisons. Our evaluation of 6 popular LLMs across 7 different languages reveals substantial bias across different languages, highlighting risks for clinical deployment of LLMs in LMICs. Existing debiasing methods such as Translation Chain-of-Thought or Retrieval-augmented generation (RAG) by themselves fall short of closing this performance gap, often failing to improve performance across all languages and lacking specificity for the medical domain. To address this issue, We propose CLARA (Cross-Lingual Reflective Agentic system), a novel inference time de-biasing method leveraging retrieval augmented generation and self-verification. Our approach not only improves performance across all languages but also significantly reduces the multilingual bias gap, facilitating equitable LLM application across the globe.
Figures
Forward citations
Cited by 1 Pith paper
-
BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning
BELO is a new ophthalmology benchmark of 900 expert-checked multiple-choice questions with reasoning, used to evaluate six LLMs on accuracy and explanation quality.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Antaki, F.; Touma, S.; Milad, D.; El-Khoury, J.; and Duval, R. 2023. Evaluating the performance of ChatGPT in ophthalmology: an analysis of its successes and shortcomings. Ophthalmology science, 3(4): 100324
work page 2023
-
[4]
Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; and Hajishirzi, H. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511
arXiv 2023
-
[5]
Canese, K.; and Weis, S. 2013. PubMed: the bibliographic database. The NCBI handbook, 2(1)
work page 2013
-
[6]
Chai, L.; Yang, J.; Sun, T.; Guo, H.; Liu, J.; Wang, B.; Liang, X.; Bai, J.; Li, T.; Peng, Q.; et al. 2024. xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning. arXiv preprint arXiv:2401.07037
arXiv 2024
-
[7]
H.; Chen, J.; Zhang, H.; Jianquan, L.; Xiang, W.; and Wang, B
Chen, Z.; Yan, S.; Liang, J.; Jiang, F.; Wu, X.; Yu, F.; Chen, G. H.; Chen, J.; Zhang, H.; Jianquan, L.; Xiang, W.; and Wang, B. 2023. MultilingualSIFT: Multilingual Supervised Instruction Fine-tuning
work page 2023
-
[8]
Doshi, J.; Kashyap Jois, A. K.; Hanna, K.; and Anandan, P. 2023. The LLM Landscape for LMICs. arxiv
work page 2023
Show all 63 references
-
[9]
Dychiao, R. G. K.; Alberto, I. R. I.; Artiaga, J. C. M.; Salongcay, R. P.; and Celi, L. A. 2024. Large language model integration in Philippine ophthalmology: early challenges and steps forward. The Lancet Digital Health, 6(5): e308
2024
-
[10]
Fan, X.; and Tao, C. 2024. Towards resilient and efficient llms: A comparative study of efficiency, performance, and adversarial robustness. arXiv preprint arXiv:2408.04585
2024 arXiv
-
[11]
Grandinetti, J.; and McBeth, R. 2024. From Generalist to Specialist: Improving Large Language Models for Medical Physics Using ARCoT. arXiv preprint arXiv:2405.11040
2024 arXiv
-
[12]
Han, X.; Zhang, J.; Liu, Z.; Tan, X.; Jin, G.; He, M.; Luo, L.; and Liu, Y. 2023. Real-world visual outcomes of cataract surgery based on population-based studies: a systematic review. British Journal of Ophthalmology, 107(8): 1056--1065
2023
-
[13]
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021. Measuring Massive Multitask Language Understanding. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[14]
Henkel, O.; Hills, L.; Roberts, B.; and McGrane, J. 2023. Can LLMs Grade Short-answer Reading Comprehension Questions: Foundational Literacy Assessment in LMICs. arXiv preprint arXiv:2310.18373
2023 arXiv
-
[15]
X.; Song, T.; Xia, Y.; and Wei, F
Huang, H.; Tang, T.; Zhang, D.; Zhao, W. X.; Song, T.; Xia, Y.; and Wei, F. 2023. Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting. arXiv preprint arXiv:2305.07004
2023 arXiv
-
[16]
Intrator, Y.; Halfon, M.; Goldenberg, R.; Tsarfaty, R.; Eyal, M.; Rivlin, E.; Matias, Y.; and Aizenberg, N. 2024. Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications? arXiv preprint arXiv:2403.04792
2024 arXiv
-
[17]
Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14): 6421
2021
-
[18]
Jin, Q.; Dhingra, B.; Liu, Z.; Cohen, W.; and Lu, X. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language P...
2019
-
[19]
C.; Yeganova, L.; Wilbur, W
Jin, Q.; Kim, W.; Chen, Q.; Comeau, D. C.; Yeganova, L.; Wilbur, W. J.; and Lu, Z. 2023. MedCPT: Contrastive Pre-trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11): btad651
2023
-
[20]
Kaur, D.; Uslu, S.; Durresi, M.; and Durresi, A. 2024. LLM-Based Agents Utilized in a Trustworthy Artificial Conscience Model for Controlling AI in Medical Applications. In International Conference on Advanced Information Networking and Applications, 198--209. Springer
2024
-
[21]
Labrak, Y.; Bazoge, A.; Dufour, R.; Rouvier, M.; Morin, E.; Daille, B.; and Gourraud, P.-A. 2023. FrenchMedMCQA: A French multiple-choice question answering dataset for medical domain. arXiv preprint arXiv:2304.04280
2023 arXiv
-
[22]
Lai, W.; Mesgar, M.; and Fraser, A. 2024. LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback. arXiv preprint arXiv:2406.01771
2024 arXiv
-
[23]
u ttler, H.; Lewis, M.; Yih, W.-t.; Rockt \
Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; K \"u ttler, H.; Lewis, M.; Yih, W.-t.; Rockt \"a schel, T.; et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33: 9459--9474
2020
-
[24]
Li, D.; Kadav, A.; Gao, A.; Li, R.; and Bourgon, R. 2024 a . Automated Clinical Data Extraction with Knowledge Conditioned LLMs. arXiv preprint arXiv:2406.18027
2024 arXiv
-
[25]
Li, H.; Zhang, Y.; Koto, F.; Yang, Y.; Zhao, H.; Gong, Y.; Duan, N.; and Baldwin, T. 2023 a . CMMLU: Measuring massive multitask language understanding in Chinese. arXiv:2306.09212
2023 arXiv
-
[26]
Li, J.; Tang, Z.; Liu, X.; Spirtes, P.; Zhang, K.; Leqi, L.; and Liu, Y. 2024 b . Steering LLMs Towards Unbiased Responses: A Causality-Guided Debiasing Framework. arXiv preprint arXiv:2403.08743
2024 arXiv
-
[27]
Li, J.; Zhang, H.; Zhang, F.; Chang, T.-W.; Kuang, K.; Chen, L.; and Zhou, J. 2024 c . Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10122--10140
2024
-
[28]
S.; and Wen, L
Li, S.; Hu, X.; Liu, A.; Yang, Y.; Ma, F.; Yu, P. S.; and Wen, L. 2023 b . Enhancing cross-lingual natural language inference by soft prompting with multilingual verbalizer. arXiv preprint arXiv:2305.12761
2023 arXiv
-
[29]
S.; Chiang, M
Lin, W.-C.; Chen, J. S.; Chiang, M. F.; and Hribar, M. R. 2020. Applications of artificial intelligence to electronic health record data in ophthalmology. Translational vision science & technology, 9(2): 13--13
2020
-
[30]
Liu, J.; Zhou, P.; Hua, Y.; Chong, D.; Tian, Z.; Liu, A.; Wang, H.; You, C.; Guo, Z.; Zhu, L.; et al. 2023. Benchmarking Large Language Models on CMExam--A Comprehensive Chinese Medical Exam Dataset. arXiv preprint arXiv:2306.03030
2023 arXiv
-
[31]
K.; and Andrade, R
Malerbi, F. K.; and Andrade, R. E. 2022. Real-World diabetic retinopathy screening with a handheld fundus camera in a high-burden setting. Acta Ophthalmologica (1755375X), 100(8)
2022
-
[32]
Marchisio, K.; Ko, W.-Y.; B \'e rard, A.; Dehaze, T.; and Ruder, S. 2024. Understanding and Mitigating Language Confusion in LLMs. arXiv preprint arXiv:2406.20052
2024 arXiv
-
[33]
B.; Quao, N
Mensah, P. B.; Quao, N. S.; and Group, P. G. C. E. 2024. Can Large Language Models Provide Emergency Medical Help Where There Is No Ambulance? A Comparative Study on Large Language Model Understanding of Emergency Medical Scenarios in Resource-Constrained Settings. medRxiv, 2024--04
2024
-
[34]
Nath, S.; Marie, A.; Ellershaw, S.; Korot, E.; and Keane, P. A. 2022. New meaning for NLP: the trials and tribulations of natural language processing with GPT-3 in ophthalmology. British Journal of Ophthalmology, 106(7): 889--892
2022
-
[35]
K.; and Sankarasubbu, M
Pal, A.; Umapathi, L. K.; and Sankarasubbu, M. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, 248--260. PMLR
2022
-
[36]
Plaza, I.; Melero, N.; del Pozo, C.; Conde, J.; Reviriego, P.; Mayor-Rocher, M.; and Grandury, M. 2024. Spanish and LLM Benchmarks: is MMLU Lost in Translation? arXiv preprint arXiv:2406.17789
2024 arXiv
-
[37]
Poulain, R.; Fayyaz, H.; and Beheshti, R. 2024. Bias patterns in the application of LLMs for clinical decision support: A comprehensive study. arXiv preprint arXiv:2404.15149
2024 arXiv
-
[38]
Ren, Z.; Zhan, Y.; Yu, B.; Ding, L.; and Tao, D. 2024. Healthcare copilot: Eliciting the power of general llms for medical consultation. arXiv preprint arXiv:2402.13408
2024 arXiv
-
[39]
M.; Do Carmo Novaes, F.; Azevedo Costa, I
Restrepo, D.; Quion, J. M.; Do Carmo Novaes, F.; Azevedo Costa, I. D.; Vasquez, C.; Bautista, A. N.; Quiminiano, E.; Lim, P. A.; Mwavu, R.; Celi, L. A.; et al. 2024 a . Ophthalmology Optical Coherence Tomography Databases for Artificial Intelligence Algorithm: A Review. In Sem...
2024
-
[40]
Restrepo, D.; Wu, C.; V \'a squez-Venegas, C.; Matos, J.; Gallifant, J.; and Nakayama, L. F. 2024 b . Analyzing Diversity in Healthcare LLM Research: A Scientometric Perspective. medRxiv, 2024--06
2024
-
[41]
L.; Cifuentes-Gonz \'a lez, C.; Wei, Y
Rojas-Carabali, W.; Agrawal, R.; Gutierrez-Sinisterra, L.; Baxter, S. L.; Cifuentes-Gonz \'a lez, C.; Wei, Y. C.; Arputhan, A. J.; Kannapiran, P.; Wong, S.; Lee, B.; et al. 2024. Natural Language Processing in Medicine and Ophthalmology: A Review for the 21st-century clinician...
2024
-
[42]
Ruder, S.; Constant, N.; Botha, J.; Siddhant, A.; Firat, O.; Fu, J.; Liu, P.; Hu, J.; Garrette, D.; Neubig, G.; et al. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. arXiv preprint arXiv:2104.07412
2021 arXiv
-
[43]
Salemi, A.; and Zamani, H. 2024. Evaluating Retrieval Quality in Retrieval-Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2395--2400
2024
-
[44]
Shafayat, S.; Kim, E.; Oh, J.; and Oh, A. 2024. Multi-FAct: Assessing Multilingual LLMs' Multi-Regional Knowledge using FActScore. arXiv preprint arXiv:2402.18045
2024 arXiv
-
[45]
W.; Tay, Y.; Ruder, S.; Zhou, D.; et al
Shi, F.; Suzgun, M.; Freitag, M.; Wang, X.; Srivats, S.; Vosoughi, S.; Chung, H. W.; Tay, Y.; Ruder, S.; Zhou, D.; et al. 2022. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057
2022 arXiv
-
[46]
Shi, Y.; Zi, X.; Shi, Z.; Zhang, H.; Wu, Q.; and Xu, M. 2024. Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems. arXiv preprint arXiv:2407.10670
2024 arXiv
-
[47]
Talebirad, Y.; and Nadiri, A. 2023. Multi-agent collaboration: Harnessing the power of intelligent llm agents. arXiv preprint arXiv:2306.03314
2023 arXiv
-
[48]
J.; Selva, D.; and Chan, W
Tan, Y.; Bacchi, S.; Casson, R. J.; Selva, D.; and Chan, W. 2020. Triaging ophthalmology outpatient referrals with machine learning: a pilot study. Clinical & experimental ophthalmology, 48(2): 169--173
2020
-
[49]
Tao, C.; Fan, X.; and Yang, Y. 2024. Harnessing llms for api interactions: A framework for classification and synthetic data generation. arXiv preprint arXiv:2409.11703
2024 arXiv
-
[50]
Vilares, D.; and G \'o mez-Rodr \'i guez, C. 2019. HEAD - QA : A Healthcare Dataset for Complex Reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 960--966. Florence, Italy: Association for Computational Linguistics
2019
-
[51]
Vrande c i \'c , D.; and Kr \"o tzsch, M. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10): 78--85
2014
-
[52]
Vu, T.; Iyyer, M.; Wang, X.; Constant, N.; Wei, J.; Wei, J.; Tar, C.; Sung, Y.-H.; Zhou, D.; Le, Q.; et al. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214
2023 arXiv
-
[53]
Wang, H.; Minervini, P.; and Ponti, E. M. 2024. Probing the Emergence of Cross-lingual Alignment during LLM Training. arXiv preprint arXiv:2406.13229
2024 arXiv
-
[54]
Wang, W.; Haddow, B.; Peng, W.; and Birch, A. 2024. Sharing Matters: Analysing Neurons Across Languages and Tasks in LLMs. arXiv preprint arXiv:2406.09265
2024
-
[55]
J.; Goetz, L.; Watson, D.; and van der Schaar, M
Wei, Q.; Chan, A. J.; Goetz, L.; Watson, D.; and van der Schaar, M. 2024. Actions Speak Louder than Words: Superficial Fairness Alignment in LLMs. In ICLR 2024 Workshop on Reliable and Responsible Foundation Models
2024
-
[56]
Xidong Wang*, D. S. Z. Z. Q. X. X. W. F. J. J. L. B. W., Guiming Hardy Chen*. 2023. CMB: Chinese Medical Benchmark. https://github.com/FreedomIntelligence/CMB. Xidong Wang, Guiming Hardy Chen, Dingjie Song, and Zhiyi Zhang contributed equally to this github repo
2023
-
[57]
Xu, Y.; Hu, L.; Zhao, J.; Qiu, Z.; Ye, Y.; and Gu, H. 2024. A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias. arXiv preprint arXiv:2404.00929
2024 arXiv
-
[58]
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 a . Qwen2 technical report. arXiv preprint arXiv:2407.10671
2024 arXiv
-
[59]
T.; et al
Yang, Z.; Wang, D.; Zhou, F.; Song, D.; Zhang, Y.; Jiang, J.; Kong, K.; Liu, X.; Qiao, Y.; Chang, R. T.; et al. 2024 b . Understanding Natural Language: Potential Application of Large Language Models to Ophthalmology. Asia-Pacific Journal of Ophthalmology, 100085
2024
-
[60]
Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.-J.; and Huang, G. 2024 a . Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 19632--19642
2024
-
[61]
Zhao, J.; Ding, Y.; Jia, C.; Wang, Y.; and Qian, Z. 2024 b . Gender Bias in Large Language Models across Multiple Languages. arXiv preprint arXiv:2403.00277
2024 arXiv
-
[62]
Zhao, J.; Zhang, Z.; Zhang, Q.; Gui, T.; and Huang, X. 2024 c . Llama beyond english: An empirical study on language capability transfer. arXiv preprint arXiv:2401.01055
2024 arXiv
-
[63]
Zhu, Y.; Moniz, J. R. A.; Bhargava, S.; Lu, J.; Piraviperumal, D.; Li, S.; Zhang, Y.; Yu, H.; and Tseng, B.-H. 2024. Can Large Language Models Understand Context? arXiv preprint arXiv:2402.00858
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.