Pith. sign in

Paper Citation Record · LEDGER

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning

As of 15 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2507.15717.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15717 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:28:49.334512Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41b878ad-7ebf-4384-aebd-e6bc462a0da6 · outbound

This paper cites Attention is All you Need [Internet].

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Attention is All you Need [Internet]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:57.276819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.035082Z digest=sha256:08908c167cacc9f7fb93c22bcd244a533a6859b7143668b09e95b0de7faa230f

Observation b7a59abd-7264-4026-868e-40f46daa170e · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Towards Expert-Level Medical Question Answering with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:45.191838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:45.191838Z digest=sha256:3189beb7940012960ebe1958097d14237e8586e2e8370846923a34aa6ac63e82

Observation 6968ced4-5dd0-41b2-8e79-eae323c24c12 · outbound

This paper cites Towards Conversational Diagnostic AI.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Towards Conversational Diagnostic AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:45.332823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:45.332823Z digest=sha256:db9d302328a2edb8e3a4b9950d342f695b332d510ebce4b555f19f120932570a

Observation 2974194f-aad3-4e9a-aa91-77ced40dba2e · outbound

This paper cites Opportunities and challenges for ChatGPT and large language models in biomedicine and health.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Opportunities and challenges for ChatGPT and large language models in biomedicine and health

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:57.159541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.436050Z digest=sha256:a13831990aa7ace124c418bf71da90d851e0d355021f53c966142b090babddf8

Observation ae2c89fd-8c14-49e8-8fe9-893bb4abe982 · outbound

This paper cites Benchmarking large language models for biomedical natural language processing applications and recommendations.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Benchmarking large language models for biomedical natural language processing applications and recommendations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:57.005286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.529386Z digest=sha256:fe5dcb2260fb42cff2bfdba27bee05d43e9831687baed3c7dd61b521baeb9868

Observation f16aec01-0270-4d23-b243-633889ed059a · outbound

This paper cites Radiology-Llama2: Best-in-Class Large Language Model for Radiology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Radiology-Llama2: Best-in-Class Large Language Model for Radiology

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:45.636946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:45.636946Z digest=sha256:74c93eeabff60c4af2fab068c02dc43228c1dacb5190a6681549bbdec0b8b9da

Observation b435dda8-be36-4702-80eb-6a71439ee956 · outbound

This paper cites Towards a general-purpose foundation model for computational pathology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Towards a general-purpose foundation model for computational pathology

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.893597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.692928Z digest=sha256:8661a88577827ad187d5c7f85e9e8123234cc7228f9b3e558b73ca654380e92c

Observation c4e2200e-a08d-424e-b627-ee91fbd05561 · outbound

This paper cites Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model [Internet].

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model [Internet]

Reference 8

Resolution
verified exact
raw_fallback, observed 2026-08-06T15:28:51.092097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.752216Z digest=sha256:62206c2ef0979eb574812ae2cd075054ffdfdb37092119a13f5bda4afac2d101

Observation d0b17bca-a972-4431-becc-bd81ef17c053 · outbound

This paper cites Utilizing Large Language Models to Simplify Radiology Reports: a comparative analysis of ChatGPT3.5, ChatGPT4.0, Google Bard, and Microsoft Bing.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Utilizing Large Language Models to Simplify Radiology Reports: a comparative analysis of ChatGPT3.5, ChatGPT4.0, Google Bard, and Microsoft Bing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.791303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.820969Z digest=sha256:a50791937cbfda733b9862c84343449c467f94e139ce562f37cdc89674b0cdb8

Observation dd25c89c-50a6-45f5-9800-8396f2691f32 · outbound

This paper cites Can OpenAI’s New o1 Model Outperform Its Predecessors in Common Eye Care Queries? Ophthalmology Science [Internet] 2025 [cited 2025 Apr 26];5(4).

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Can OpenAI’s New o1 Model Outperform Its Predecessors in Common Eye Care Queries? Ophthalmology Science [Internet] 2025 [cited 2025 Apr 26];5(4)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.626722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.911562Z digest=sha256:bb47e7c3efb9d4c1ec2574454e6e8c4e58278afb8c3d362970c1f40379763529

Observation 4fbba0a1-01a9-43d9-acd1-4396a2de224a · outbound

This paper cites EyeGPT: Ophthalmic Assistant with Large Language Models.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning EyeGPT: Ophthalmic Assistant with Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:28:50.835793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:45.969695Z digest=sha256:dd4f72079116c8c143e59197feefbe2fe436ae4bc7e6e05f8bd1ca96d8b2db09

Observation dc9b8830-4855-42d4-8a87-377e0f522fe7 · outbound

This paper cites Can OpenAI o1’s Enhanced Reasoning Capabilities Extend to Ophthalmology? A Benchmark Study Across Large Language Models and Text Generation Metrics [Internet].

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Can OpenAI o1’s Enhanced Reasoning Capabilities Extend to Ophthalmology? A Benchmark Study Across Large Language Models and Text Generation Metrics [Internet]

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.497357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.099803Z digest=sha256:0eb86766d9698fe08a206bd95e048bf517d124829a8e3ecc60c21a3cf9848ddb

Observation eb89e847-6645-4f25-a93c-09548d54c317 · outbound

This paper cites ChatGPT-Generated Differential Diagnosis Lists for Complex Case–Derived Clinical Vignettes: Diagnostic Accuracy Evaluation.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning ChatGPT-Generated Differential Diagnosis Lists for Complex Case–Derived Clinical Vignettes: Diagnostic Accuracy Evaluation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.339531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.227312Z digest=sha256:f89fb1e063692d4325eec37da614d7c26cab168c45b1cf1d19eade29f080e839

Observation 77f3be25-38e3-4b8d-af47-cf236a59ef85 · outbound

This paper cites Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.164737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.311462Z digest=sha256:73ec9d028b134901a15e1fbeca6f90e4c70c4d289baf2803dfedcec10735c26b

Observation 0612924a-dd29-47d7-a6e0-0b232010179e · outbound

This paper cites Can off-the-shelf visual large language models detect and diagnose ocular diseases from retinal photographs? BMJ Open Ophth [Internet] 2025 [cited 2025 May 15];10(1).

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Can off-the-shelf visual large language models detect and diagnose ocular diseases from retinal photographs? BMJ Open Ophth [Internet] 2025 [cited 2025 May 15];10(1)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:56.057025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.394293Z digest=sha256:43be6b704a6f8046ca33d7d58853f819c52a07060ad5e546c584c210bf581897

Observation 5a512d6d-13b5-485e-b7eb-0988035bb7fb · outbound

This paper cites ChatGPT: the future of discharge summaries? The Lancet Digital Health 2023;5(3):e107–8.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning ChatGPT: the future of discharge summaries? The Lancet Digital Health 2023;5(3):e107–8

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.890375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.434939Z digest=sha256:4aee304c66063b4d415c831a463c357783c930474911acb77ea622f72954781a

Observation 9523c4c7-4528-4ebe-ba60-3136efa9f739 · outbound

This paper cites Large Language Models Seem Miraculous, but Science Abhors Miracles.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Large Language Models Seem Miraculous, but Science Abhors Miracles

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.584987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.505475Z digest=sha256:4e82987d39cb39d85e41b0a55e91553e16f3b745fc177b478ae443d0f6549c7c

Observation ae5f0ee8-f75f-4af2-9fa2-603115a10bf7 · outbound

This paper cites Medical Ethics of Large Language Models in Medicine.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Medical Ethics of Large Language Models in Medicine

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.484975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.581662Z digest=sha256:2f09401402ecfc897ffeff4c319182b69673bd5a39f310464a062fa74906af17

Observation 64f78800-3061-4837-8c3b-de7ea129d70a · outbound

This paper cites Google for Developers.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Google for Developers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.354156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.668047Z digest=sha256:b68227e64a63d0b23ea6d4419af386208746d3fc1970e51cefd1d28701ba61a9

Observation 82a1fe2c-2b66-4048-a245-fd54a38c2915 · outbound

This paper cites Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.225606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.736571Z digest=sha256:0ab5f7da59063e8f665b19040a439aea87c2f8a8cf1035eaf3f272d432e6b9ca

Observation ab00acb6-03f0-49cb-918d-09ea297d8ad5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:46.831673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:46.831673Z digest=sha256:2373233c37aaf07c5e8419505b6dffc92186709d7e93a26436e67e1d42b05a59

Observation d8afb170-bfd4-4613-84da-88c6a6f8b200 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:55.030073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.891283Z digest=sha256:ae5c93c5bd974ba53625b52501d7851ac5a7fa4b2abd85f244177d7786d15870

Observation ceb9921a-cccb-40af-90b1-b0da72ed42a0 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.911878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:46.949387Z digest=sha256:66493808172ece153f801edcce00bab8d01fe7ec78bb512252846185efdc2e40

Observation 6c297985-bf6d-4bb5-9c5a-ad39439d5b36 · outbound

This paper cites MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.766328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.030758Z digest=sha256:d763bda6dba0f4c66d90fe8e266f7001f57bc0fcb42e6f42fd36afd752d2753c

Observation d8c35cad-0cd8-4378-9ce3-dd0d82383e72 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:47.108780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:47.108780Z digest=sha256:79771629d462afef3698785fee6d4ddbceadbcdef833f29932fcf44faa5c96e5

Observation 620c27bb-70d6-4918-8fb5-5eef302943b4 · outbound

This paper cites EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:47.162607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:47.162607Z digest=sha256:181c4bdb4682599c05a2d1bb243763658048ebb42e2b79a2ebcc9a9e7e6264b3

Observation 793f4a7b-ebd2-484d-b992-0539ba088dbf · outbound

This paper cites Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:28:50.529350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.251903Z digest=sha256:89fad121235fd00267144f8a62ef8add79e52b7529ddecc5f9ccc4c91314107b

Observation a8106e3a-06e1-46d6-bf98-58024df53215 · outbound

This paper cites Unveiling the clinical incapabilities: a benchmarking study of GPT-4V(ision) for ophthalmic multimodal image analysis.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Unveiling the clinical incapabilities: a benchmarking study of GPT-4V(ision) for ophthalmic multimodal image analysis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.660686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.339856Z digest=sha256:7589e78768d3c11c4e277c1000f2080642c07ecad4737e2d10303c01bbb9a8b3

Observation 7dab4c86-1fba-44e8-97ed-0aa0aeaf3785 · outbound

This paper cites Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings

Reference 29

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:28:50.372241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.412379Z digest=sha256:2f689d51ab260cad6da33442cbd88569d7ce2f91ddba65632744a6c7367b85eb

Observation 7cc5c6d7-369d-43b7-81c6-7381732dce9c · outbound

This paper cites Evaluating Large Language Models in Ophthalmology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Evaluating Large Language Models in Ophthalmology

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:28:50.050972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.476669Z digest=sha256:aae2ff89921801b96777a804d3dd14eea4dffc4990fd9113b944eb26dd3e4301

Observation 71bd979c-52ef-432c-96cc-8e3fa3c7fbc4 · outbound

This paper cites OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:28:49.862916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.559357Z digest=sha256:46427cb5b2d3abfd7dcdf90eec2d89483fcf9af02b31ae7a88c2e4ef72be71c6

Observation 7a40c24a-e822-4dee-85cc-54da13ee7c41 · outbound

This paper cites EYE-Llama, an in-domain large language model for ophthalmology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning EYE-Llama, an in-domain large language model for ophthalmology

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.534407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.652656Z digest=sha256:3d121a92a2e0dd64213dcf5377d56334a9a1465a575f52d2b611eb77ff54b4f2

Observation ac32d7b4-1998-4ac3-9e4d-29eb344278e7 · outbound

This paper cites BioASQ-QA: A manually curated corpus for Biomedical Question Answering.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning BioASQ-QA: A manually curated corpus for Biomedical Question Answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.420785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.743157Z digest=sha256:b6472c887e31c4d1b1c347dfefd1d75040fe8a12e81330136fb40a1d45f3be27

Observation a14f5cc9-a325-4875-a478-471344818985 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.253238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.803904Z digest=sha256:50b55c58a77ab05756234d155c694db0bde4c149ed803b0090c9a711b5076c8b

Observation 2feb1c9a-705f-40ea-94fd-ce61c8dc9cdb · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:54.129247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.857163Z digest=sha256:22a79376c81c38662b0bec3554c696eca3453113b4c1abfe15280ea763305366

Observation d80f5379-265e-4c45-8b36-9a23073f0a87 · outbound

This paper cites Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:53.878466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:47.913858Z digest=sha256:d0320e1b127f99679068dd0e3cc71ce1f51e862329c4eab58c8242b6b9763582

Observation 89a81b06-6227-4b91-8831-4b97f4ad7987 · outbound

This paper cites How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:53.578158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.008049Z digest=sha256:a56882cb49dfb7777f8dcb005c66ff768c37031925b3fefc87054f993ad2f7e1

Observation d4321c0a-7ddc-4671-b69d-38d1c30a9009 · outbound

This paper cites Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:28:49.676971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.085423Z digest=sha256:c1ee10134f64f4cd78c09e37ea4363a5620382bd4e5e8d1d46a3ad93fa0b73c5

Observation 56f9c3ce-b564-4ea3-9fa2-20ff9cf01447 · outbound

This paper cites ROUGE: A Package for Automatic Evaluation of Summaries [Internet].

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning ROUGE: A Package for Automatic Evaluation of Summaries [Internet]

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:53.374598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.171460Z digest=sha256:4f25fed3aa494de888ca953f569e5e364ebddf4e959e229a0b02edffd5016084

Observation 2f114fce-6175-467b-9f64-0ea339905859 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning BERTScore: Evaluating Text Generation with BERT

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:48.277853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:48.277853Z digest=sha256:f3ffc6c1c7ffeedba96326e302fe148e9f110d79d3aa4bd93ca149f4c0b781de

Observation 803eb354-c9b5-4c9a-aecf-cf64380105c3 · outbound

This paper cites BARTScore: Evaluating Generated Text as Text Generation.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning BARTScore: Evaluating Generated Text as Text Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:48.347345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:48.347345Z digest=sha256:1b3387c9f455e4b4eb479227cc663c56e96267c160b34402fc81efe98a01adf8

Observation 005638be-e921-4c09-bc34-fda54f712046 · outbound

This paper cites AlignScore: Evaluating Factual Consistency with a Unified Alignment Function.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning AlignScore: Evaluating Factual Consistency with a Unified Alignment Function

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:48.415628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:48.415628Z digest=sha256:0d62e4a8fa5f535f75cbad164717bca867a7b38ebd657593ad390da1f8782775

Observation 901cef58-08f4-4bde-86c0-d97e2fd9bd10 · outbound

This paper cites A visual-language foundation model for computational pathology.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning A visual-language foundation model for computational pathology

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:53.077371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.535484Z digest=sha256:876c8daafcf69bfbbbbdca3b69753e2189f491242a424119a7ac52a095b6abae

Observation a3d77380-73a4-477c-a265-12412b9e81ff · outbound

This paper cites Developing and Evaluating Large Language Model– Generated Emergency Medicine Handoff Notes.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Developing and Evaluating Large Language Model– Generated Emergency Medicine Handoff Notes

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:52.795191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.666664Z digest=sha256:3da36e1f14357a753bd0d97e00fbbc2e5d4b8ac7fe0585c41249a9149f0ecec9

Observation a8c674eb-fc2d-4269-8c39-fa32849d718d · outbound

This paper cites CPMI-ChatGLM: parameter-efficient fine-tuning ChatGLM with Chinese patent medicine instructions.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning CPMI-ChatGLM: parameter-efficient fine-tuning ChatGLM with Chinese patent medicine instructions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:52.523882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.703934Z digest=sha256:00afc199ea79b7cae3bbae446e27ad1195c23c9a649f984176f263018275042d

Observation c9a1c9a2-34c0-41da-a883-64f5401f8426 · outbound

This paper cites Reducing hallucinations of large language models via hierarchical semantic piece.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Reducing hallucinations of large language models via hierarchical semantic piece

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:52.356278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.790885Z digest=sha256:a142664ea90b4a16585df3a89d7e8d6c73043bf4da2555e51a1a975e08b26ffd

Observation 6c1cbc6e-4805-4862-b218-907ab14c75b3 · outbound

This paper cites Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:52.100594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.835919Z digest=sha256:d0b852e3c5948b5628b4aeb1ce2ed63528b15106928f482f5a4f2ae163a98675

Observation 8ddb41ec-a6c3-471f-b196-f0edbf68b762 · outbound

This paper cites LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:28:48.930283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:48.930283Z digest=sha256:229738c121eea2e77f6e4b32584ec6323e944bb83cbcf159a3a9a56b31fa816c

Observation b030e2a6-9b2d-4496-a218-a08b76dbfc17 · outbound

This paper cites medical QA dataset.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning medical QA dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:51.799503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:48.982187Z digest=sha256:3ffd760e7b13f5e926b6583e765f7ce4c4156c95190ad42132429d3ae15448a9

Observation a5402410-20ca-46c3-9d2d-a7920e64ef30 · outbound

This paper cites Each question offers four options, out of which only one is correct.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Each question offers four options, out of which only one is correct

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:51.668008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:49.197532Z digest=sha256:764e4ecd4fb897c129f33eb4bf614868997d4f93e45a9669ec939a8ba7574a19

Observation af61e65b-595e-4430-b810-8a7cd2432b2e · outbound

This paper cites glaucoma,.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning glaucoma,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:28:51.495773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:49.272071Z digest=sha256:ebac6ab57ae457036ff213871c6be98b5bb7e5815a3b63cbaa9c5e9d5655b5bd

Observation c2becfc3-52ba-434d-9744-9d5071f96789 · outbound

This paper cites an unresolved cited work.

BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:28:51.338945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T15:28:49.334512Z digest=sha256:3e7eb9e4f9a75de78cd899570407a104453ca66849c17eea3acfdf166352efb2

Pith citing papers

No inbound Pith citation observations are available.