Pith. sign in

Paper Citation Record · LEDGER

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.19217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19217 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:10:46.701581Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0cbf479e-e58c-4054-a24a-01bfb744648e · outbound

This paper cites Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Computed tomography and magnetic resonance imaging: past, present and future.European Respiratory Journal, 19(35 suppl):3s–12s, 2002

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.485949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:42.598520Z digest=sha256:42817f3754764b10be0bd3701ae93ec67cf2e945558dfe66063d0cc185de03d7

Observation 9e6857e0-9f0d-4a9e-b64b-987df1fcd7a5 · outbound

This paper cites Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Should we be concerned about the rapid increase in ct usage?Reviews on environmental health, 25(1):63–68, 2010

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.469459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:42.647398Z digest=sha256:7fd3d879776f503e9d65ad76ea0f3dead3006a987e064f8fe242208e41eae406

Observation 1786067b-301c-4b76-864b-451275c3d7a2 · outbound

This paper cites Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cognitive and system factors contribut- ing to diagnostic errors in radiology.American Journal of Roentgenology, 201(3):611–617, 2013

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.454129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:42.739674Z digest=sha256:154f41473db9746a16cb56d964975c23c41668de6dcdb003e6d64129a92a8a33

Observation 2a811af5-e706-4dc4-974a-dacc0cbcc068 · outbound

This paper cites Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Automated radiology report generation: A review of recent advances.IEEE Reviews in Biomedical Engineer- ing, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.433752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:42.826120Z digest=sha256:3755cb979534130a3f329234c8409de83ce05081c5adab6b28796ea0520dbc5b

Observation 6ef77ae9-f7cf-4339-bb01-206237bc6697 · outbound

This paper cites Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Comparing diagnostic accuracy of radiolo- gists versus gpt-4v and gemini pro vision using image inputs from diagnosis please cases.Radiology, 312(1), July 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.417716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:42.915903Z digest=sha256:52c4c727ed302347647dd0a6a615ff13ac97997b9fa69fa9f518187762fc7be4

Observation 9ce878c1-d2f8-4c0f-a405-be2a867a5b76 · outbound

This paper cites Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Evaluating large language models on medical evidence summarization.NPJ digital medicine, 6(1):158, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.368002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.011449Z digest=sha256:1563668eea59bd4e73671327d62cccd18a4306459b011736ffdfd6ffbafca06d

Observation e1bfba31-07e9-48d4-aa5e-dd1b23790d72 · outbound

This paper cites Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Embracing large language models for medical applications: opportuni- ties and challenges.Cureus, 15(5), 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.234294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.097684Z digest=sha256:d69587cf0d37d02f72a886b659b5b36b36b5dd9cbe74a7e9839f95daf3fec9a3

Observation a1cfe936-bb90-4e1f-9b69-6e983c5c4bdd · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:51.108888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.193737Z digest=sha256:cd2c28579d06485da83e8336e034ecee247da81d187793093989500fa08aace5

Observation edb9e078-c00c-4196-aa6c-2063a1b601e6 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.293038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.293038Z digest=sha256:1d0a3649ec3a901964e6e5387384433507123894bc5e5edfe711d54f36d32c85

Observation fb21cd16-96b6-45dd-99f6-433cc62987a8 · outbound

This paper cites Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.786538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.381150Z digest=sha256:efea9fccd77a0626feaaa5f28f8e3a45efdbe5bf9a9efa244e29fb7145c24244

Observation 428a250c-52c1-46a2-b798-e1b4bd39fbcc · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.479305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.479305Z digest=sha256:664a86f27c3522ae2ca082544c55a85d0fd08c3e1363eaa3019cea78fa95bd75

Observation 115ddd63-1894-4bef-9a33-8707ebcfab85 · outbound

This paper cites Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.479972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.555156Z digest=sha256:4cc8434963fdf6add312fcab9b54c15a2b83128733ce70e590ae9113273a9003

Observation 88b74f26-1912-41a0-ab9b-ccc68fbf59d6 · outbound

This paper cites Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:50.210620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.645745Z digest=sha256:4988dd3504b0bead05795bde00733845c812088c9f8c70e20b4ed0795e5c9c18

Observation 05f1c7c1-6992-4164-925c-2b7b3d5f5976 · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:43.721022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:43.721022Z digest=sha256:141063fcd7d7b56190d6472e6efb81cf454d795e74c9988c47f66cc1a4803a23

Observation a8d1fb6e-ba17-4062-b407-3ce167e99ed7 · outbound

This paper cites Mme-survey: A comprehensive survey on evaluation of multimodal llms,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mme-survey: A comprehensive survey on evaluation of multimodal llms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.811271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.818300Z digest=sha256:26aa1a68190c445a96ab6ae8ba9907c7d8adb110e92de1b46d42b7a2e83e8505

Observation 69e2268f-8bd6-4168-b4ab-fcec64620eab · outbound

This paper cites Kim and Liem T.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Kim and Liem T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.430981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.903843Z digest=sha256:5fd6922eb3d1f7970359c859a08c5a53e81baa6eca83ef1b1964c2a780d823ee

Observation c25235c2-3a65-424e-9262-fa9a93711142 · outbound

This paper cites Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Recovery at the edge of error: debunking the myth of the infallible expert.Journal of biomedical in- formatics, 44(3):413–424, 2011

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:49.078302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:43.983906Z digest=sha256:9ed944987c10f31ba4cc800ed06d490905cbbc9187b7c3ae48c073ece924db6e

Observation e6c3db7a-8e7f-4548-898c-e984a7b7deae · outbound

This paper cites Overview of the mediqa-corr 2024 shared task on medical error detection and correction.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Overview of the mediqa-corr 2024 shared task on medical error detection and correction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.799582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:44.067488Z digest=sha256:f6bf1da0ed988831c3c979e29fab3d14100d9cf65126a43be842bea9a794756f

Observation 06a40d3c-bd45-40a2-a5e0-0983b2bbd85b · outbound

This paper cites Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Potential of gpt-4 for detecting errors in radiology reports: implications for report- ing accuracy.Radiology, 311(1):e232714, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.626354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:44.141513Z digest=sha256:28dc27b6afc610f4d7911b58bdb1793dfeab874ccce69a924ae595064548a4af

Observation 18e7c4a7-cc8a-4549-9bf3-db27cfb420f7 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.228630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.228630Z digest=sha256:4d153d10ccf18ea1a021517834a51e6ac6fc57ba369f8778307d6327b354e221

Observation 83010895-4fda-422d-8ff1-67939e9412c1 · outbound

This paper cites 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.308483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.308483Z digest=sha256:7007dcfd97c6b7f0dcce4a751bcf6d3c21b3ee9ac510e374c17f7d2b575cb3e3

Observation 5353d755-11e3-4d34-a00d-9f1d3d9cb799 · outbound

This paper cites Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Med3dvlm: An efficient vision-language model for 3d med- ical image analysis, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.447620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:44.712221Z digest=sha256:dd030a29592b74f46302f9d6027babc7e9592d07584e775b7b6057a0e58cf429

Observation 0382bbba-26ec-4f69-8ff6-65e5914ada57 · outbound

This paper cites Medm-vl: What makes a good medical lvlm?,.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Medm-vl: What makes a good medical lvlm?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.276586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:44.801787Z digest=sha256:819cead4cf1e3c3ea2cbc7606cbe0350acc19913e02084f533ae345fdef9e1f1

Observation 17dd16c3-2a9e-4bc6-b893-7ac6135e1c63 · outbound

This paper cites MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.901214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.901214Z digest=sha256:da2f5dcec4bd5fd78e4e9f0aa223929ff6b31a9906c02aef1e8d093a3bc42170

Observation 33204f4a-b552-4226-9903-d3757868ced9 · outbound

This paper cites ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ReXErr: Synthesizing Clinically Meaningful Errors in Diagnostic Radiology Reports

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.037970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.037970Z digest=sha256:0bf6321018bafa19667ba0805d77e95de3ee3ffa23c77212e7ca386b9dd7e7b4

Observation 44d7a051-144e-4d14-950f-827898976d04 · outbound

This paper cites Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Mimic-cxr, a de- identified publicly available database of chest radiographs with free-text reports.Scientific data, 6(1):317, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.149757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:45.146921Z digest=sha256:4ebc096ace2f53f972ac7fb8020eaa1eadeb8146ce1af9f97b3e7d525c3d8523

Observation 63cfcc53-e338-4855-aab8-75092ce45f50 · outbound

This paper cites MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.285853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.285853Z digest=sha256:ad4e89780049d040dcc090a6ba2524140635015df11d0cad4e2e0fe8fcb52960

Observation bdb16b3b-681a-4534-a31f-e49f2746f9b5 · outbound

This paper cites A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero- shot detection of abnormalities.CoRR, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:48.034543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:45.400446Z digest=sha256:6977b6ced27ae56fdf94554bd85be9901d75876965e2fdb3369207e288480cad

Observation 4f65962b-c11c-4f1c-b76e-b002564a7ca5 · outbound

This paper cites RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.482676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.482676Z digest=sha256:b5641c0312bcf3de6f22eed25065ebee1b95fadebf12494f7b3cecdda6dfb345

Observation 6f7c1a52-eae0-465f-8324-1dfff9514139 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.904119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:45.564153Z digest=sha256:f27277fa0ef1052de2fa7e92b67757eeb84b3faa85657273841f79ba71729e78

Observation 99978041-19d1-4172-b7a5-b4633f686d55 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.655074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.655074Z digest=sha256:d483c58d997fe3d7215510580259763a5dfd603d324651d52a8132ce2622bf89

Observation d873694a-61ff-4e2e-8518-a284e4869ca4 · outbound

This paper cites The Llama 3 Herd of Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports The Llama 3 Herd of Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.736107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.736107Z digest=sha256:2b8b21c3ab115ae4932a3c22b32192d69268594b5d2e9e0331d8c612703d327e

Observation 21572ae9-4807-4a72-b881-70ca51cc44b3 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.868971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.868971Z digest=sha256:2e6ac0ba7da3ca371b6a49bfbce1ddd8cc9e304e5c4febf81f09ace762a1891a

Observation 470dc869-5bc2-4946-ac5e-dde649d80ed8 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:45.988146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:45.988146Z digest=sha256:b1db277982e90d422eed0990e2912f50d2b4ce0816b4b295924da3b59a2ead88

Observation 6d3e0d87-338f-4a8a-abce-dad19db84edf · outbound

This paper cites De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports De- veloping Generalist Foundation Models from a Multi- modal Dataset for 3D Computed Tomography, April 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.089541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.089541Z digest=sha256:6f24d20007e2d7f0817866c6a528e8d5b210e5ab88460c67786d94f3cd9a115b

Observation 89e9e2f8-049d-44d9-a7e0-021817e7a208 · outbound

This paper cites ROUGE: A Package for Automatic Evalu- ation of Summaries.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports ROUGE: A Package for Automatic Evalu- ation of Summaries

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.764041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:46.201273Z digest=sha256:a3bac3e185cc37686ee0e7ade37e63b516a30cfd3dcc31784fc96597397f46f7

Observation 49635171-5368-4672-9ab1-54707ea48709 · outbound

This paper cites Bleu: a Method for Automatic Evaluation of Ma- chine Translation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Bleu: a Method for Automatic Evaluation of Ma- chine Translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.559786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:46.283549Z digest=sha256:b3f395af3bfc66290eb54940e4d7031cedac84afa7c4dcf23e5fd16713790a0d

Observation cbd352a5-33a0-412f-8bd0-aeffdc2de515 · outbound

This paper cites METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:10:47.393684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:10:46.443896Z digest=sha256:96151072bf267b551329de8ae7cd7152f96bfe2d5ecd12ad693735bcd21efaa5

Observation 5fff4125-b51c-4d5c-ba59-8a7cedb4ce4e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports BERTScore: Evaluating Text Generation with BERT

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.566102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.566102Z digest=sha256:b8b0ee65eebecd113810c30698cd9b41683c7b4d05ff1979bb384d4931ac6a43

Observation 6e5a2bf0-69a9-4a05-8e98-d63b4cb7fd0a · outbound

This paper cites GREEN: Generative Radiology Report Evaluation and Error Notation.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports GREEN: Generative Radiology Report Evaluation and Error Notation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:46.701581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:46.701581Z digest=sha256:786edaee66c9210355466e714a0f4d5ae14c53bcf8458acf21c126196b0dbc35

Pith citing papers

No inbound Pith citation observations are available.