Pith. sign in

Paper Citation Record · LEDGER

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.14544.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14544 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:57:51.244796Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:57:31.114754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed17e8bd-9b45-4cc5-b20f-4ed2d5ae7b98 · outbound

This paper cites Vqa: Visual question an- swering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Vqa: Visual question an- swering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.715301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.178667Z digest=sha256:14cafb8ef6bbc21dd7c472389dc62edd126ab0c23dffc51d0d81bc5086fae175

Observation 917e16fc-b3fd-496a-a4c1-591ffb87165f · outbound

This paper cites Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.182017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.182017Z digest=sha256:1935c2e0d0f16adfcf0f874bd56a5c609141fd577eefc4411ff2d89e260c0981

Observation a5c61c48-ebcc-41e9-bced-86d5c1f9d7f1 · outbound

This paper cites Exploratory data analysis.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Exploratory data analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.708624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.185072Z digest=sha256:f8293eb4658f2184e515f7f17e234bef398a705bd2a5e107efac3dbf6415ee49

Observation 19671860-56d1-4ab7-8524-a1604c3096d2 · outbound

This paper cites Mapping medical image-text to a joint space via masked modeling.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Mapping medical image-text to a joint space via masked modeling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.701691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.187745Z digest=sha256:e9762fcdaaf215375f046197517d0adea5e22acb334b1594148f4901521b3acc

Observation 4452c226-e998-4c73-9871-6ea9c39b2f39 · outbound

This paper cites Vision- language transformer and query generation for referring segmentation.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Vision- language transformer and query generation for referring segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.694595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.193007Z digest=sha256:bce0ee7b6f086c653336ebb05b037d29ae2b80b07d419d5cd320bf33aedd2ee1

Observation 0c7a4cd1-cc89-4c54-b125-36b774af5824 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.687329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.195959Z digest=sha256:c7207a5d491ef1e40f3d374628544532d64e8fb783d71682d50243b5241e4b4a

Observation 1722c8a3-45d2-4679-8dc0-42c7e77181c4 · outbound

This paper cites Bridging Multimedia Modalities: Enhanced Multimodal AI Understanding and Intelligent Agents.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Bridging Multimedia Modalities: Enhanced Multimodal AI Understanding and Intelligent Agents

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.679840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.198637Z digest=sha256:27731ec1678198c1a62e372c456b98560ad6aebfb0b85a93d3671e6736515920

Observation f0beebc1-9999-4b7e-97b3-738c0a256f0c · outbound

This paper cites Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.206063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.206063Z digest=sha256:0427df8cca6a760bdd986feefc8ef65cbecab9d4d3fa5d87d70e0eaca5c3d2e3

Observation 86233827-70c9-4c20-85fb-799eb5c066fa · outbound

This paper cites Hicks, Vajira Thambawita, P ˚ al Halvorsen, and Michael A.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Hicks, Vajira Thambawita, P ˚ al Halvorsen, and Michael A

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.376598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.208731Z digest=sha256:145c43dfbb919a9fdaa6e35a82aedcf4c921c632ecd42c27d3b37e05fd78a2cd

Observation e741dce1-0bcc-4023-93c1-4a08f0304e2c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image under- standing in visual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Making the v in vqa matter: Elevating the role of image under- standing in visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.665240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.211245Z digest=sha256:a69155db8c576739deb18c31f65692c432ec05c9a4e43950c9e22134c33472a4

Observation b96e36bc-40d1-4992-98fc-e17246fb0c95 · outbound

This paper cites Unk-vqa: A dataset and a probe into the abstention ability of multi- modal large models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unk-vqa: A dataset and a probe into the abstention ability of multi- modal large models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.657327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.213584Z digest=sha256:f15c5a2c72df713d179e29be2df9ffac80a56f8ba7ae4c03ac2f4eff3b49e11c

Observation e42d92b9-b616-42be-a79a-447daed339a4 · outbound

This paper cites Unboxing the black box of attention mech- anisms in remote sensing big data using xai.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unboxing the black box of attention mech- anisms in remote sensing big data using xai

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.648487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.215845Z digest=sha256:1c1293f0b3156a6c0f256c1e28299a40993703b310abefde257bbfea346f642f

Observation ad9fce89-40e2-408d-a07f-17520ca1cfe4 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.639196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.218218Z digest=sha256:0c348d377a064ad6c5fe78a22db927f77e86e8f97de766e282c4f6c0808547c2

Observation 52fce76b-5efc-4398-a560-1fd2926a21f2 · outbound

This paper cites A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:51.314658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.220736Z digest=sha256:090d72a2591aa3349000eee239725d068b234fe407761d1ccd6ca49b4163b8e6

Observation eec7c4a0-b5c2-4e0c-9742-c6425111f476 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A dataset of clinically generated visual questions and answers about radiology images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.629812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.223538Z digest=sha256:9bb817a4f36ec5e39102553921cdec9b6402eb0ebe21be1851f098ee1941c03e

Observation ad66af8a-1587-48ad-bde3-2ae49004b792 · outbound

This paper cites Blip: Bootstrap- ping language-image pre-training for unified vision-language understanding and generation.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Blip: Bootstrap- ping language-image pre-training for unified vision-language understanding and generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.620877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.225869Z digest=sha256:d72085be4319e39bf982c873741dc64105ac33dfad4367ae18f2ffd0f18204aa

Observation 9ceb7cfb-2943-4faf-b52d-6e4488ce4327 · outbound

This paper cites Contrastive pre-training and representation distillation for medical visual question answering based on radiology images.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Contrastive pre-training and representation distillation for medical visual question answering based on radiology images

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.611629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.228230Z digest=sha256:7d83a6553fe606b4cac7ea65496e3a21ee05eafcb6d8ddf571016a7d078c30be

Observation 056a0cdb-ad1e-4a66-b930-761fe3d5ccc5 · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical vi- sual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Slake: A semantically-labeled knowledge-enhanced dataset for medical vi- sual question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.602511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.230666Z digest=sha256:0cb41532ca008c3222b9bdbbe6792c5e7cafc9e4c654497744f1c864f8c49878

Observation 96e95a63-8207-4aca-8982-c9ecab8c1d1f · outbound

This paper cites Improved base- lines with visual instruction tuning.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Improved base- lines with visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.592236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.232938Z digest=sha256:8aab9d6ad27b5de4e2bab3c10e853231bcb42a87236c565bd3b84dd46d7bdf29

Observation 07196891-9335-457c-8edd-c9ce5cd95e30 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.235217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.235217Z digest=sha256:0812c01a5e54bae0b789de0268a7be8be59e5e5f167ced99f7dbf69212dc832c

Observation ec70c0d9-9ccf-41fc-883a-365e7fb63635 · outbound

This paper cites A survey of efficient fine-tuning methods for vision-language models—prompt and adapter.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A survey of efficient fine-tuning methods for vision-language models—prompt and adapter

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.575920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.237728Z digest=sha256:0f467152ce44577da94da9fdd2135b955b348be8ff86b393d4de94831e046126

Observation a61747b3-7918-4215-8f59-8b2ab5cda85c · outbound

This paper cites Multi-modal concept align- ment pre-training for generative medical visual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Multi-modal concept align- ment pre-training for generative medical visual question answering

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T15:57:51.271628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.239896Z digest=sha256:73da6062ad1f4934714e7baab42beba9514247e4d8a75fcc72e223ed0959a3da

Observation 952f75b6-7fa7-4b3e-90e0-63aea0e300d9 · outbound

This paper cites VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.242123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.242123Z digest=sha256:6bc1ae939ee05674cceac6657806e4a4e91a3444f88fc7e239c0911ceb684ca3

Observation 2b306bd7-60f4-46a7-abd9-7848f0f9f7fe · outbound

This paper cites Medical visual question answering via conditional reasoning.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Medical visual question answering via conditional reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.565319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.244796Z digest=sha256:a681fe3d7a20d7c029922f8aebd788298a6295c37ba70c227bc19b79a8d9aacb

Observation ef7d4920-9ac2-4e48-bbe3-717594c224ec · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 699

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.672540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.201059Z digest=sha256:d44b7ebd6f59dd026185ebe565d377b86bd37c0a212c8bb05f7e7244f5040a28

Observation 9f4fee1c-b771-45d2-8480-77092804978e · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 2023

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.473870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.203586Z digest=sha256:15fe33ff0ad1a7578f20600118f42932e52317f458faf047a4364b2050794a72

Observation 031d157a-5c45-47c9-a918-6be118d30dd0 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 2024

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.556209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:57:51.190555Z digest=sha256:bb7673ab66516b75dd8fc3d95493dc8ecedc4c937aef137b415b045aa6ad36a5

Pith citing papers

Observation db7da7da-3c37-4341-aca2-c6eee986dffc · inbound

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA cites this paper.

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T16:57:31.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:57:31.114754Z digest=sha256:e6c29e7ca2d30f1f3d27d8835459b14ac00203d567e00f897cf14966761a1238