Pith. sign in

Paper Citation Record · LEDGER

Mitigating Image Captioning Hallucinations in Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2505.03420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03420 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:56:59.651886Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-14T19:18:41.987533Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T19:19:23.960000Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9192742d-3976-4fbe-ab73-e28885520b77 · outbound

This paper cites Deep multimodal data fusion,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Deep multimodal data fusion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:57:00.038161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.531662Z digest=sha256:82a713eb1d963e3a92b9ea581d0c506f78df1dfd1c31f73ada6015ecaf00553c

Observation d5aa29a2-3f2a-4b68-8404-8675e139e461 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Mitigating Image Captioning Hallucinations in Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.536994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.536994Z digest=sha256:c8048415adc50b0c19184992cef52bc74a9468b5386b04bfabb387edbd8a1e2b

Observation 34d83191-f3c0-42d5-9360-26c34cb07cac · outbound

This paper cites Understanding and im- proving layer normalization,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Understanding and im- proving layer normalization,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:57:00.021551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.541897Z digest=sha256:373c99f4c676042ffab7d7d329a005046ff3c1484720179e42ec7a7897f665ad

Observation 9b92461d-49ae-4512-b422-22cb60193096 · outbound

This paper cites Miti- gating object hallucinations in large vision-language models through vi- sual contrastive decoding,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Miti- gating object hallucinations in large vision-language models through vi- sual contrastive decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:57:00.005904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.546613Z digest=sha256:09d073ba298cd8224dc12d34fb9e230fb0410f1f630c3d1d46c55c49fbc3eb4c

Observation 6f444e48-278a-4ee3-8657-e0798669ec31 · outbound

This paper cites Chatgpt (mar 14 version),.

Mitigating Image Captioning Hallucinations in Vision-Language Models Chatgpt (mar 14 version),

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.990042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.551388Z digest=sha256:4d028711e1b95b8e72cea25e41836c2a0d08eaf589d0bdd756a103c45c632af8

Observation 3e9917c7-ff18-47a3-9391-07295bf406ed · outbound

This paper cites Gemini (dec 1 version),.

Mitigating Image Captioning Hallucinations in Vision-Language Models Gemini (dec 1 version),

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.974569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.556196Z digest=sha256:3e2b723944c578804fa71bedcf7d796b58c3d1372b0d223ba5d59c775b86ab45

Observation e31d3e08-dcf8-4f0d-8338-be19e0c43ba9 · outbound

This paper cites Visual instruction tuning,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Visual instruction tuning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.561176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.561176Z digest=sha256:f35f2cf1258d89a49821f8ea3f4a5cc875632ce04f9644a27b1b4de582158e48

Observation 29cbdcbe-db29-490d-80b9-d7b74546e00e · outbound

This paper cites Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey.

Mitigating Image Captioning Hallucinations in Vision-Language Models Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.565756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.565756Z digest=sha256:15e63d3fec93ac38f0e3249e288e8c391afbd13363832f7caf8b91b18a435a17

Observation eb8593ea-9e9f-4365-948c-56daf9349773 · outbound

This paper cites Checkguard: Advancing stolen check detection with a cross-modal image-text bench- mark dataset,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Checkguard: Advancing stolen check detection with a cross-modal image-text bench- mark dataset,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.949734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.570836Z digest=sha256:313e63ef90a214d180c7667457182e7c24d760386536233f11e2b470397ab031

Observation 3706723b-57a6-49dd-80e7-521b83e9937f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.575309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.575309Z digest=sha256:7109bcf2266b0eb8b6003aa076b4d5b7b4d9657b577cdf025d9705dd231a5a37

Observation 7985b914-b4fc-415c-8710-8e6cbcf52fc6 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Mitigating Image Captioning Hallucinations in Vision-Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.580122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.580122Z digest=sha256:fadc292818efda66ba6d1e4427e42e088dc91d5e078f124fdf658c0d40763007

Observation c3f7cc08-52dd-4edc-b6af-5d73a12415d9 · outbound

This paper cites Prismer: A Vision-Language Model with Multi-Task Experts.

Mitigating Image Captioning Hallucinations in Vision-Language Models Prismer: A Vision-Language Model with Multi-Task Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.584659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.584659Z digest=sha256:09deedbf3784fb710a517c459fcf7cbb36ed95294b374684046ba474df736c72

Observation db143b52-2a62-4b9e-9754-a5f7f58e7ccf · outbound

This paper cites Detecting and Evaluating Medical Hallucinations in Large Vision Language Models.

Mitigating Image Captioning Hallucinations in Vision-Language Models Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.589473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.589473Z digest=sha256:a59251b26118a72fa87f864b34428f224239553ccbf030fbe4de2792fb8104b3

Observation 87f8fbe4-685e-4e72-9017-5a9c1cf87f6f · outbound

This paper cites Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models.

Mitigating Image Captioning Hallucinations in Vision-Language Models Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.594270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.594270Z digest=sha256:c1b762667b9c9b1ba918c1fa83fbb3d2999fe0bd7e8f6dc2f60084fe93af5887

Observation 41d52a2c-e9a5-4aab-8fd0-8d98ff9a5318 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Mitigating Image Captioning Hallucinations in Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.599232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.599232Z digest=sha256:f5894f25a56fad596f6757d1f6d40c6aaab94b5203addbc45f0b1a73ba5cad42

Observation da6123eb-4ac2-42c4-ba18-76a4564b578f · outbound

This paper cites Mitigating large vision- language model hallucination at post-hoc via multi-agent system,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Mitigating large vision- language model hallucination at post-hoc via multi-agent system,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.925103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.604154Z digest=sha256:6c62ed0a19f5d18036898e4c0e003ce602264a82d005d59b60b8a5694a75efeb

Observation 68806463-f33a-43b8-b587-9650d5077bad · outbound

This paper cites Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback.

Mitigating Image Captioning Hallucinations in Vision-Language Models Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.608837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.608837Z digest=sha256:47c1ea657fe915103a6dd18752601da5e70380af4dfe6cf197560d1205d145b8

Observation 51cf1c56-bb8c-457e-bb1a-9ea54c16cfda · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

Mitigating Image Captioning Hallucinations in Vision-Language Models Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.613666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.613666Z digest=sha256:5344c5479da4a39941fee5924119aeb86e8afcf7669f20e681a2addcdab280d4

Observation aa4fce13-c512-4e9a-96e8-73f9169dcb6e · outbound

This paper cites Best-first beam search,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Best-first beam search,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.908868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.618927Z digest=sha256:fb70f96350e2a5add2926db8cd11bd9e5d9d7acb9ee7ee7a37caaee0a1477189

Observation 0221d03e-0330-4ba7-8dd9-7e905a4a4143 · outbound

This paper cites Policy gradi- ent methods for reinforcement learning with function approximation,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Policy gradi- ent methods for reinforcement learning with function approximation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.623670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.623670Z digest=sha256:d2512e8643d7b3b3a07973d5dfc65789aa4a1c86a2372d45b16fc69a438c6076

Observation 9f5770c2-cda7-4918-aa57-c8586404d700 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.628180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.628180Z digest=sha256:033282d7685937523350e1e2576f82b4028927116b430ef240974e60ea36a5d8

Observation c0673f22-919b-43f4-afcb-e33adbc8a4b1 · outbound

This paper cites Attention is all you need,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Attention is all you need,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.632912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.632912Z digest=sha256:35f2f539558c7504e7fffa44c6bacc81f4be93b2504dd5df1589c27075c5f9b3

Observation 001a194b-d610-423b-950a-c12364efba95 · outbound

This paper cites Facenet: A unified embed- ding for face recognition and clustering,.

Mitigating Image Captioning Hallucinations in Vision-Language Models Facenet: A unified embed- ding for face recognition and clustering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.637425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.637425Z digest=sha256:494f66fefbbdd33abee7ee4e02a8be7148ca803d7bfa3694b6e786df3c9f6d67

Observation 21735214-d299-4dc1-852b-bc675ad92b1a · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

Mitigating Image Captioning Hallucinations in Vision-Language Models AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.641620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.641620Z digest=sha256:e0e002e45fc9370d5135dc694659593b66a35f59fa526326c38e8ac63feace95

Observation 2f25cb7e-3f8f-464e-8e3b-b500674241fd · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

Mitigating Image Captioning Hallucinations in Vision-Language Models From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.647249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.647249Z digest=sha256:5b4122ab12acdb8be3b9ad3a9d1ef72d18342f7f3d086cb87e826671210baf20

Observation 2b9f4120-1c75-4397-9b33-e92dd12e1989 · outbound

This paper cites Llama 3.2 multimodal (version 2023),.

Mitigating Image Captioning Hallucinations in Vision-Language Models Llama 3.2 multimodal (version 2023),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:56:59.855141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:56:59.651886Z digest=sha256:dacaceb4ff7265eb6e39f42c39adbb0279d178e854b6b50463b2e2a3a4c870ce

Pith citing papers

Observation 1cc9f4ee-fb94-4a71-8c7b-05835ce49510 · inbound

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models cites this paper.

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models Mitigating Image Captioning Hallucinations in Vision-Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:19:23.965381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T19:18:41.987533Z digest=sha256:46b45739ba9303bef55e0552c29280057676687ab451c382ada11e04503c4155