Pith. sign in

Paper Citation Record · LEDGER

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2501.12206.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12206 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:28:42.275751Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:44:54.981867Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:34:01.801514Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cebe75e3-5d93-4e2c-a9d4-cb9422b86fc3 · outbound

This paper cites Agla: Mitigating object hallucinations in large vision- language models with assembly of global and local attention,.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Agla: Mitigating object hallucinations in large vision- language models with assembly of global and local attention,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:44.013954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.682350Z digest=sha256:6b719ded89a746f7aac02ac80dcec8930036df6b04888dcfec33e76998aa3e9b

Observation b03c8fa3-eaaa-4928-b31e-5721df12e862 · outbound

This paper cites Nikolopou- los, Hans Vandierendonck, Deepu John, and Bo Ji.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Nikolopou- los, Hans Vandierendonck, Deepu John, and Bo Ji

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.994796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.688182Z digest=sha256:60d766bc4fec31de20de062ecbb22b541979ecc1896b4ca04482b11dea5fafdb

Observation 6036a9a4-08fc-439e-9e78-a8ccd741d871 · outbound

This paper cites Qwen Technical Report.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:41.693418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:41.693418Z digest=sha256:08331aae9e9ad2177a56fc6932ab994ca3c8251017aef518b6bff4ff736ae19e

Observation b9e054b2-8ecd-4667-8a6d-94ff044c3fcd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:41.699528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:41.699528Z digest=sha256:1c9b0b42ba845ba7542f61d97da96fbad9835b01b8433e26cf33cabbcf9e9e91

Observation ecc5df0d-b8da-4357-a1fd-8e86a9138651 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.972223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.737287Z digest=sha256:980f1e2a14d6976e22f0820a065ed8ca064bca540c1f1565e89bd37cfb9f3404

Observation a8b4e75c-16e2-4b74-ae90-6c3fcb0dd92a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:41.874186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:41.874186Z digest=sha256:8b064b09ba975b1f62d3c93c17d4584695fcc9c5438d017676bd2d516cff609c

Observation 9bbefe5c-b46f-4afb-a089-f07b530e149c · outbound

This paper cites Vision transformers need registers, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Vision transformers need registers, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.917685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.895226Z digest=sha256:4d9a49b4ae5a6e41448c9fd80a158940f1469d9235a40477524dbc2c77bb1814

Observation c6275b5e-986c-42db-8241-63e895ad78c3 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, 2021.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model An image is worth 16x16 words: Transformers for image recognition at scale, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.721770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.938134Z digest=sha256:8f78d0b13b0e4200776f7f829f388f0e2d0a15feccff90d18fefd11159dfc1ab

Observation c826dfec-a4a3-4aef-b2cb-89d898e6a202 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Multi-modal hal- lucination control by visual information grounding, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.669896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.978123Z digest=sha256:d83c1bdac9996a1651e2f866141f3fa242ec05cf044aede280281323473d0a89

Observation ecf3c9d3-fd65-4e4c-aa05-f955cc1b0e7a · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:41.984011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:41.984011Z digest=sha256:3d922080709615ef6cf1116b94313728bd2a9b9c13e760a5ab753060b58cd40f

Observation 87fffa6a-56c4-4da3-a662-aa0ae5ed28f8 · outbound

This paper cites Cogagent: A visual lan- guage model for gui agents.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Cogagent: A visual lan- guage model for gui agents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.652632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.990211Z digest=sha256:08cff34f64d859aa2f06db549a897a6e5271d78d79f86adf3b93a22105bf426c

Observation 3b17fd8d-0cfa-4ed7-abac-91afe3a32851 · outbound

This paper cites A survey on evaluation of multimodal large language models, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model A survey on evaluation of multimodal large language models, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:41.995721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:41.995721Z digest=sha256:6feb379d185bffd03dc033204ea3bde4b47e302e58eccc4757c879979a238423

Observation c881b9e8-5aee-47a1-bd0e-54d4f3633b19 · outbound

This paper cites Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.001774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.001774Z digest=sha256:b9ec9c89040df671d6017b3df5cf9e55b3493999fb8fea8406806cc700ed8beb

Observation f8fc29ef-cc3d-4e1f-bcf3-966aa8fdf8c7 · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.614059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.007830Z digest=sha256:509b51dd01c76abfb3967ef138a9bc8aa369ccda5463748790d36cb2f1d69db3

Observation df7ad794-dd1c-45bb-a9a1-05a52ef16438 · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.596824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.013537Z digest=sha256:59883ae5ea5f35e7f9b77bb55610d7c90c4b64bbec8f606424b5342c8248a5f3

Observation 17e08fe4-9d2d-49ab-b61e-da5919140776 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.019215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.019215Z digest=sha256:45392da7c0a11c70381c66149018d8de134076826e53ecc1f822bb527123d5ea

Observation ded805de-481a-4317-a48b-b0be529cc5be · outbound

This paper cites Code: Contrasting self-generated description to combat hal- lucination in large multi-modal models, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Code: Contrasting self-generated description to combat hal- lucination in large multi-modal models, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.567529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.025397Z digest=sha256:3ab062ffe80781f8f9068c7eba3bf051b0d0f23741ebce37efebb978c6d336e6

Observation 16919685-580a-4978-b028-0a81cdea7ffd · outbound

This paper cites Building and better understanding vision- language models: insights and future directions, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Building and better understanding vision- language models: insights and future directions, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.452924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.030398Z digest=sha256:d71821bc0664859557cf5bd904a260642af3d167bfc740b45a741689344916fb

Observation b824524f-581f-43b2-a4ae-95953f9b0cd2 · outbound

This paper cites Reference-free hallucination de- tection for large vision-language models, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Reference-free hallucination de- tection for large vision-language models, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.320198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.036040Z digest=sha256:09127cfc3f02b850054b1294c830f79f97aa545f648db1ac32bfa516f399ff88

Observation 81768820-ce7f-4443-94cc-71e13ddb8891 · outbound

This paper cites Vlm-eval: A general evaluation on video large language models, 2023.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Vlm-eval: A general evaluation on video large language models, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.176831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.041792Z digest=sha256:c4fd403d248cd4a412fed56e5cef7ff2a5292e8e3a3c910385ad45b5a2b7dad6

Observation 269d5e76-cc37-408a-b3e5-0e827378708e · outbound

This paper cites Evaluating object hallucination in large vision-language models, 2023.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Evaluating object hallucination in large vision-language models, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.161168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.047052Z digest=sha256:398cce8ac55d605097d806ce4e35b2c5cd1100a1a0998dcecd782894533c075b

Observation 6c514fb9-1a90-4ae0-bd2f-629d00a3821a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.052071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.052071Z digest=sha256:d19093969403ab87dd97da8f5b30f2cec695f42584b9587496f9e902707caf17

Observation a99d5419-917f-41d6-b745-d4b5133f2125 · outbound

This paper cites Gpt- 4 enhanced multimodal grounding for autonomous driving: Leveraging cross-modal attention with large language mod- els.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Gpt- 4 enhanced multimodal grounding for autonomous driving: Leveraging cross-modal attention with large language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.144719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.057574Z digest=sha256:2b62a00ceaa7b4adfcca18ae1320501580d1580fb3b948108987f5d277415be6

Observation 3a5c302d-b6e8-4d6f-a00c-7a4e0e22ad90 · outbound

This paper cites Microsoft coco: Common objects in context.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.063233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.063233Z digest=sha256:ec0f742b0fee54dfdcba91890bf9b2e7a149141f6feb78961016643e1679c96d

Observation 62026cff-aadf-4dc3-8cee-a70033b396ff · outbound

This paper cites Visual instruction tuning.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.117189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.068072Z digest=sha256:3e5f4a17f052f709886ee926c1311a40437c86e512f52dfcc4b25d2f5bf1af94

Observation 057b4ae5-1c0b-4f76-b895-e19a0cc7fbb5 · outbound

This paper cites Visual instruction tuning.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.072757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.072757Z digest=sha256:8cd4d7acea20d12fab1b9b87a0d87047aa2749366829a829ba31ed328fca5c3a

Observation 45101b8e-64f7-4863-8dc9-6ff86868437d · outbound

This paper cites A survey on hallucination in large vision-language models,.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model A survey on hallucination in large vision-language models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.078333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.078333Z digest=sha256:16ba5018b8156c77efe79a07584d82e77df4f777406552afb31901f9b22881a8

Observation 3c1a433d-e008-4f0b-b33a-998f8934cb7d · outbound

This paper cites Paying more atten- tion to image: A training-free method for alleviating halluci- nation in lvlms, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Paying more atten- tion to image: A training-free method for alleviating halluci- nation in lvlms, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.078577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.083660Z digest=sha256:454ccc513b4ac505c5be2752984b66da8387fe23cbb94d6947c17d8a93b1aec2

Observation 3df155bd-5f2a-4dda-9762-67548722f339 · outbound

This paper cites Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.088812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.088812Z digest=sha256:32af35efce3638cf5ef7a4e7a42769f763619a368f0531f1ffcd2669b054ffeb

Observation 7fe2a868-7d6f-4edc-8553-eaf011d25820 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.094407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.094407Z digest=sha256:516240972a3c31194f8b896b307866a8cc926a1c7b2cc04e00f1ebb2d7cec786

Observation 784ff819-22cd-4800-ae7c-13a03565c709 · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.062004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.113158Z digest=sha256:672e4d62ac5a6821ff72c6cd1d7e8c3fdfb16f09f8c439bfe64e4e6451c43b14

Observation 0fb71868-bf1b-48c6-afbe-7e6992b059ed · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.133517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.133517Z digest=sha256:68286a8b29599237f6d0124781abbb114a447e31ea12a6000dcc5591473a75c9

Observation 2204cb83-7859-4382-bb79-c3226c277391 · outbound

This paper cites Object Hallucination in Image Captioning.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Object Hallucination in Image Captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.155224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.155224Z digest=sha256:1ccde77645d46d7ab6898bb8cdaf94bddb0dc7ef427d0bca686520d5e5f960be

Observation 181d5f5f-f57e-47c4-a8ee-4ce6e1ce459d · outbound

This paper cites Arık, and Tomas Pfister.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Arık, and Tomas Pfister

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.029973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.208042Z digest=sha256:87b09fbe4f44e102a5dd531d0a27d9027efb8834f84dbb9be079bbd7c8d90d81

Observation 1bf4b930-e462-4796-bb99-0cc3d3b5724b · outbound

This paper cites Visual text meets low-level vision: A comprehen- sive survey on visual text processing, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual text meets low-level vision: A comprehen- sive survey on visual text processing, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.010949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.213290Z digest=sha256:73d8c0598a35789297016c0681faa7caca784ac32c3da4647fcc2a7c076bd971

Observation 5b3667b1-1698-4f56-aabb-15a579959005 · outbound

This paper cites Mitigating entity-level hallu- cination in large language models, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Mitigating entity-level hallu- cination in large language models, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:42.992611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.218200Z digest=sha256:0d43cd71432b4cf3edb8c371df286680a13eddcccba98fd3a459968da2643960

Observation b7b701dd-9b22-4faa-a669-2733cd166fbc · outbound

This paper cites Evaluation and Analysis of Hallucination in Large Vision-Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Evaluation and Analysis of Hallucination in Large Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.223116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.223116Z digest=sha256:009f20b915c765fcbe64cf833a464c309b99a2d3887471cd359074f122bca774

Observation 13b72026-dccb-4dec-ab73-2229e4ff33ff · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks, 2023.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visionllm: Large language model is also an open-ended decoder for vision-centric tasks, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:42.972940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.228495Z digest=sha256:b9d582a9dd5c961d37e3039853d760e6db36332bfd356418ee7155b633e8861f

Observation bfbb2228-5f6f-48df-968f-3b24e1cefcf1 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.233604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.233604Z digest=sha256:66370d4701cff35c17cd5de56ef37514fe7e6205ef9adefad39b31a7c35a90de

Observation c8cf76d3-9626-476c-805f-9332bd5a1df1 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.239225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.239225Z digest=sha256:2f568379a718073598b6cde1199f7aeb3976a014d0bd99a1bc9499ef1dc3acb6

Observation a692ded4-65dd-4451-a780-7b0152e20d4d · outbound

This paper cites A Survey on Multimodal Large Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model A Survey on Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.244622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.244622Z digest=sha256:d145296537d8e8165e45e2cffc2b02ddb3ed94cd688a0333e62517ba24b18f62

Observation ad07a1d5-f159-43d2-afb0-feda7fa30c31 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:42.765073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.250215Z digest=sha256:ada306ce1626a17502339fac72620ff737cbc4314255270a23b0832cc6d6660c

Observation ba2d2d30-3afc-40e3-9c57-92dae7f0463d · outbound

This paper cites Sigmoid loss for language image pre-training.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Sigmoid loss for language image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:42.677767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.256496Z digest=sha256:965f17952c0c25d0598e456c195251724a5b8ab89e122d95fe5150b5c2fb0895

Observation 6d6e63ca-ed06-4655-a3c2-e71fcbd7a45a · outbound

This paper cites Vision-language models for vision tasks: A survey.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Vision-language models for vision tasks: A survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.260931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.260931Z digest=sha256:4b2a1ecfc8af8fe8acc131c503e7b42f19cfac799d2a726351e0b16ec03c0f02

Observation 868ac95e-4c65-42cb-8a23-46aacd7d3f1e · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model OPT: Open Pre-trained Transformer Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.265647Z digest=sha256:2df4ebc898f70f308a458cd35bda7ad543046e8513a6b3231a5830124b8f0da0

Observation 7a7b035a-aa99-4f35-86c9-ec5c70dc255d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.270697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.270697Z digest=sha256:f9801df8d6d08040f862706f00502e9eda7f99f845192705f41d5f8d628869ef

Observation cbf2c9d0-5529-413c-8428-3b77b3aca602 · outbound

This paper cites Ibd: Alleviating hallucinations in large vision- language models via image-biased decoding, 2024.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Ibd: Alleviating hallucinations in large vision- language models via image-biased decoding, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:42.572620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:42.275751Z digest=sha256:48db723250adaf6a2ca52af64f76d4f9097035d9a5b9f2b0c3aaec950b1b480b

Observation a24d36d2-2478-4698-b7c6-c8fbed12cf15 · outbound

This paper cites org/blog/2023-03-30-vicuna, 3(5),.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model org/blog/2023-03-30-vicuna, 3(5),

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:28:43.951583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T17:28:41.784398Z digest=sha256:2c3f04e4856ee0593db0467d92c879ee791da6facf5b7f24af84f6f392f9418a

Pith citing papers

Observation eff686ce-38f4-432d-871b-66cc90968a7d · inbound

CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:54.981867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:44:54.981867Z digest=sha256:52ac87341730f63ee1cb6dc73e8a38cfecbd010129da167db53c2a8ae13e0ded

Observation 55ad7066-1187-4b0e-98f2-ea60b2651ddd · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.882276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.882276Z digest=sha256:22cc928ce5c1800889409d9d0056f47fa9002982cd2c684c1256ed5208b139f3

Observation 992b2855-7ec2-4c27-94d7-51fccf0ab630 · inbound

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering cites this paper.

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:50:40.327897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T18:05:38.705480Z digest=sha256:8e62f6d6f85c39681e4c574834e252108458f36328b611b5b2593f69d5cf6c20

Observation 5a8682e6-8c52-484c-a40a-7d4c8aa504b5 · inbound

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.803714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:28:35.155099Z digest=sha256:ec434b6109724bee0f00eb5a0c76a68579bfc1e4be324cc7672e1958e0dddfb7