Pith. sign in

Paper Citation Record · LEDGER

Visual Hallucinations of Multi-modal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.14683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14683 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:19.398932Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.276137Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bf07f7af-9c6a-4294-a552-7cd19cb80e1f · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Visual Hallucinations of Multi-modal Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.817257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:18a4ac0ccf4dfaec3eb1ff5a99d8b729bae8301f4297a22eca74ab369e18290c

Observation 920249d8-c86b-4e6f-b7c4-f0a867405d35 · inbound

ChartLens: Fine-grained Visual Attribution in Charts cites this paper.

ChartLens: Fine-grained Visual Attribution in Charts Visual Hallucinations of Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:19.398932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:19:19.398932Z digest=sha256:41084cabcca907a507a6ad1a0ef4e274db92f704c00daf5b747f3d148dc6b137

Observation 862b44a8-62ee-49d1-b434-f352787d9f27 · inbound

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model cites this paper.

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model Visual Hallucinations of Multi-modal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:57.776021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:57.776021Z digest=sha256:a8d1b69122c72bde5f3e9760aef5e05bc86e8c1d3fcd9383d901f6edc73f9efd

Observation 467d2227-2375-4c1c-98e8-3e3dbcb4cde6 · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Visual Hallucinations of Multi-modal Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.573134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.573134Z digest=sha256:bb4b6de3a539fc3bdaf19c8ba55e3f7698fe081bf038def4b21996153d6b2b18

Observation bb510fe0-5957-4869-8512-70d7d52aad82 · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:44.314634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:44.314634Z digest=sha256:3456f9705a53302bdaf5461639644c96e449549a00a750cf5be254d0e70f581c

Observation 441644d5-d89c-431a-83b2-81a162f614d1 · inbound

ReFrame: Rectification Framework for Image Explaining Architectures cites this paper.

ReFrame: Rectification Framework for Image Explaining Architectures Visual Hallucinations of Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:59.754622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:59.754622Z digest=sha256:7deb56da979af5b535be8eda1fc136c9d3b660acd516e7648bccd5e96b97af01

Observation 490c5c7a-309b-4952-9dd1-f72c7038452c · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual Hallucinations of Multi-modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.374403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.374403Z digest=sha256:38d802eedb33364606778f6a3f0fd44b0a30d55d328e42c81c3f38b844456352

Observation 7c577c89-7a9d-4e85-a3cc-d7e6067ac6e4 · inbound

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs cites this paper.

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:12:23.562660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:12:23.562660Z digest=sha256:db73cb4b843ab1d738c37548af600c5a13793cfbf90b5e6130c403f7dcdcea70

Observation 22d363ff-ec68-4982-b136-9c4ba7dffbba · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Visual Hallucinations of Multi-modal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:44.432144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:44.432144Z digest=sha256:cd231a9eb7c156cae8da787638ddae53336ae9f903e08f9c743cb4f8ffb632b9

Observation ec4e3dde-8ab0-4c98-b4cb-7f4132f7ba17 · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward Visual Hallucinations of Multi-modal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:ba35ac1c222aae2c6c00af15206ac68cc2eacd4740ecdcc2c8733963a5165fa7

Observation d8d4980c-0baf-499f-9a1f-34b238e14fa3 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Visual Hallucinations of Multi-modal Large Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:05.974790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:fbbbd7ae6f093eafd3029a2d0f5b8a5e5cd76caa396bafef9a30c216ca2b5018

Observation 4671d905-cb61-4458-ae74-f0ea239706a4 · inbound

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs cites this paper.

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:05.264219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:55:07.172228Z digest=sha256:3c8a9a2c8ac5bf94f8d54c4a434e2d79d9b8f241be0f4e6d0c95866928fc817d

Observation d80d52b0-69fa-4bf6-b027-362ec319cc4b · inbound

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness cites this paper.

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness Visual Hallucinations of Multi-modal Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.607805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:21:50.254449Z digest=sha256:4d5a711d6acc4652fedee04ef95d251e5184191234e31d539e43bd0ddaf35aee

Observation a6fdc414-622f-4fc3-82fa-f89423275927 · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Visual Hallucinations of Multi-modal Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.581523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:3e67e7157b9c485d46259fab4d4d7f0c4e2125a5429a59d9ec34d01eebba9342

Observation 89f0e3dd-7a66-4585-8667-6330cf9a416a · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation Visual Hallucinations of Multi-modal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:29.278185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:c76cb10fdf3e08822c1aa05725de70bc6f00dc2aa01cd52d0dc22098228f53fb