Pith. sign in

Paper Citation Record · LEDGER

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.12787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.12787 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:02:51.368119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.441655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fc46da7-1f83-493f-b96a-7791f73734d0 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.231593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:44eaf2f0bd8b796688546b958ff782374fcaafa518be20f91735f4e1e3525783

Observation 459e9243-efca-474f-a464-1109ebd917f7 · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.368119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.368119Z digest=sha256:b1a8ebbae53bd04f9b78847a643fd287ef9cc78cab7776eb0f2a20b36ba8d3a1

Observation a9af95a0-7df2-4c73-8114-69135c20ee99 · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:22.181322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:22.181322Z digest=sha256:a811015bf53b97dc824ad8f88b526681db25795071c230df772d07dc87546d65

Observation 430d8b40-28c3-485f-a683-4845b1a48d5a · inbound

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models cites this paper.

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.761053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.761053Z digest=sha256:c8ee4973ec58f695bb028765af92478e3a742fb2e743cedad6d03becef0ea854

Observation 833c9254-754e-4661-92f0-21e24bb4c782 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.280415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:af25bafc70961c128d1da428c97dbf39965eee020703460c8a66f1358c059ebc

Observation 70a1b00b-d496-44ea-9da8-771c397e9079 · inbound

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models cites this paper.

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:06.449230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:06.449230Z digest=sha256:d8ff31997c21b1ec7c1c409cc1d65084c42a0eccef0c54c45bdf8399a54d1d11

Observation 3b48526d-9f14-44c8-a57c-2dc57b7492f2 · inbound

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination cites this paper.

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:22:19.020115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:22:19.020115Z digest=sha256:c8b1723ac8c856564d5e4fa01baf06618c2ceafa593470698acf7d8c563617da

Observation cb2f6a73-a98a-4d9e-8f62-d4c276e99bd9 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.026513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:9c1405c670b229268d9919b3e8abfdb447c76195ace8bd6ef7d52525c3403a7d

Observation 3dba48bc-6a31-4745-8e45-20087f00ad48 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.174926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:831ed3065c992f7c17d449af383851d4bca70220930856f9346ab9c22e76e34b

Observation af26b459-ad70-4ae0-926e-665a5b91c5a1 · inbound

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation cites this paper.

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.948394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:04:50.483329Z digest=sha256:981c439402239c7e2c918abcda83907b446d90133c7aa7a99c23f3675ac0e468

Observation 4dff1acc-d9ae-4ea4-a3f6-1993317d2aba · inbound

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models cites this paper.

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.236037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T07:28:54.731659Z digest=sha256:e13137f3cc860c03058af27b8a9eeabea277f7f19782ed0b944ef11170778bf4

Observation 3a63a04c-2f71-4ad5-81c3-040a4a07d6ed · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.444991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:9b5c8e9e613c5527ffa10abe80ca71c39b30eba3c17172a6370db4acf7347e4f

Observation 4195d20a-1f75-4cf4-8761-4f07df3a19d3 · inbound

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding cites this paper.

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:02.491723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T01:18:13.657975Z digest=sha256:d0e40f103aea1686e8bf0e27266d9d29e34386fc3e3c4fabb67a858d7e64d2d5