Pith. sign in

Paper Citation Record · LEDGER

Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2410.03659.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03659 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:38:22.063110Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:56.589135Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1aa22ad-0436-4163-8de4-e9a8d6e439aa · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 181

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.878537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.878537Z digest=sha256:7af3018527eb8c483410d74ce987b6f9408a9b67e894ec6c0461b687e0775e79

Observation 7dc98abf-b225-4c39-85c8-175ef4573d84 · inbound

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models cites this paper.

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:54.646826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:54.646826Z digest=sha256:3b30e365c42f294ee648f9cc084fb48012c86747a33b5eff7a3598863999cbfc

Observation f860008b-cb30-4061-a79d-89b0ca0f50af · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:28.298796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:28.298796Z digest=sha256:3ed035d005de8a8f1af59985b2e67ff47ad337a6d8b081582402b7381490ab6a

Observation 47527f2e-8ffb-4862-b043-2ec7340325b1 · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.971423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.971423Z digest=sha256:a7f72f0cd5ffccb37009e38b49fccc07a424a9fbb9ea5285e101c0beb2a54bb6

Observation 9c47deb8-34ed-426f-b729-1e8ce3c23b4d · inbound

How Do Vision-Language Models Process Conflicting Information Across Modalities? cites this paper.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.903561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.903561Z digest=sha256:77b419c860c54c7acb2e2a4ea58b66f31a4e47625075df58025b8e9261157763

Observation 63c38f93-dfe9-4510-9ac0-fee91f0bec99 · inbound

Robust Multimodal Large Language Models Against Modality Conflict cites this paper.

Robust Multimodal Large Language Models Against Modality Conflict Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:35.018852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:35.018852Z digest=sha256:e16b7f60e9cfbe357e8de8e75b0b1a729511efe237951b09a4949bbf4b7d325c

Observation 6cf7bf70-5853-428f-80e8-95f8bf768da5 · inbound

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models cites this paper.

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:12.654015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:12.654015Z digest=sha256:f616048d7e44b3bda4d955fd1a41635fda2c3cfc441416995e25d7431159b253

Observation 5b1684b8-c769-4354-aeb3-4894714bd8a3 · inbound

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation cites this paper.

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:41:26.010367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T10:00:25.846913Z digest=sha256:321a278020abff948a1d0c8a51c94920ac89c1b70272597cc4fb705ec8c07e36

Observation 82b8a289-9b1d-40f3-8d53-a20d3a6238df · inbound

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination cites this paper.

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.503260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T06:25:43.369455Z digest=sha256:b3007c7be82332f3801af4bd3db575778e911af9f122fd4f9a1fb19d54457628

Observation bbd94af2-b3eb-4169-8e70-07eacc036cfe · inbound

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? cites this paper.

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors? Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:27:22.638397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T20:23:18.667677Z digest=sha256:b508bde36b8e41c8ef6cf691efbdd8c037792be93fa03e09fbbb4f61525bc383

Observation b5f96521-0c87-4d8b-ab48-5a3057d25141 · inbound

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models cites this paper.

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.208924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T19:02:12.613002Z digest=sha256:3fd6e136804763708cdae5f8972a5edb47ce9f237c768a39e57c58a6c455ac0a

Observation bb3526a8-6af3-4eac-bf2b-f76977ac22b1 · inbound

MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias cites this paper.

MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.591891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:28:50.021432Z digest=sha256:e71cd9029009e5dc28d1f578a4c96f612a586c63816fa07fda2b95b75d29c558

Observation 924cbbd3-b504-4636-a0fe-85da7e477ac8 · inbound

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models cites this paper.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T10:51:18.853885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:51:18.853885Z digest=sha256:8321abaf2a18a3b6a5da6ac73f12eeb169d5dcbb6e2e1c7be933686ccb839733

Observation 04db053e-d3c9-4ac9-86c3-f893bf014338 · inbound

Linguistic Context Recodes Visual Representations in Vision-Language Models cites this paper.

Linguistic Context Recodes Visual Representations in Vision-Language Models Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:41:48.311245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:41:48.311245Z digest=sha256:0fc732e072af158f62f57086c79873a378dd1194a233a0c05b0c0c69c4e22fe1

Observation 46970202-0f83-470b-8d36-25537a4ee3ce · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:49.171426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:49.171426Z digest=sha256:af6f90cc165f179e10f719f2a977ab8fbf3507f9f126b4d2b7134e4595f913f8

Observation 41026aa6-ae7f-4e62-b832-f97f6220f113 · inbound

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation cites this paper.

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:38:22.063110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:38:22.063110Z digest=sha256:45c1928fdd3e05ee527c59d9989008c04fb836f0431d38d380fd00085735ac53