Pith. sign in

Paper Citation Record · LEDGER

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 1 of 1 outbound references and 12 inbound Pith citation observations for arXiv:2508.10552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10552 v1

Coverage vector

measured 1 of 1 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:25:06.520011Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:47:44.047025Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.274843Z

Reference resolution

1 of 1 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8907c983-dd9a-4091-a5e0-15aeec607250 · outbound

This paper cites When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models.

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:25:06.520011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:25:06.520011Z digest=sha256:e6535d311216811762db02e85277b612d530f1273f20140d993d8ed8624f07c1

Pith citing papers

Observation 8907c983-dd9a-4091-a5e0-15aeec607250 · inbound

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models cites this paper.

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:25:06.520011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:25:06.520011Z digest=sha256:e6535d311216811762db02e85277b612d530f1273f20140d993d8ed8624f07c1

Observation 6afbbff3-4ac5-4c18-a358-39f24d832a0d · inbound

Token-Efficient Multimodal Reasoning via Image Prompt Packaging cites this paper.

Token-Efficient Multimodal Reasoning via Image Prompt Packaging When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.621054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T21:52:25.713970Z digest=sha256:e069ef524264cd0b4288805ff2097c3c5a559b5e7d4c265c9cf9eeea0f341d48

Observation e203d505-c702-443f-9dc0-1889ae26d7f0 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.771486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:bc246541e53d169893454184168f4a8457cdb5b937fe7a068de65e5a7a706c63

Observation feadaba6-dedb-41a3-8d0e-df473df2b503 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.190847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:02eacd25271532afdf4da155dffc3be3d6b28e50fcc4bd098669457ee57db5e7

Observation 4822d922-35d1-4da0-bc30-a3e6167eb3a4 · inbound

Information Router for Mitigating Modality Dominance in Vision-Language Models cites this paper.

Information Router for Mitigating Modality Dominance in Vision-Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.163886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:d4a91ce6546ff1f8b18fdf6c9e3ca102aa00c431081a526974ee417b899f28a7

Observation 1e21e9b8-db5c-4a07-b01d-92667dfef6aa · inbound

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment cites this paper.

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.568028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T21:56:38.004300Z digest=sha256:93d0bc8e12b89400481c9a89758c925f6149e04aca43cd2fcbe61dfc5b7d3853

Observation 0557267d-94f2-42bf-8900-6dedf3e3b4ab · inbound

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation cites this paper.

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.940714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T12:04:50.483329Z digest=sha256:0343557d22a1c18883adf7f427dd995fa0b7b4545384ceb6d56c1055f549ba4d

Observation d88d3522-0088-462e-860d-1a2fb0ac740a · inbound

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs cites this paper.

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:39:24.641008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T19:36:40.901674Z digest=sha256:7f0617dbf0cae8cf8e4ec5e73af099624e24ffae84ed9d942a4c9da06a14b71c

Observation 108ed84f-519d-4932-99c9-927fa4411a9b · inbound

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models cites this paper.

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.276916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:28:54.731659Z digest=sha256:a042527162bdc1294f18be12c0390f798315f537959e4ec085b6e7b3aa4930e5

Observation 3743324e-f58d-479e-9d1f-e5a99017c5e8 · inbound

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection cites this paper.

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T07:04:04.812401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:04:04.812401Z digest=sha256:e8ff0152daf85d326847cc248ed87177bd903a729620ae096c25d8fc8f1e2972

Observation 1292bf7c-cbee-4843-8715-43181321eafc · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:38.657616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:38.657616Z digest=sha256:594e54263d89204a00049c89ad4ab8739cdd09dfd3943991144d3b2157d2b3f9

Observation fe3929af-e1b9-4de8-a0b4-6ce2f51ee134 · inbound

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination cites this paper.

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:47:44.047025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:47:44.047025Z digest=sha256:cb0c82d9cca9910261cfeaecc0506cfb990702555e7dd4dae9c0034f25a3a88b