Pith. sign in

Paper Citation Record · LEDGER

Are MLMs Trapped in the Visual Room?

As of 18 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2505.23272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23272 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:27.130704Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T08:20:56.561166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T08:22:10.936561Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10197d36-6930-412d-943b-607fa623afb5 · outbound

This paper cites Advances in neural information processing systems35, 23716– 23736 (2022).

Are MLMs Trapped in the Visual Room? Advances in neural information processing systems35, 23716– 23736 (2022)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:28.081107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:54:24.966379Z digest=sha256:d276b62ccd70715336c7bc96e50cd2ab99d17ed1b89727e6e0f9cf2f3fd46ecb

Observation 6a90386f-f6b1-4cf6-9e9f-6d65c061194c · outbound

This paper cites Qwen2.5-VL Technical Report.

Are MLMs Trapped in the Visual Room? Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.054856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.054856Z digest=sha256:c3887376887fb25a45b072fd479eb7d9765716e25a8a58525d142b962627cf73

Observation 3b4ddf6d-1c1e-4d6f-9e42-8b6807b3a149 · outbound

This paper cites Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper).

Are MLMs Trapped in the Visual Room? Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.159760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.159760Z digest=sha256:554ad04bf752ed74769bbdac719d78fbc666fbaa2a26887ee307d4df4f9f5246

Observation 5139d084-0e47-43d3-a0c9-ec38f6bb92da · outbound

This paper cites Advances in Neural Information Processing Systems37, 110805–110853 (2024).

Are MLMs Trapped in the Visual Room? Advances in Neural Information Processing Systems37, 110805–110853 (2024)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.926495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:54:25.251838Z digest=sha256:92d768bc5fbbe2fdad019ff9617bd5f8fd6696fc244766686b2fcd59325962aa

Observation 19863e85-f32f-4496-a8aa-4b9b282be700 · outbound

This paper cites an unresolved cited work.

Are MLMs Trapped in the Visual Room? Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.350577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.350577Z digest=sha256:1d0f31cf9648f59a6afd65348249fa876bb9f242ef3139759c7cdf78ee77e525

Observation 59ba3b7d-bae5-4f2d-9f27-a7b0bdcedd10 · outbound

This paper cites GPT-4o System Card.

Are MLMs Trapped in the Visual Room? GPT-4o System Card

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.466576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.466576Z digest=sha256:75eedc791e12769a8b7a1a679c7fcf5818c904cd2c55f13a2298603235141818

Observation 02bc6d2b-074d-4321-8b2e-9ef466e960d9 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Are MLMs Trapped in the Visual Room? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.598847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.598847Z digest=sha256:acff6d1d5aae8b7d9948333fcb45ab5d10a658a57ad96e94ae85c42645172e46

Observation fb71c79b-557a-4685-8682-775a215e124a · outbound

This paper cites In: International conference on machine learning.

Are MLMs Trapped in the Visual Room? In: International conference on machine learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.728615Z digest=sha256:3c45283ee90cabbd7b5245908a8bc4787983ff1ed9ab51f96b75fdf2dc949e4a

Observation 9eb782d3-f035-447f-843a-565f68f4b0ac · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Are MLMs Trapped in the Visual Room? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.856176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.856176Z digest=sha256:138f7c7df3d2ca4fb0bde8c3eab718fc6220eba7664fe6b860dfb37355fa1a36

Observation e29d5e5b-71d6-414e-9023-a1cc80353618 · outbound

This paper cites Generated Knowledge Prompting for Commonsense Reasoning.

Are MLMs Trapped in the Visual Room? Generated Knowledge Prompting for Commonsense Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:25.942571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:25.942571Z digest=sha256:db225263a849f3840cdb5a2fa462bb3bc05d507351bf4adb9fba7538fa8cb063

Observation 0d3d752f-9ee3-401d-9e15-2d0d2aa687e2 · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia.

Are MLMs Trapped in the Visual Room? In: Proceedings of the 32nd ACM International Conference on Multimedia

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.742685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:54:26.039290Z digest=sha256:020e9aa25e4605860690ef4e592ec41970ec6dc21302666f1109dd83340f2edd

Observation d7c945a3-792b-4025-a7e2-1e1504c5284a · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Are MLMs Trapped in the Visual Room? DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.194743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.194743Z digest=sha256:4df5694f24545b1230acaabab2da46827b3e2709a7daa6a1c70c66ea616d0912

Observation 68e865bb-5934-4a6b-ad97-0c9996bb4c98 · outbound

This paper cites MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System.

Are MLMs Trapped in the Visual Room? MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.308818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.308818Z digest=sha256:15ed6527399f2677ea63920464f2271182b94c2d7ce17defbfd0964af409c12c

Observation baee13cc-5c43-43cc-b383-46088c667f8e · outbound

This paper cites In: International conference on machine learning.

Are MLMs Trapped in the Visual Room? In: International conference on machine learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.367430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.367430Z digest=sha256:3c0486e8880c5bd7ea0e5c27e821b09958665cf4c8aae1177b26676c724f600b

Observation 7b9972cb-ffeb-4639-9573-b00babf8f469 · outbound

This paper cites Scholarpedia4(8), 3100 (2009).

Are MLMs Trapped in the Visual Room? Scholarpedia4(8), 3100 (2009)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.597716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:54:26.439968Z digest=sha256:70664ceef99c0a1d521ecb767502c886405ba02bd3f4df6e57010f8ac03c1361

Observation f815efce-317d-4461-a7f1-7c8f898e212d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Are MLMs Trapped in the Visual Room? LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.520235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.520235Z digest=sha256:0def8c2d165cd83e6083758138c53ac30b4ac55634636dcc7cc9d3fdac942535

Observation 7e8d331a-8437-4602-8717-6fa5d7c5c7e3 · outbound

This paper cites Information Fusion 103, 102132 (2024).

Are MLMs Trapped in the Visual Room? Information Fusion 103, 102132 (2024)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.649855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.649855Z digest=sha256:3ba1bb2b44ea01a5f08dc5dda69ec75ac0beea22792b6ca667521fec7070bb54

Observation 00371fb9-2991-4069-9f45-b31c4e4dd40c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Are MLMs Trapped in the Visual Room? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.702049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.702049Z digest=sha256:3f24b18974fe930cbd203ed6bb880f97df987e5583ee28653c94b5c165e17e0e

Observation e8df1fdf-8a55-4fe2-8821-1465ea6aeebf · outbound

This paper cites Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis.

Are MLMs Trapped in the Visual Room? Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.786207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.786207Z digest=sha256:3c62e279f1a95563dccbbb90401cff92d60aa48842230117a4df9610394bf668

Observation 7456e1fd-f150-46b0-9bba-6f41c97a0ce7 · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Are MLMs Trapped in the Visual Room? DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:26.879990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:26.879990Z digest=sha256:4870602c5a303da905f1f9f5f384b08406d176ac692e256b55ffff6e19e297f4

Observation d057ae2d-c8ef-41de-b82c-769ac8f6f1c4 · outbound

This paper cites Advances in Neural Information Processing Systems36, 18794–18805 (2023).

Are MLMs Trapped in the Visual Room? Advances in Neural Information Processing Systems36, 18794–18805 (2023)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:27.444845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:54:26.971368Z digest=sha256:1f287b41e9f02a4f1eea329ac5f960fcd95700f87daf98861316aa34b17d995e

Observation 9279b93b-17e7-469f-a099-4376b71c4232 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Are MLMs Trapped in the Visual Room? Multimodal Chain-of-Thought Reasoning in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:27.034519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:27.034519Z digest=sha256:854d694f45382f03d14d57e3bd3b1ca2977966e322a61875b387ca522d82f42f

Observation 38fa8348-8c66-410d-89c0-fb7767f6dca1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Are MLMs Trapped in the Visual Room? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:27.130704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:27.130704Z digest=sha256:353d1c2abcca5cd98242a83c3c52e9f1311fa0f93895d0c183644f104e7f327e

Pith citing papers

Observation 51b99c23-2f2c-480c-9eaf-133d03052004 · inbound

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection cites this paper.

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection Are MLMs Trapped in the Visual Room?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:10.940154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T08:20:56.561166Z digest=sha256:1168b0f743ac2052dd3433d22647894dce87c22407cc135799f1960d5fac769a