Pith. sign in

Paper Citation Record · LEDGER

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.05626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05626 v3

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:49.290784Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dace842-e939-43af-b761-3563d2ff32e4 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.692845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.185329Z digest=sha256:269d0ade07e03abacb56acd906237815eb70c7b1adb36aaf62843374d61b450b

Observation 0d8090c1-07b4-4f97-ad1d-4b078e715398 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.679539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.190221Z digest=sha256:b63c7daff5305dd232506a4b23a52a43325c0ba0687802a265f2081a72706558

Observation c9a3fce8-87bf-422a-ba77-fa9d7cacfb40 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.194642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.194642Z digest=sha256:d044fbe0385217598d0866cb8228c6f4fc7d4059ad2ef8883c6ac11bd2d20205

Observation 83cb824a-d8a7-444d-adf8-035a30653275 · outbound

This paper cites From colouring-in to pointillism: revisiting semantic segmentation supervision.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models From colouring-in to pointillism: revisiting semantic segmentation supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.199068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.199068Z digest=sha256:d501ca2b5a3d8c87304f9aa9a67d23b87a3564ff1656844418e8e9dd11322541

Observation c231a48e-3a5c-4215-a6d1-6c23b9219b01 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.204156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.204156Z digest=sha256:1aa8fa9a67344d47ef37217ce1b6c5cab051eb6366354e875ff2e4c934595307

Observation c9d423cd-3234-4f39-bc5c-59a7505f519c · outbound

This paper cites The Llama 3 Herd of Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.208740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.208740Z digest=sha256:07a02d2ada278710f0e5b74ae04e444c384d625880447fd0d9e298e96d3bf814

Observation 2ce9a543-525d-4b16-b5f1-56ab96fb1fc6 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.666781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.214294Z digest=sha256:57f75e2736b0d1251f86441b6aea11af41a756b7b13068ad186d3f05a47ab7e1

Observation e565e41f-649b-4594-8d4e-23f0bb47a7f3 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.654002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.218511Z digest=sha256:f5af18bdfa6a1a967b3e745672d2dfc0682f870597a8a85129f05cbd7f42a586

Observation 058f94bb-38e4-4873-8a6f-cdc5645d5cc8 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.222650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.222650Z digest=sha256:88318bb2f3fad8523b686dd8f879f51ad5cac4016fceb76b2b653e6455f2b7b8

Observation b9807f5c-88de-4f85-9435-c86598a4344e · outbound

This paper cites Visual instruction tuning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Visual instruction tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.640996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.226922Z digest=sha256:bda4fcd5abda95d9c96871a1ffb73b11488c649e517464956876a58fe5ba4f7a

Observation ba66a892-f2be-4524-a828-e81cf705a085 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Learning transferable visual models from natural language supervi- sion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.230819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.230819Z digest=sha256:777843fd046bc764a30fef839e2b041619a7e46aeada31988e6ab2cb13fc12bf

Observation 6fc286c2-30da-43c5-af8c-a801e47ba9f7 · outbound

This paper cites Vision language models are blind.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Vision language models are blind

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.619327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.234940Z digest=sha256:8095e83d451c429d95c20e8a0a0be0494611dccca6294aafc3abde9008c76ac9

Observation a3d89823-2811-4467-8393-ecaefd6298ee · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.605236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.238981Z digest=sha256:f575cbee6534a0fb6a172a597a63d0ed2b8f9845fa25d1fc2b8c2a8a40421136

Observation 3999e065-746f-4cf2-9d14-73ee2e0163bb · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.591033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.242786Z digest=sha256:ea273274627082ffe2ccf9a544469f03203f7f1c4102e2a19041cda3780cbdca

Observation 013f002a-5d4e-4f4c-ae2d-b02673928a89 · outbound

This paper cites Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.578106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.246827Z digest=sha256:98ccb3f1d0c3096ba1a1e2da28e47b522b9faffc9b12b8cf03a44aed48802c0a

Observation d6876a80-5e66-4a3c-89eb-4126c3eeee25 · outbound

This paper cites An empirical analysis on spatial reason- ing capabilities of large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models An empirical analysis on spatial reason- ing capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.565288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.250843Z digest=sha256:1b6e20e15ab818c236112b36c73f5909be84c361986e91e288ae3a37227342cb

Observation 9efa1615-e701-4c5e-8364-7242022bc0b6 · outbound

This paper cites Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.254574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.254574Z digest=sha256:4acb7db7c0d676b4ca8f0a31a86344d0f3e7d8c76b040fc477049d4de7834dea

Observation 18a6c9d6-e108-462e-a853-f4af6e022a58 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.552250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.258487Z digest=sha256:8f33656c796e44c353a28d398f6595606cf7e16bf222a7143877e9bce86952e2

Observation dbca669c-abaa-4d6c-8a43-a6df314ccbcb · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.262424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.262424Z digest=sha256:f7da9e910897483818c0308323136744d55453f5efbd7172026f06ac04915045

Observation b5c0014b-a35e-438f-8a7d-38f912c2dc6b · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.530294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.266213Z digest=sha256:89a4247c7ef4fba591e3e0e36e2930d251d563556cf69e4083de58c944cab882

Observation 2ab09e69-0a89-4155-980e-9eb6e2ac21cc · outbound

This paper cites Cogvlm: Visual expert for pretrained lan- guage models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cogvlm: Visual expert for pretrained lan- guage models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.516505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.270221Z digest=sha256:a225261eb6f8f908aafcb6bb6d8456e0b1e6c605d371a1eb36bc725053ed6e81

Observation d4faa040-7b49-45a5-bd45-318db2895f37 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.274106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.274106Z digest=sha256:1cd427651ea39979a70c5ea5b7250ecb523959142d2131ce3e4a5a9040eda1b7

Observation d865d26b-9e31-4d0f-be5e-ab97492a1744 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Llava-grounding: Grounded visual chat with large multimodal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.503755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.278600Z digest=sha256:1c9df4bb151d809e5abbfc14bcebdd3eb42b46e50ec48eb897222ad6043b1a04

Observation 1401d02a-b2a4-4778-9fd6-99ef7e1098a0 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.282590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.282590Z digest=sha256:746af3292c886a125120aaa37a402adb2dc5aa7ab471e288b88b9a6959ead599

Observation bd48bbfe-147f-4335-808d-61f7fae9aff1 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.476853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.290784Z digest=sha256:aee210649d643508d4e4c0eb8cff8de070860c8a0ff96e449b80e9f5a2c13c6d

Observation f920a5cb-03e8-4ee9-a0ea-a7d27d415fdb · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.490235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T23:06:49.286769Z digest=sha256:a2fee4193d07b287db3caf48fc2ec778f87a9987599601bae3b2d27740321c09

Pith citing papers

No inbound Pith citation observations are available.