Pith. sign in

Paper Citation Record · LEDGER

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.05626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05626 v3

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:49.290784Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dace842-e939-43af-b761-3563d2ff32e4 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.692845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.185329Z digest=sha256:ea2a9cfbfb07f1a3836966c3404fd3850a8a863b63191fe0f981cf10cbcce4be

Observation 0d8090c1-07b4-4f97-ad1d-4b078e715398 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.679539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.190221Z digest=sha256:f629c891d93be21a34a048137da78460c87acc652718e0b9481b1e0ccd616594

Observation c9a3fce8-87bf-422a-ba77-fa9d7cacfb40 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.194642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.194642Z digest=sha256:4717a5f05cf4cd4e21db1a9d1ab374b9cdc3d3eb77960d23d1e1a7dc972bf95f

Observation 83cb824a-d8a7-444d-adf8-035a30653275 · outbound

This paper cites From colouring-in to pointillism: revisiting semantic segmentation supervision.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models From colouring-in to pointillism: revisiting semantic segmentation supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.199068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.199068Z digest=sha256:312f712b16916dc5baf3d280bdc36e23b2369fbd01445adab48c50f5a7e033ed

Observation c231a48e-3a5c-4215-a6d1-6c23b9219b01 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.204156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.204156Z digest=sha256:fbd0e51948b890de2675fc65c037faf499e11d59840f3f0a09e70b155a42c21b

Observation c9d423cd-3234-4f39-bc5c-59a7505f519c · outbound

This paper cites The Llama 3 Herd of Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.208740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.208740Z digest=sha256:196bc3696bfef51b85f06e2914570393ca4d0735d34dfde016b1c6faf7425ab0

Observation 2ce9a543-525d-4b16-b5f1-56ab96fb1fc6 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.666781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.214294Z digest=sha256:36c5bbb2a1b1fc55c9aaecba5f8ebd61b4e4d2a083bc34383f7a6eed9da24e3b

Observation e565e41f-649b-4594-8d4e-23f0bb47a7f3 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.654002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.218511Z digest=sha256:eae67826225ec76e29eac10bd7a041d2db7cb30bb1f4087d1ff19102e55ed0e1

Observation 058f94bb-38e4-4873-8a6f-cdc5645d5cc8 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.222650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.222650Z digest=sha256:6f53265a52202b479f2dc5b469fea4560c3b14c56f2e1ab3d618427ff1d998ae

Observation b9807f5c-88de-4f85-9435-c86598a4344e · outbound

This paper cites Visual instruction tuning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Visual instruction tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.640996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.226922Z digest=sha256:806fc037e131c6a6cf6e3d26add3e6708966f6311da097edd13b4422da0fe84f

Observation ba66a892-f2be-4524-a828-e81cf705a085 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Learning transferable visual models from natural language supervi- sion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.230819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.230819Z digest=sha256:c64c6a921d1d0ac4644f3a52a9a65ace775767dc6bb532416bf6435c47d87df9

Observation 6fc286c2-30da-43c5-af8c-a801e47ba9f7 · outbound

This paper cites Vision language models are blind.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Vision language models are blind

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.619327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.234940Z digest=sha256:f4536bba510e1960a4985a7e21f77f4f4d2910dd1836814d51f17a51c7415c1a

Observation a3d89823-2811-4467-8393-ecaefd6298ee · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.605236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.238981Z digest=sha256:7830084f80c802e7f5692ac535fbda306658cc730373e0fb2987a84e3f71c4fb

Observation 3999e065-746f-4cf2-9d14-73ee2e0163bb · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.591033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.242786Z digest=sha256:3d899932a2d92f18309526c6fa051efcad4fdaa9a7f252a203ee6181d5b7f979

Observation 013f002a-5d4e-4f4c-ae2d-b02673928a89 · outbound

This paper cites Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.578106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.246827Z digest=sha256:eb5758618481805bcdca484fe591d1f80887d100b7c00ac844ec85df6ccaa52e

Observation d6876a80-5e66-4a3c-89eb-4126c3eeee25 · outbound

This paper cites An empirical analysis on spatial reason- ing capabilities of large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models An empirical analysis on spatial reason- ing capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.565288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.250843Z digest=sha256:035422460580c0862a3ed58818e72d3a5d3d45d6b1b09af39d5e57498ea68ce2

Observation 9efa1615-e701-4c5e-8364-7242022bc0b6 · outbound

This paper cites Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.254574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.254574Z digest=sha256:79d0b1059467ef3e6f63c552e47fbba66d368493192f03a08b16ddb2e529f2a3

Observation 18a6c9d6-e108-462e-a853-f4af6e022a58 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.552250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.258487Z digest=sha256:02d3e254fd553285fc2109d17b8fd8544e75ec2a7adb86fe41c84e28ee1c87f7

Observation dbca669c-abaa-4d6c-8a43-a6df314ccbcb · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.262424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.262424Z digest=sha256:388d6cd4e60b1d664c6ee04c774aef226b104c67da3240a2017054e4a2275fbc

Observation b5c0014b-a35e-438f-8a7d-38f912c2dc6b · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.530294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.266213Z digest=sha256:2c09b3984df551b23beac01e8acdc6ac7efa73f316cafa193ae8d5d533ab189d

Observation 2ab09e69-0a89-4155-980e-9eb6e2ac21cc · outbound

This paper cites Cogvlm: Visual expert for pretrained lan- guage models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cogvlm: Visual expert for pretrained lan- guage models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.516505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.270221Z digest=sha256:1126f030c14c46f08826ce7ff71f024c35192ab1f2fbddc9300a40a230102d3a

Observation d4faa040-7b49-45a5-bd45-318db2895f37 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.274106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.274106Z digest=sha256:7ed4670fe169986b355faddea049b64bd6a1ac017f7c1ac3470c30b21f7795f3

Observation d865d26b-9e31-4d0f-be5e-ab97492a1744 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Llava-grounding: Grounded visual chat with large multimodal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.503755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.278600Z digest=sha256:71496ca6e2f574a9fdb764b5f04325ae5341aa9274f4b7f8dc0dc5f85e3cb965

Observation 1401d02a-b2a4-4778-9fd6-99ef7e1098a0 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.282590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.282590Z digest=sha256:862ff991fea8e25a3ca136c9d1abf773e610bf2ded949ab8c983125951c21d83

Observation bd48bbfe-147f-4335-808d-61f7fae9aff1 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.476853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.290784Z digest=sha256:ae7177b457456e3046891ca45250aa681276b072757fb19aaa01780e45c99920

Observation f920a5cb-03e8-4ee9-a0ea-a7d27d415fdb · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.490235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.286769Z digest=sha256:0d24b5722aa4fcc810812ecc9c793c61a2f33c3e3dcf1f6f224bdba06a91bda4

Pith citing papers

No inbound Pith citation observations are available.