Pith. sign in

Paper Citation Record · LEDGER

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

As of 21 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 4 inbound Pith citation observations for arXiv:2505.17316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17316 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:09.476244Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:57:29.460876Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fadc0ca7-aeda-4650-b84f-c888ac9a3105 · outbound

This paper cites Visual instruction tuning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual instruction tuning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.433783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:04.097733Z digest=sha256:a2834b37e92140c4fe2e0bc4b413589a1b0e0b188cea0e063170d0ac34e49bda

Observation 29de9b4f-7a6c-45b0-ba66-65863a79d134 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Improved baselines with visual instruction tuning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.277054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:04.208405Z digest=sha256:9ab05e53cd10cfdddd56faaf07eeee3795ddc45dc0e1b36b21346e65963b4474

Observation 208045f1-7904-41c4-a316-77bcd68edc2c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.379623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.379623Z digest=sha256:a7c3224216b2ad770d37cd054f66ad8f541bc4f602898af34a7e88f559fc1153

Observation 70584995-12c7-48a0-89ce-68f725df6f8d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.539750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.539750Z digest=sha256:3a9602d36bbc081e7e2fb0e5ac8457cca036f0e203e411a65ae39bbdc705984a

Observation 5db061ba-bb92-4291-80e2-ec1c3ab061f2 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.677406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.677406Z digest=sha256:44cf990276de561137f818c6b3b9b3610b3174fd0f09eeca4ebdb6adc986618a

Observation 98609cdd-071d-48fc-a888-967810b5d0dd · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Flamingo: a visual language model for few-shot learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:04.770528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:04.770528Z digest=sha256:0ff63bbe85669f5b9bd7210936a770a8844748da08e49cb18e341c1054d6f30e

Observation bc4869a1-3879-4476-9c1f-272048cb17a6 · outbound

This paper cites Language is not all you need: Aligning perception with language models,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Language is not all you need: Aligning perception with language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:14.117267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:04.917798Z digest=sha256:bbbc4b38636b630f4c9a5c45eb477d215ae200b895cc9ee54e9044157b3fa49c

Observation 47117c81-1b56-4a5b-aae9-9bfb30599547 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.094450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.094450Z digest=sha256:abf5c01c22cbf3bf517bfa6a7103900fe0b651c9dcf8d26f82a30ee2d83ddea6

Observation 09fdddc4-ddbc-4e47-8c00-b13a5032e0aa · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kimi k1.5: Scaling reinforcement learning with llms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.964943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:05.186005Z digest=sha256:4744e5edf743975941b8171fd0aeba77e3d6fa4932114213df3fd79405eebd5e

Observation 61a44add-9445-4537-aaa6-470852bb6cc7 · outbound

This paper cites Multimodal transformer with multi-view visual representation for image captioning,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Multimodal transformer with multi-view visual representation for image captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.839155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:05.270372Z digest=sha256:d0aa9d66019783e3fa80acec773c318fc3f5dabd9a0c2e42919837d1f643765f

Observation 12569a1d-b954-4dba-951a-aa8085cbd105 · outbound

This paper cites Vqa: Visual question answering,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vqa: Visual question answering,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.384694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.384694Z digest=sha256:1a602660575068809399313d68d51723e253bf6863149ecf0494a25f549131ab

Observation 8bb64434-aec4-491f-8c0d-89bf212f3a5b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.476390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.476390Z digest=sha256:ac1b0bda095cbecec9f5d8f25ea76394da877b8dc46c721e3f5450f9dd105ae0

Observation cc12dd15-7ac6-46de-a797-54b4ce1bddf1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.597599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.597599Z digest=sha256:c79b06906c74f6be169c4231a909ee8d5f0298640b316cdb08eb93aaa3a75071

Observation 53446d3c-8f1f-4f37-b759-51d558bf7049 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.690123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.690123Z digest=sha256:09796741a9ebb4d2cb8ed2202b1d74c86405a86226b9eb7cebab023cba827153

Observation 29156912-f77a-44c8-ac02-3ef445607f04 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Glamm: Pixel grounding large multimodal model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.659627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:05.769581Z digest=sha256:2e0edc537c0775530717facc86b7b20dfea15bb1af1d235d4f66ca2e9acb6cf4

Observation 099610fd-b9dd-44de-855d-e0271fb59288 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Hallucination of Multimodal Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:05.858115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:05.858115Z digest=sha256:f8dd4a4e5908bef3f05bd47bf38dc5736d72f974d53852d82433a6b8a6278086

Observation d751b048-6e58-4c5e-9b9d-f1acb8e9589a · outbound

This paper cites Honeybee: Locality-enhanced projector for multi- modal llm,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multi- modal llm,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.495833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.013664Z digest=sha256:219db7bae7ed790b4545274953aada15fc4fc921d97e74ffdb560546771d4b58

Observation 3bcaf127-cac2-4f33-aa77-d8a119ec9186 · outbound

This paper cites The Platonic Representation Hypothesis.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Platonic Representation Hypothesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.128909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.128909Z digest=sha256:2383f596338e16a5d5ecc0451267c3be97c4fdbc30f1b4b6b8f16ec7dd8c2be8

Observation 29f11f75-9d76-409a-8d86-51f6192b84f6 · outbound

This paper cites von Neumann,Mathematische Grundlagen der Quantenmechanik.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models von Neumann,Mathematische Grundlagen der Quantenmechanik

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.345684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.209608Z digest=sha256:455235112480bc783a8914eeba4ed7d3ff57bacded9cc0186ddd665d373f0a40

Observation 1cd1b476-c3c0-49af-a2cf-202ca427a365 · outbound

This paper cites Linear algebraic structure of word senses, with applications to polysemy,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Linear algebraic structure of word senses, with applications to polysemy,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:13.022576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.426208Z digest=sha256:118e802dc45cb9b50ea23570bad67137648379871a9ff5ce3e365068017a50ea

Observation ced2b0b2-19d7-456f-875c-eeb48feff724 · outbound

This paper cites Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.502588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.502588Z digest=sha256:5ed09977ac86fbfc60412f89aada7c54ab5a07b5885f60ee923445bb06827b2e

Observation ba7c44c2-883e-48e9-96ee-9444c5df2e11 · outbound

This paper cites Signal recovery from random measurements via orthogonal matching pursuit,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Signal recovery from random measurements via orthogonal matching pursuit,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.858259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.619408Z digest=sha256:79658499585bff0b45237f91ed3467288b092c93ec721d93b52a37867f405871

Observation bc3cf9d0-2f91-4990-82cc-56bd5650f420 · outbound

This paper cites Recognize anything: A strong image tagging model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Recognize anything: A strong image tagging model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.717265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.690111Z digest=sha256:8af7312ac4809f30f0f18dfa49fa2723e33358ecc31333b051b2842e9339c085

Observation 8fedc27c-4a1b-4068-8304-468498f53fef · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.505247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.753385Z digest=sha256:519dac0b996fe1b439e647d4090d176a179c339eb784bd15886cf8b9f0722cd3

Observation 69cb3100-d9ea-4117-95a9-8016afa821b8 · outbound

This paper cites Segment anything,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Segment anything,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.817193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.817193Z digest=sha256:cc166454f44ee7c54fd5eba93689d009af997386da1cc3dea886c5e2adb71c93

Observation f5cf6810-da43-4cd8-81d0-acd914e552a2 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.857116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.857116Z digest=sha256:8fdb58debde9826b54fea80943afbf1a1f8e673b5599d6259b88de1ca5e63ffd

Observation 26a3b850-f28d-414e-a021-fcf65bd986ea · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.924738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.924738Z digest=sha256:a0980c8e4eef2ab6edf2a7a6216fac890ec85b8a1b3f0d44b0c631670715dd6f

Observation a61da63f-e94a-4644-8fe6-02badd35fdaa · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:06.984546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:06.984546Z digest=sha256:30500df0b9010c3f6b804fbf275849281d1a9fc8d155842bb1ab06df32b24d1b

Observation a07487ff-fff9-47ec-8cc0-69e5657edb1b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.035989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.035989Z digest=sha256:d3de8525473062c285165c113740f6b7dfe73e08c0f673cf493a8004959dc5a0

Observation 294bd963-4843-4895-a63b-ea8567de710b · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.161010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.161010Z digest=sha256:3b8051aa759531dfe8f609e44734f5bfc6eecba3a91415636917615875583012

Observation 410e58b5-1c87-450d-9d58-7086b04f6ba4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.253612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.253612Z digest=sha256:08c35ca03ec8e34e73b2e7d000e9aeaba6bc8eef593f8457729bde2d2316c921

Observation f78e1c9d-00dd-436e-bc06-0b9b921d9ed9 · outbound

This paper cites Law of vision representation in mllms,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Law of vision representation in mllms,

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:52:09.875483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:07.363562Z digest=sha256:c52e83ad1fd73c82da100939b6afde7d5f8243c9f6f662190672235ff7a86b8c

Observation 58d0ae74-b066-4fa9-8c35-3430164cc0ef · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Learning transferable visual models from natural language supervi- sion,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.324474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:07.459692Z digest=sha256:ae4a520704b54c5ffb080ef30d8c48ef3057c78dc37296f25447fb536b9e01ed

Observation afd9b67a-ee7f-4db6-a5d8-cecc6819e3b2 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.553797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.553797Z digest=sha256:fc77e6d1ff349bd0a6edab11bcd9c27d4a0c88af96526b05facb5634ee1340ed

Observation d5c0e1f4-08fa-4fae-a4d5-f324eebe1557 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.702315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.702315Z digest=sha256:f1a1763e7e47689b7605e4aa01b1d89299b974139d43e70bf88769dd18bc334a

Observation 258e03ae-c032-4490-886e-02438a42f2c5 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.801312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.801312Z digest=sha256:08300e5959e10720f981806d020b8106dec700f74ff1e70669262ad7704b8976

Observation 15f7dcb8-de70-4ef8-a31d-a938af74c5c5 · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.914884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.914884Z digest=sha256:939b64739c3bb4b39caa8b58baf7103d2fde036640b2d4655d324b7a47a0577f

Observation ffb7bfab-4a76-425f-9254-89705a2460c8 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Honeybee: Locality-enhanced projector for multimodal llm,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:12.176356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:08.003998Z digest=sha256:ec3edd602611584524a4cf3d1a293b3efbb9b2e30afdd7a266c1a3c7d14b1335

Observation a4e8b62b-1cb5-45b0-aee7-596158711fb4 · outbound

This paper cites Microsoft coco: Common objects in context,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Microsoft coco: Common objects in context,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.137009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.137009Z digest=sha256:3f382f320bd78601f2dc33d7c40c74bcc42357d52084e7bc271c781f9341679c

Observation 66925e60-7f15-4ea4-b05c-42d93478d3d6 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Referitgame: Referring to objects in photographs of natural scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.991080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:08.278205Z digest=sha256:8b4a9e1d779538ad3a058c0f050f0b4aea99b34b658e0eb2f15b09cf2b3f5ba6

Observation 766fa7c7-55c0-4cef-81ca-9e6e07e5deed · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Generation and comprehension of unambiguous object descriptions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.792342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:08.425913Z digest=sha256:944fc1b2020f919d97708dad6655b2ffddaddcd6c7b6f990f7b8d0e98d8228de

Observation 519e912a-9cd7-48be-99a7-28ddc9d0707c · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.552384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.552384Z digest=sha256:e7affe13d6f7e6e28069fca5867b9a6928839eeb69141ed7772b30529604e6bc

Observation 94b6ad84-73b1-435c-aefb-38348b595dfe · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.567486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:08.630284Z digest=sha256:2799af3d220dccb46ec052ee3167bcb2cabf564f5d9523c4c62bf41998b1ee45

Observation aebb9424-4cd8-4cc6-9db1-d85de7550613 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.711601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.711601Z digest=sha256:d353afe138908bda261b193cb1ce1cdb31c9cc804480474b602fb3c3941105bc

Observation 3ee107fa-bfd8-4446-bb2a-7f31b967b7e0 · outbound

This paper cites Instruction Tuning with GPT-4.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Instruction Tuning with GPT-4

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.800402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.800402Z digest=sha256:a93aea6f2efdc4f753993e68a1a9bb269f345bf174a58a3dc42bfd1494b792e2

Observation da6a0ac9-3e8b-4782-b358-5e15b6c0d4cf · outbound

This paper cites The Llama 3 Herd of Models.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:08.905998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:08.905998Z digest=sha256:2760b233ddaf7348c0e177f93565c48bc8e7fccd2a92e4a6c1d8ddf946a98890

Observation 55246c22-573a-4a8c-a38f-94e1528d6620 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.305627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:09.039392Z digest=sha256:0c8ea1135c852a7183b405df3268b848afb85c5ac57d9702a33f90ddc4044e79

Observation 717aab0c-5678-414b-b5a8-09c0685b4fdd · outbound

This paper cites Towards vqa models that can read,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards vqa models that can read,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:11.043224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:09.193236Z digest=sha256:266f12dff055c7ab06ef3a9330c9819a515de14aadc35b34a1eee9db5785196d

Observation b5e92fd3-23ec-473a-ba67-f160ea063043 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.787542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:09.291633Z digest=sha256:a61055e15afc2d8da52b1875a0363b985668478e0bba88f5b6079261eea682ee

Observation 9143f5af-b801-45d4-b5fa-ced2989a23a3 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Ocr-vqa: Visual question answering by reading text in images,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.559675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:09.394927Z digest=sha256:5851a77581eb7eb19748792b7c4acb4b2d76b67f3df40ae4500d945f8379507f

Observation 894462a2-6627-4fe4-a7f7-ad0abbf5ec84 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:52:10.298358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:09.476244Z digest=sha256:07acabd495e1c315e01c9d7dcf8530ab6ad94f1e2632cf1d21f11151d5062ce4

Observation 81ffb6bc-c061-4c86-96d1-c38e50e59701 · outbound

This paper cites an unresolved cited work.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Unresolved cited work

Reference 1932

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:13.179783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:52:06.306580Z digest=sha256:ea490e807328b00269860cc84b311b6b626fa5927f992855c903feece6e01bc5

Pith citing papers

Observation 016138de-0e18-46f9-a76b-1f2029ae70c9 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.689929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:b9380b82ed24236be1baede0dfbd9eda8b8c263de119f503837bcb82b80a8a29

Observation 9c925f13-3f9d-4834-b9d5-bd95ebda16bf · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:20.436198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:c3ed76b5f3e39a7cdf477b99ff72b966b327b9f15f178c0fde1a5c33a7c5af6e

Observation 5774306e-2e1f-44d5-b230-98bb897ca4f1 · inbound

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning cites this paper.

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T04:35:49.701061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:35:49.701061Z digest=sha256:e55f7e83ff254edc7e97755b760455919afcd954139803a82f207f357f438bff

Observation d6ac425d-f669-4b50-ae94-43c4517bc5f7 · inbound

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment cites this paper.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.460876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.460876Z digest=sha256:b3c85945e9fa1cdba94ac39ff1457949249d8e9ac19f833af5c94ae70d21b33c