Pith. sign in

Paper Citation Record · LEDGER

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 7 inbound Pith citation observations for arXiv:2505.10917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10917 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:59.827793Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:14.574922Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T18:45:00.174957Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd3bee83-896d-4216-8c3b-ddcd40ce3d0f · outbound

This paper cites Qwen2.5-VL Technical Report.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.750212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.750212Z digest=sha256:d87dd564a8f59983380295d50d1984e79f33fb414c14e855c7e8764e9b2b8049

Observation 342fac08-527d-41f1-9fe7-e2b424d58ddb · outbound

This paper cites Locality Alignment Improves Vision-Language Models.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Locality Alignment Improves Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.762345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.762345Z digest=sha256:35349d20893351fa08c29ae280a638dcce935ee1d3b0efc83ee0c930993c70d1

Observation 329ea386-3d6a-4592-aedb-2301564890e6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.765372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.765372Z digest=sha256:8fe7ca9517c8084aab20fa660d9b8b0efac2a3ffc5793679ad5f5a9f3b2829d6

Observation 7c7994bc-54a1-4c17-b3c2-0b80f0dc217c · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.778856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.778856Z digest=sha256:6b2e550ebeb58f2f1b616de4287d76a9ae1f10e978f6569bf387d4648bbec023

Observation bde1c36d-28a0-4fac-b6a3-6f62240ba564 · outbound

This paper cites Microsoft coco: Common objects in context.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Microsoft coco: Common objects in context

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.076558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.782811Z digest=sha256:3da2727f82c5b8d40b51cc91caacd8da681487bb665d6142e796ebaa249a9588

Observation 83052d8e-d279-41be-bdba-76e4c43e60a4 · outbound

This paper cites Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.787597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.787597Z digest=sha256:aee7eb19937f3ac9a8e822e7f358988421149cdd4d46443624b5a8c7e8bc3ad1

Observation 7c04ad6c-ac6e-40ff-9420-f6372fff69c4 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.799314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.799314Z digest=sha256:f29163f1337787606e1857a75755d9f2b35dfcb537d184cf901e04009e452d85

Observation d49ee1b3-6510-447e-a474-bce1d897dde8 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Reconstructive Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.802425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.802425Z digest=sha256:e3151d0c4fc2f3d41ca6f0aae2559a3af07d031817838d53bc4bb3aae5d9b6e9

Observation 05891336-2efd-4c02-8940-a9f9b5a05cb3 · outbound

This paper cites Accessed: 2025-01-.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Accessed: 2025-01-

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.047059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.805316Z digest=sha256:3e8dd4a757d2bbc30147f75ab735ea19a600258e658a6e8589006e81c2f1dea2

Observation 8fb35b7c-0c4d-425d-b401-c2eca526e82b · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.813004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.813004Z digest=sha256:9df4f46ac363d3bccb8d3dfac7f6e3a3f6664a0549f1b49101851b20921b32fc

Observation 835875ee-adc9-4e2b-90a5-24ec269979a2 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.815927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.815927Z digest=sha256:f429b1306b053b035911869574105653f0ee3e7e93044834746de27f3d8e1cfd

Observation 601caa60-7141-4474-91d3-18c64dc2e16e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.819535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.819535Z digest=sha256:a56f8a15213ec6b766818d6dcb4e79cf2865097371583e06b4b39c4845f9f962

Observation af89c01d-7931-4292-aa51-c115dad10ff8 · outbound

This paper cites an unresolved cited work.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:08:00.028481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.823432Z digest=sha256:d1cdb80a32336c0e0434243c86365e0a355d1079a8bd894ef6dd822ce87af86b

Observation 67fd94ee-9da5-4dbd-b54e-bf48ab174532 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.809431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.809431Z digest=sha256:b9f3955d72cf051730b2612a9f0337eafb32a6ac740ec4eb966c2e7bb2297515

Observation 15a76928-e198-4444-8fac-33a077137e0d · outbound

This paper cites During fine-tuning, the batch size is halved to 16 to accommodate increased model complexity and dataset scale.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization During fine-tuning, the batch size is halved to 16 to accommodate increased model complexity and dataset scale

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.011517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.827793Z digest=sha256:939ef9cdda0c0541ff96043584152d60a96f4e20460247f72d34d91c5d6885d9

Observation ab2a4859-a5e2-46f3-b4e6-b34b5d06e44d · outbound

This paper cites A diagram is worth a dozen images.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization A diagram is worth a dozen images

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.089592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.775310Z digest=sha256:e7ff6447bcc7b058ae6583292ec4a49be60e62d55da8af8d780632ee2793a1dd

Observation 0cf762b2-b22b-454f-ad3b-2c2397003bc1 · outbound

This paper cites Kimi-VL Technical Report.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Kimi-VL Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.795956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.795956Z digest=sha256:6f74f9b6881d85d42c1c1678d4bbff8abac035074a3464d5032cec28590d85e2

Observation 1692588c-b51e-4fdc-ae1b-c2cdc1f82b46 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.768232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.768232Z digest=sha256:f83286ca91e52df953b032a536554db66a197f48ea553cb1da63c3836f522800

Observation 4b1c5bce-cc5b-4b38-ae96-e1a43230aa3c · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Ocr-vqa: Visual question answering by reading text in images

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.063149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.792420Z digest=sha256:388273d995fb4dbfc961c333d77befdabddff7509bec288eafc263950b0234d3

Observation 95c087a8-7383-4bad-b687-73f6fde120a9 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Referitgame: Referring to objects in photographs of natural scenes

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:08:00.101799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:07:59.772019Z digest=sha256:42d91e2af6e6979d8983735d8910d39d33370dc266ad89af505b7e4fbb946c9f

Observation 3458492c-aa16-4c70-8214-bb981a6c7456 · outbound

This paper cites Gramian Multimodal Representation Learning and Alignment.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Gramian Multimodal Representation Learning and Alignment

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.757737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.757737Z digest=sha256:5ecb52011dee5e715878ab3ccdf8c59cadb0f86cba902194fd54d1a7b3549df3

Observation 9885729b-f0c2-42a1-91ad-f1aa82b1d2b3 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.753923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.753923Z digest=sha256:c6d9c99716e67c5d5f7f6e6f2e035c1459992fe9287233238b4ae5e9dcf1f8f0

Pith citing papers

Observation 1a1910c1-7522-4081-8d32-e8820e2d1460 · inbound

From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents cites this paper.

From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:14.574922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:47:14.574922Z digest=sha256:06a17b1b177fe3a4249c70a64a85481af36b6c075b9a36a2a1fd5a6e8d35fa0f

Observation 3e4f6cf0-67ff-405f-88ce-3095010eaaa4 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:00.805264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:00.805264Z digest=sha256:d75e31987e33ceadac672d90ebabf8746f97d0f26cf9cf876c0056b9aa3f5392

Observation 67c046c7-93fe-4e25-8faf-bfd9bf4d1412 · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:52.569023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:52.569023Z digest=sha256:efd7c7c88b5a2fa0f49dbbae244ab083fac4a8b338c607d07486aa26bfb23c49

Observation af063bff-0d23-43a7-bdbf-c9a6ab11253d · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.446584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T11:38:01.947806Z digest=sha256:d1ab8a953f6e656039568bad42bcdd9ff08f7bb11ef829e3825a320117500cc5

Observation c0ab3302-a89a-4ce4-8e50-5d72fee52e11 · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.177236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:42:46.422756Z digest=sha256:6da8bd29ec4f6b1dc30c8c1c0f85617a7b475283ed02dbef3727c068e3010a7f

Observation daa69f7a-4758-43b6-8c68-3db61f49861e · inbound

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition cites this paper.

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:26.088290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:40:30.038337Z digest=sha256:cb187163b2efef0d775a55c1b34785c81b12d4ce2497f107395a6577e9183133

Observation badd34c8-4d75-4c39-b6c2-9bc6af3ccad1 · inbound

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein cites this paper.

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.712771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T07:28:52.030720Z digest=sha256:5c3d439646c640e02e31d27641b1afd272e774d1fb792d2e85fffe12baff14b8