Pith. sign in

Paper Citation Record · LEDGER

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2310.00653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00653 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:54.590335Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 217ca226-394d-49d1-82e7-fb462604f334 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.309406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:15f9b1b85e6b5006677cf01889b61c71ec3fc3a02aacbe09aa4eea68e051b55f

Observation 6aa60318-4518-4baf-b743-feb9d0bdf296 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.307160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:7a4f96ed5442541e5a224e39aa5205a8dda0f791ccdcb831fd8dc242a2816bea

Observation b5641bb9-40a0-4262-8502-9a954cfc4b04 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.087932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:115bd26f63d48ff7240ed5860bb4911806de0d4a2edfb6ed2569b38015a06612

Observation 0b33969a-66ec-4006-8894-d29737ff13b1 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.993709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4e0271948c26ee8f096fc3baea3eb531dbdae3121bc60895766963e4c0433b35

Observation c745772d-88be-4c47-9f00-6d7516d1e0ac · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.819520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.819520Z digest=sha256:a46eb77eed9007bf61877c6cf06e9454ae1efd5094adb95ef91fb02b0f19434e

Observation 1edcaefa-fd54-4104-8ec9-14bc52363ac6 · inbound

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs cites this paper.

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T11:53:55.481346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:53:55.481346Z digest=sha256:a65ea3d72b7b06a0b5b99140a5875b2eccc00ce22bad8a553f8b3d8030556bcb

Observation 82484be8-65d5-475b-8c38-226399ffc1f5 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.490358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.490358Z digest=sha256:354eac45d189dca07535564e6a1c34d5eb07aa28bee97a38bf50aafaf340b2f6

Observation 7d7ebf00-79a0-40bb-aa53-1d77bebc41fd · inbound

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization cites this paper.

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:54.590335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:54.590335Z digest=sha256:6c052c2b551318668181c8451564b6f0bfd99f4e51760f9a08f85194f7a60b7b

Observation 02f89778-fd8d-41d2-9518-107419363a5e · inbound

A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision cites this paper.

A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:20.914042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:06:20.914042Z digest=sha256:25ca6afaef2185032d6f244e05aaa5c17ba708d16df81a5baec9e4182231202b

Observation d24988dc-67bf-42b4-b1df-8f20cb1da434 · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.580874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.580874Z digest=sha256:4e3d8fa00addcfe5ff7b07676fa475a92ed5fc2eeafb7cf11aa4e9b78eb5e5ee

Observation f30a3137-a4df-42e0-871f-31751ce3591c · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:27:39.259198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:c9be20ea6fdbd46994b2252da10ce4b54d9f09641ae4b36d96d1fa36d7364e0e

Observation 965a76ad-644f-4ffc-87e6-c643293715e6 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.417157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:22e3c5e52cb2d903a54447e92220853e212ebfdc9a344782d3de72464e122182