Pith. sign in

Paper Citation Record · LEDGER

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

As of 8 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2509.06994.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06994 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:11:20.896514Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25d29911-204a-4527-85e4-62c23e6e7ad1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.756776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.756776Z digest=sha256:e6ba291ec1b45814d7e2b0d9af0b225f929ee3521ff5fea9a3caefba451b7193

Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.762326Z digest=sha256:cdd383388178ff877ea6e20d5ecd465de99e453d8bea6f87ad58b9de5ba35b84

Observation 3dc30fb9-740f-40dd-a8f9-893cf054f866 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.767242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.767242Z digest=sha256:db31567cd068539b001f4f2da14ed358e69bf635339e029cf5078772b6935b1a

Observation a95b629f-7b0e-476b-863a-9eff8deca0af · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.772877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.772877Z digest=sha256:0ce75a49ae6d009e6d8548d5ce104dc7f12d7623190884d33c2a7bfc741b12b1

Observation 125b24fd-37ce-4e92-8b67-0a961774bd46 · outbound

This paper cites Visual Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Visual Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.778714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.778714Z digest=sha256:eba757c8ccd626dc3aec5af81dddc9e8dc1850dcd78a65f298cd20e09834ff32

Observation 25f20eff-fa39-4a1d-8d7d-3b9923386019 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.784030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.784030Z digest=sha256:7c5fa04978af0ec53363e620eee92e68a8bcdce37e6483edcdd256086601b827

Observation 57ce2f51-1090-450f-8057-65b228fe2d21 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.790509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.790509Z digest=sha256:6734c60ed654550331c117d0c54eb21cff3cb4302f751e438c2ddd44f2a438fa

Observation b49cb21d-d80a-4e59-ae9f-78a74c7026e6 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.795205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.795205Z digest=sha256:301f30db09dcc12d2494537800020976512fe4cc3163a5e4614e8de909bea285

Observation 28ca6039-9c4f-4e14-b15b-487c61ac902a · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:11:21.766317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T11:11:20.800168Z digest=sha256:94b0ed22900faf88e708da48f78117975afdd775f9a0510fb52e050cc7b5a230

Observation 9ae5c169-ec7d-478d-bca4-ac87619a245b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.804751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.804751Z digest=sha256:50aa9efb3330979d6623b498186aa73b654211df4db16680016c5b3f50a17ceb

Observation f7ab3902-57b5-4892-9090-bb475f8b8cea · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.810277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.810277Z digest=sha256:0f86ad86ed8a88460500c52fd5f42be6ae178b3376507ae12a83b7e50e06d82c

Observation 81de8b30-b5c4-49e4-a8f0-2efdf15f0475 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMBench: Is Your Multi-modal Model an All-around Player?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.816367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.816367Z digest=sha256:f29dae2dbb0e5a7fc4ce8f49a8253d9d8c05211f124a266e48213b85acf678ec

Observation 92a070ab-54e3-4f19-acb7-99389f5a1150 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating Object Hallucination in Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.821484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.821484Z digest=sha256:6f0072184a147b43dde1af30596e263c667a3e1526d2ad88b4625ab5d99fc395

Observation 79822f8d-424e-4bce-8edd-cbd1201db123 · outbound

This paper cites Towards VQA Models That Can Read.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Towards VQA Models That Can Read

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.827420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.827420Z digest=sha256:763aee464adc292ef552642a179b30a095ba237d809e2dc6a0f1286d1573fed9

Observation 2b5709b0-a04a-4bd6-87d9-71b46c381cc6 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality DocVQA: A Dataset for VQA on Document Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.832097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.832097Z digest=sha256:927ee85029868b292e091b68de7c34694cb483c84466da3fdd4cce965a84096e

Observation f7336672-e16d-4c2b-91ac-d8155104ff79 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.836873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.836873Z digest=sha256:b6de2e10382512caab289f65ffef1423f6832011cdeaea99716540c894a27d42

Observation 071e3566-2153-4815-ba45-743320f94247 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality A Diagram Is Worth A Dozen Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.842285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.842285Z digest=sha256:1b44261ae9c4bbe6ab8fc9c4f7173c966e0c6a3feecc3f18437686d74db30acd

Observation 61f6fb42-0646-4ab6-9d9d-70cffb7bfef7 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.847152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.847152Z digest=sha256:6585a8a9d627ed1a6e70777a47e104b242986d72d8841f57a128bdf1174a2cb9

Observation 593e1ea5-8e24-414b-bfa4-a69f352623d6 · outbound

This paper cites Redmon, S.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Redmon, S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.851633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.851633Z digest=sha256:54765e53da4ae7971cfdec9046d9eef21085459fc0b6361da801a01bbb32643b

Observation 5e0164a0-d264-4499-a12d-187ea2623436 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.862036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.862036Z digest=sha256:540d828179656b57d219166819807433fc3ebe9b9e2ddca135ac19fe62778b25

Observation fe9e5a2d-8e5d-4476-a818-28d0902e21fa · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.868016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.868016Z digest=sha256:68358bc94c6e2c500bcb8c44f49755f9de18ee0954dd896cba1fffaffaf44562

Observation 4035c6db-6ec6-4dd3-9421-f7f623789f78 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 22

Resolution
verified exact
raw_fallback, observed 2026-08-05T11:11:21.162277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T11:11:20.873379Z digest=sha256:7d6066c1cbb44760e11e1e230861e87b7190779261f7e1fd9726ebe4bb178057

Observation adec65b2-bbe4-4c14-90bf-310ef96b25da · outbound

This paper cites Qwen2.5-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Qwen2.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.878592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.878592Z digest=sha256:a9e1c780c40d680a7562e28fd60222ddcc3dd13c9562a4c85e27b7bd6ae0522b

Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · outbound

This paper cites MiMo-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.883362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.883362Z digest=sha256:d8e49489f49bb555effb175d44cda383e874d954b36aded0ce54d2bebcf94136

Observation e788a6a7-88ad-4c27-add2-a97b5fd6fea8 · outbound

This paper cites KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.890509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.890509Z digest=sha256:e552445f3199d9aa23b99de7480fc92ce76dd70ef8c3807897dcdfc1d2bc2bb5

Observation ea359d5e-3ff9-41ae-9495-eb45cfce5b40 · outbound

This paper cites Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.896514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.896514Z digest=sha256:debbd143f5ec5765168e87fdfc5640e45a49d8a7f802529aae9cc2f72cedf8af

Pith citing papers

No inbound Pith citation observations are available.