Pith. sign in

Paper Citation Record · LEDGER

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2509.06994.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06994 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:11:20.896514Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25d29911-204a-4527-85e4-62c23e6e7ad1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.756776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.756776Z digest=sha256:9be989b22c899b88d6eaea05f33aa4700015ccc5ba1b68befb9974c91fafac8f

Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.762326Z digest=sha256:e6275d8ebcf4af72f437207440fbdcb786bccb16c79183787a25b1b107d964e0

Observation 3dc30fb9-740f-40dd-a8f9-893cf054f866 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.767242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.767242Z digest=sha256:570ea306c9456d69d8510473ea8ca2e7c4dc6d441795381a5ea8ec72e463d3bf

Observation a95b629f-7b0e-476b-863a-9eff8deca0af · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.772877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.772877Z digest=sha256:9886532336a79f419ed9f3af70e856902283ab86c3d82f78645df7a1faf5e2e4

Observation 125b24fd-37ce-4e92-8b67-0a961774bd46 · outbound

This paper cites Visual Instruction Tuning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Visual Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.778714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.778714Z digest=sha256:737119491be92e74ec4a943fb58a5d456d55479b38ace390e3cd6e39436b3d05

Observation 25f20eff-fa39-4a1d-8d7d-3b9923386019 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.784030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.784030Z digest=sha256:04728dc435869f3d663646d45a8556adb2dec93fd34351f694e30e7d69dc17be

Observation 57ce2f51-1090-450f-8057-65b228fe2d21 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.790509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.790509Z digest=sha256:94530394a77354c2eb317918f311a93aceb56dd8617899c6dbdbec4099113ce2

Observation b49cb21d-d80a-4e59-ae9f-78a74c7026e6 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.795205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.795205Z digest=sha256:67a7a44b84ad54b3175679f4323ae663136accbe72483562fb4b336dffcd9bef

Observation 28ca6039-9c4f-4e14-b15b-487c61ac902a · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:11:21.766317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T11:11:20.800168Z digest=sha256:b802474d58e2f8434161ef6694bbcda9290f446411563b8d30011ed0b8857139

Observation 9ae5c169-ec7d-478d-bca4-ac87619a245b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.804751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.804751Z digest=sha256:75fd9664b67feff7d7c9a7c3ff8c3a9294a9b3723e149623018a1300394660d8

Observation f7ab3902-57b5-4892-9090-bb475f8b8cea · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.810277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.810277Z digest=sha256:e317ee5ec504cfea9fb590dd4f58beaebc9d4579ddd5d413c0fff8a59888aeb5

Observation 81de8b30-b5c4-49e4-a8f0-2efdf15f0475 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MMBench: Is Your Multi-modal Model an All-around Player?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.816367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.816367Z digest=sha256:4cf9d2372f799e494cecd4621428d2a989d1f3bdfc51f8156da7dc47dbafbcbb

Observation 92a070ab-54e3-4f19-acb7-99389f5a1150 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating Object Hallucination in Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.821484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.821484Z digest=sha256:c274c0d6afec90f705be851f052a5e9d1d25fe84dc3f6b870114dd441f131f0d

Observation 79822f8d-424e-4bce-8edd-cbd1201db123 · outbound

This paper cites Towards VQA Models That Can Read.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Towards VQA Models That Can Read

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.827420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.827420Z digest=sha256:d0a20eff22757f2d6a8f2fa4c88f4b4497bb9df9e9cbfb95adc71f98793b957e

Observation 2b5709b0-a04a-4bd6-87d9-71b46c381cc6 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality DocVQA: A Dataset for VQA on Document Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.832097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.832097Z digest=sha256:de9bb45fc82a44decd62cdf7631ab16587c1d3c87e07cb0977527b901bd66b51

Observation f7336672-e16d-4c2b-91ac-d8155104ff79 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.836873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.836873Z digest=sha256:01c6bffbcee8771c179ce09eea5105c330aee4e50b0bdcd9aa7923153d6c2029

Observation 071e3566-2153-4815-ba45-743320f94247 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality A Diagram Is Worth A Dozen Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.842285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.842285Z digest=sha256:d9b262d5f7ce7f199315a8c0eb6598e823956c3ffbfc0945ac3d35ef7a7a5fae

Observation 61f6fb42-0646-4ab6-9d9d-70cffb7bfef7 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.847152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.847152Z digest=sha256:0fe664b49e0eca069509ccb933cb896bcf1dc9dec0b55895ef13ad5bc2fc5dcb

Observation 593e1ea5-8e24-414b-bfa4-a69f352623d6 · outbound

This paper cites Redmon, S.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Redmon, S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.851633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.851633Z digest=sha256:71fa7d7402d2adc30988d3b061920f1b4b57f23dbac72a568d50d3cba64178fe

Observation 5e0164a0-d264-4499-a12d-187ea2623436 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.862036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.862036Z digest=sha256:64e2e6f56c4165a69e34acb30931bb5f33b2e7ba69f315e658e962c49f731fb9

Observation fe9e5a2d-8e5d-4476-a818-28d0902e21fa · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-05T11:11:20.868016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.868016Z digest=sha256:be67408f3196e6a6653b9f31bd305f03b6e7a1b91f026db15e1902725a9bb641

Observation 4035c6db-6ec6-4dd3-9421-f7f623789f78 · outbound

This paper cites an unresolved cited work.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Unresolved cited work

Reference 22

Resolution
verified exact
raw_fallback, observed 2026-08-05T11:11:21.162277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T11:11:20.873379Z digest=sha256:8d46020233fe8c8741410e3fabba22d84bbc011d32ea6df36a9a14d381e672bc

Observation adec65b2-bbe4-4c14-90bf-310ef96b25da · outbound

This paper cites Qwen2.5-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Qwen2.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.878592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.878592Z digest=sha256:e7d12eae4d936374629f2bd3a8a3654656b38ca8d0df6e24554e006d186fec4d

Observation 0c8d68e1-e4a3-49ab-ae44-cdf495f6c7a2 · outbound

This paper cites MiMo-VL Technical Report.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality MiMo-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.883362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.883362Z digest=sha256:5b241b3fde1b9ba7ab79740b51e9d46b52912ac7654f708935375f51a5832d74

Observation e788a6a7-88ad-4c27-add2-a97b5fd6fea8 · outbound

This paper cites KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.890509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.890509Z digest=sha256:03dd5570b6e30bca959d71911330897efe9babd55b14c4b6a05bf8ee3311b0a0

Observation ea359d5e-3ff9-41ae-9495-eb45cfce5b40 · outbound

This paper cites Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.896514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.896514Z digest=sha256:c7105d5a0bf9e0916d7553df64713085fae95397b8620a724156d4fa269fdca6

Pith citing papers

No inbound Pith citation observations are available.