Pith. sign in

Paper Citation Record · LEDGER

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 2 inbound Pith citation observations for arXiv:2504.13123.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13123 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:28.185701Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:02.739280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:39:09.371232Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df80247f-9dbe-4d6c-8759-a1859b2f72d9 · outbound

This paper cites GPT-4 Technical Report.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:27.999999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:27.999999Z digest=sha256:edaab208010c44586bc248b1f796e9ad726d140b01d381332000b039ab541d84

Observation 35ba0703-4018-4ea1-819b-dbc4735ea5ff · outbound

This paper cites an unresolved cited work.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:28.793655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.021547Z digest=sha256:f5e4f3c6d751b6b5843bb0d1562ef798e74ae1b958a0e9f175872eaab6ec9878

Observation 1b793377-d092-43fd-832c-35b96981fdca · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training PaLM: Scaling Language Modeling with Pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.041488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.041488Z digest=sha256:0d30975fc795a47de934d881aea2e15e3c2a3a75a8830d9a342cfadd016dad34

Observation 53b1445a-8dd7-40c6-9a1b-e65051dd1f96 · outbound

This paper cites Write and Paint: Generative Vision-Language Models are Unified Modal Learners.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Write and Paint: Generative Vision-Language Models are Unified Modal Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:19:28.569746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.047324Z digest=sha256:2d50858e6fff79dfbba7ff15bca9abcb316c68781e6cca31a39cde6916d0859c

Observation 1c7148b0-ee69-4ad2-b3a7-cdcdbb6b9c9e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training An image is worth 16x16 words: Transformers for image recognition at scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:28.756486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.052963Z digest=sha256:db6679a464e6115bd877e9e73bbd954a91b51800c8e19fa548b860e162f50ed1

Observation 982c87cf-2d6a-4270-aa18-a6d232f17a26 · outbound

This paper cites The Llama 3 Herd of Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.058134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.058134Z digest=sha256:1c38514ddbbe4d8cf79bc29f42d6a1a91926df6e3a48c0e4e612e09735e55aa9

Observation 2c1fc6d0-4b07-4ffc-8f2d-25f315d1a5b0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.065979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.065979Z digest=sha256:7e116ea7aa6b9033b08c48d16b34c8a90baf025f79b88b30f781856b7a399d4a

Observation 4dfdb7e5-e9fa-47a9-9f5e-be16e34f524f · outbound

This paper cites CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.070658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.070658Z digest=sha256:1001a732e395f8279cd17d3e972e7696c1c337a7ad4d91863bcab9a3bc95cccd

Observation 883350f9-ef5a-48ba-a37c-fd4de2882fbc · outbound

This paper cites Visual Instruction Tuning.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.086605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.086605Z digest=sha256:cfe945c04a8c9e180cba93ff0a08cdc139673ead9b57184436f78e5301541c39

Observation 20ab5fa4-0c8a-4954-84a6-01cdac79a04f · outbound

This paper cites an unresolved cited work.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:28.717588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.096890Z digest=sha256:345dcaf2a8a589fdf077f26af56b89777a0aeccfb9a488ba8fd69d97221edb3b

Observation bd060094-ba59-480a-aef5-77476307b95d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.101940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.101940Z digest=sha256:54b9ce3b14f263ec99290c410f81ccc7c097607627f18d94cffde8504155cf1b

Observation df753a51-f896-4d75-a577-43a4884a0720 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.106999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.106999Z digest=sha256:d14f75d7f0f84e883186f8c5af871dd0f34a2ef8a9545eb96eda4192c60f3fe9

Observation 400eec56-5304-47d1-a992-312df3bd0b83 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.111387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.111387Z digest=sha256:4855a0218cec2f4c8cef39691cf9cbfa8890dc36c7b8d7beae038bc1bcd4f2af

Observation 284a5f42-329e-4ffb-bbd7-4d565552ee9d · outbound

This paper cites FLAVA: A Foundational Language And Vision Alignment Model.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training FLAVA: A Foundational Language And Vision Alignment Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.122614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.122614Z digest=sha256:6f733cf9ae8d5f03877a77b28f51caf0feb353fb4cf8eea5823be098a9f71b5c

Observation 7f2388d4-4ad5-4f52-a172-f5362daad767 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training LLaMA: Open and Efficient Foundation Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.127343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.127343Z digest=sha256:e482e4902e61b27944dfaf53f12f29351003c8e49fba14d6f29edba4867778ca

Observation ac68f509-cb66-40d8-9000-f7d2aa8c660e · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:28.697173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.132176Z digest=sha256:80a55b1c6916410b643492a05a9d5e97c40f095ab9f8e33cb13d8684a39ebe71

Observation 1813411b-6ceb-48eb-b25e-a1be8ad0c85f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.143671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.143671Z digest=sha256:5f9a52cff92e5647d5cd2282ea08051659c4a497fea1a2a2ece81daba501990d

Observation 2809b66c-517e-4068-94bf-c1d847f00f4d · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.148409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.148409Z digest=sha256:bfad47ee769dce4bb7f2ac7da77bfe403fc3a1c69690d681962a5e50ab764bdd

Observation 2427627d-fc08-415c-b78b-22deac0bb1d7 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.153545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.153545Z digest=sha256:305cfb5565dc6df66fb77ae1db9831f9543445ec42c4823047931e022a52676a

Observation 3f33a485-a52a-49a4-a327-a3fcc13483bb · outbound

This paper cites Qwen2.5 Technical Report.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.159366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.159366Z digest=sha256:3027281ff869879ea5c4b869f3760a4a4f0685a883082a2703149673db1ab256

Observation 28a89139-89db-467c-8bdd-cea58f085994 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.169326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.169326Z digest=sha256:696878726a7ed61a6adaf86a6ee5aea1048411fd62168b24a6cd3b22f5bdcd74

Observation 4f592739-6d90-468a-857f-003e5b5c2459 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.175482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.175482Z digest=sha256:a6b4916be8366014b2b4479d36844eca264fee7a794a2aca6b5616f8128cf6b9

Observation 9a84a82e-fcc5-491b-8251-9570a7698409 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.180873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.180873Z digest=sha256:b94c4963c3963adab26ef8fd24b63c0175a915d3c7f47057a1d03e9c92d6d7a8

Observation c2cc7478-edd2-461d-8a31-25c568c76904 · outbound

This paper cites an unresolved cited work.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:28.679701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.185701Z digest=sha256:7567ca95f6be5a7526d708b76acd5f9a659fe081ae6028a151f3400378932186

Observation 8888dba3-c773-4b41-916d-3092aacc2eb7 · outbound

This paper cites an unresolved cited work.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Unresolved cited work

Reference 2011

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:19:28.739448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.092094Z digest=sha256:42ab929bfb81c9b7b153e836e6558d109cb7b030ddf569d2b4d7584159aa93d6

Observation 3080e3a3-abdd-42a7-bdea-6fe2c804baa1 · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.138230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.138230Z digest=sha256:a7887e7d977cb70454640be277463a6dd65b842011b02edb0cb2f606e6e6619f

Observation 1b5fc0dc-625d-4310-acdd-c5a1d37cbf51 · outbound

This paper cites Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:28.774852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:19:28.028615Z digest=sha256:bcc1c14dfd157446a15fbc57d965843fc429d3580a0fa50fcbff0e20f8311b47

Observation 31d2a297-0def-479b-939f-871f9e0c4932 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.034798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.034798Z digest=sha256:67aa8483e4e003bd87d500899f311953247b814669553644dc3c8933c6125bc2

Observation 1c67cd18-fafd-4628-869a-1946f968ed7b · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Hallucination of Multimodal Large Language Models: A Survey

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.013553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.013553Z digest=sha256:fb9179605bb9de2c533ce6027181d929a1c68f681c8f8919a4eb886611def47e

Observation 5b1744e7-b6dd-42ed-bfec-f381e83df06a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.007407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.007407Z digest=sha256:18c7a81cfd088fc9bed8e282c9b0b697b0b95fc9c4cfe98c2be4912d52cf9557

Observation 42d773f4-123a-4294-9cb3-a3e46099ffce · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.080863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.080863Z digest=sha256:3ede7611059ac41f622fcacfd9dd2723f66af8c8ac741c5a272b197d035219ca

Observation d6a61a0b-f893-490b-9aff-a2210c217a17 · outbound

This paper cites GPT-4o System Card.

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training GPT-4o System Card

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:28.076070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:28.076070Z digest=sha256:7dfdcb64776ebd400914333f37253c49ed8192c89a3943069f6d7c145144e9c3

Pith citing papers

Observation c188f1b7-3ace-4697-8ca8-53ccadd7f3ac · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:09.445066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:39:02.739280Z digest=sha256:ccdeb594647c0269c667516674ded935f21e2d50e9aaf7607fb65f1d799950ee

Observation 3a734a36-931b-4707-b5bf-109b8efadd8c · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.550837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.550837Z digest=sha256:31b78a01be6275c63043069944362bc404fe85b09935c55f2752a6d76346224c