Pith. sign in

Paper Citation Record · LEDGER

Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2410.02740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02740 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:34:14.251378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:05.116399Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a833fbd6-1add-4ddf-8647-8e803f910412 · inbound

Multimodal Autoregressive Pre-training of Large Vision Encoders cites this paper.

Multimodal Autoregressive Pre-training of Large Vision Encoders Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.609197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.609197Z digest=sha256:e64cccb0e090de4e79f153d018c154c22695d880792012ad57da38073bb6e12d

Observation 01a791ac-b095-4a8e-b28d-7419fd44d011 · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.471331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.471331Z digest=sha256:04a998cb1e4cc97b9957e1ab1159d9d1513cf8d205945ab86f8d8a1953f8e6ed

Observation 690ccd34-cd30-4a63-bdf3-caf85380e7c7 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.119696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:efac4bee14742626b117858606f4ede3458f137460f27b0b737785b61a48fb8c

Observation 1592e1d7-4a1c-4987-96b0-93cab0612cbe · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.285204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.285204Z digest=sha256:8fd6c941a445c0f7ed0904e5d1b1c5e2eba34473b90512782479f56c31b816df

Observation 2648b533-3d63-4736-b6d9-3a9fd2ba2285 · inbound

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers cites this paper.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:14.251378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:34:14.251378Z digest=sha256:187d2be001da18864a2f1d8765f632a9e26d75c79a02ff0ead6d58cfd6c4f064