Pith. sign in

Paper Citation Record · LEDGER

Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2410.02740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02740 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:34:14.251378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:05.116399Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a833fbd6-1add-4ddf-8647-8e803f910412 · inbound

Multimodal Autoregressive Pre-training of Large Vision Encoders cites this paper.

Multimodal Autoregressive Pre-training of Large Vision Encoders Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.609197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.609197Z digest=sha256:148e3b3035d774285c1c36b517b010311ca1b5bb4002261413d7447a25adbf5c

Observation 01a791ac-b095-4a8e-b28d-7419fd44d011 · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.471331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.471331Z digest=sha256:a8a3c8ed7435b8f51e9fb188c4ab0cbb99948a156f92d4a8485d4be58492d217

Observation 690ccd34-cd30-4a63-bdf3-caf85380e7c7 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.119696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:eb94fd7ad06c30c7a7dc6633f4e8d47daa38d1464441264487d2b0f80476c210

Observation 1592e1d7-4a1c-4987-96b0-93cab0612cbe · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.285204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.285204Z digest=sha256:e89b6934420f171820001dd6e2647c09f468e902db64a1ebbef78ccea67151f0

Observation 2648b533-3d63-4736-b6d9-3a9fd2ba2285 · inbound

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers cites this paper.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:14.251378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:34:14.251378Z digest=sha256:13f15a50c7470f9a7840b9c3ec050416974cc60f42665e4ca97cfe19394f52bc