Pith. sign in

Paper Citation Record · LEDGER

CapsFusion: Rethinking Image-Text Data at Scale

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.20550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20550 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:27:56.303341Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.507046Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b687545a-f19e-447e-9c3a-50662f4f4cad · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding CapsFusion: Rethinking Image-Text Data at Scale

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.751043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:f581d676b32aab688ec9e2ab6f763fe3f75d40b52d1697fdbdc4419117af0bba

Observation 2b097277-fe44-431b-a487-4ff794e18289 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation CapsFusion: Rethinking Image-Text Data at Scale

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.115040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:493309ad718797ed9e1786aa287fc4a93a08727ceb42e214d7121a4b14866c50

Observation a7276b04-4dc2-40a0-9297-790122fa7487 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding CapsFusion: Rethinking Image-Text Data at Scale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.488355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:f48ed8d949d624bdaac0d860f2b2a239e2c8b02bc6b8405238c68184ce814f70

Observation 126d5ada-14ae-4a02-936e-105d0ab7aa12 · inbound

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability cites this paper.

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability CapsFusion: Rethinking Image-Text Data at Scale

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:56.303341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:56.303341Z digest=sha256:6775d07194e315dde11e9c0f466fb82a1b8d7995e3927227b39a9a8acf59cf3f

Observation 282cb1db-1d9d-4e8f-86a7-92e62621b0dd · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapsFusion: Rethinking Image-Text Data at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.450760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.450760Z digest=sha256:106b1cf4dafb6431a4bfad69b8e40d67b257760aef24cc4242380dd69caaf4c6

Observation 85b5c121-2c3c-4e77-9e74-d93c9127a5d4 · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models CapsFusion: Rethinking Image-Text Data at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.543788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.543788Z digest=sha256:5ac50fa7d9c153ff262b581a5f04297074c001a636ee1965953e820ba8875d64

Observation 3191fa67-74b5-4f33-b0bd-8457f0ec7b40 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data CapsFusion: Rethinking Image-Text Data at Scale

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:36.678781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:36.678781Z digest=sha256:3cd221caa1220212d4e3ba62c1e102bed0595124d419df7e2e03c54eacbbf2a3

Observation b3543965-c233-40b3-83c7-965407277f72 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning CapsFusion: Rethinking Image-Text Data at Scale

Reference 250

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.508435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:03452d976288a998e40f89dfed873b56d82e1d82f318ca54ced0b80a733c7408