Pith. sign in

Paper Citation Record · LEDGER

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2102.08981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.08981 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:56:46.142751Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T02:48:44.958361Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5e6d4d93-75e6-459b-9683-df9a65d5b809 · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.919850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:e11fc668bebe99b893b719f51c35b36a6026792a71a43b142f323721030c2431

Observation 4f631775-823a-447f-ad2f-eb68bc24f782 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:44.961776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:36d69f7de9727b1fb84df77f6401256f0f59194cad3eece8003b083b424b2309

Observation cc89895b-463a-4be5-adbb-c83152f070e5 · inbound

FILA: Fine-Grained Vision Language Models cites this paper.

FILA: Fine-Grained Vision Language Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.142751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.142751Z digest=sha256:80752642932fb95be755440e0dd945f3ae7e8df2e666b619545d6d1df6f48dcc

Observation 4b1d9c74-8136-437f-94e8-736a37685071 · inbound

Native Segmentation Vision Transformers cites this paper.

Native Segmentation Vision Transformers Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.530639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.530639Z digest=sha256:8bda7ae4b75b7f273a3508794274a539ca884a9c0e8cfa86c09d34525dd19e4c

Observation 24bc5e7d-237e-4aea-a8b8-1b81580e35b4 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:59.063835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:59.063835Z digest=sha256:55df3addffb06ac97eeb8294d13ecf22a3fcb9deb963105d2ff512f8120e6221

Observation 2ad43a37-cd7d-49f7-9fb6-bd9a39efa155 · inbound

Entity Image and Mixed-Modal Image Retrieval Datasets cites this paper.

Entity Image and Mixed-Modal Image Retrieval Datasets Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:45.661946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:45.661946Z digest=sha256:f7438548a13522158453886e733cb7d626855e24e6848018dfe81fb3f9aa908e

Observation 8dc5c7c3-b4f4-4086-9b32-c5b43d10ab15 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:59.977186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:59.977186Z digest=sha256:2a4d826e536b7eeb19950dc5efd33d423ffb94ce54f18ceffc17ec7341351e52

Observation 896a3b3e-2969-4853-bef4-6ce1b59f5d62 · inbound

Info-Coevolution: An Efficient Framework for Data Model Coevolution cites this paper.

Info-Coevolution: An Efficient Framework for Data Model Coevolution Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:32.928691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:28:32.928691Z digest=sha256:f057a6f4ae5ab6d1163a07e9d3b742885af81603e0134b329d6334229ff2f19b

Observation 56413fac-9e7e-4d3f-84a1-11cbf214d70e · inbound

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation cites this paper.

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:00.098365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:00.098365Z digest=sha256:909cd64c565a51ec4adc251f4c17b8c58a307ff1b90fab1b922b167b6950d44f

Observation b797cd48-dee6-449b-8469-a669a19acbf1 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:56.083452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:56.083452Z digest=sha256:0ea7e93965b931653835c12f6f7a1b39ba42c44a8d9b7942aab56c21357ee190