Pith. sign in

Paper Citation Record · LEDGER

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2202.03052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.03052 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.072176Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T18:15:14.590918Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2984da1e-7e1d-4986-91ec-658eb3347da7 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.292817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cb27554f087088f0e9170d820bfcb2fc3e70fbf4da9f014fbbf0d1f2ab0ac81c

Observation cb6d162d-7fc2-436d-bea7-e30364cc991f · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.391341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:0c6632120dcca88ef58dc5d1363057dd3f2ae9e839311fb892085c9f23e79f02

Observation a9cbffaa-0912-4a70-9df9-fab7bd93e46e · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.723201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:0fd2df8431750c9c12f3000af2fba76e3be7af571afa0e81b9a8e4da2182f3b1

Observation 75c31969-c348-48af-9347-9793304c8979 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.240416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:66d13d926f6f86bf0e9c45f5f145e00fa9183a97fbcaecaa66719b74c714847f

Observation 3d71af40-d070-4805-a9de-c6fc64f9c9ec · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.595172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:5a47adb6a2de617d06a0d16e677b87a0013304ed53995a4580ac2a3ba6fb463c

Observation fd024c18-b775-46ce-af79-d6fafcb83187 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.175626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:6d048b1a1bdde001cb51e93c886040d282eabebd0532bbe92f11cac04ffcd4c3

Observation ad699342-6b75-4888-a562-47e8d54a4a59 · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.072176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.072176Z digest=sha256:96ec978c03dabebebf02feafb1a80f6535aa95d03e07c2585603ee66c37515c4

Observation cfd59bb3-1d50-4eda-b3ef-a63716c5f551 · inbound

VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall cites this paper.

VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:57.436866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:57.436866Z digest=sha256:0c79b76af391f2857451349af097ae805e315a31511d7f6e33404444eb42fd61

Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.037290Z digest=sha256:a596237d81e107a0202dc93f27d15c0b5ab7e699cee6ea236dbbaa7090a347f5