Pith. sign in

Paper Citation Record · LEDGER

UNIMO-2: End-to-End Unified Vision-Language Grounded Learning

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2203.09067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.09067 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:32:24.340654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:43:12.567797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45f6560e-3897-46a5-af50-566ea777affd · inbound

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples cites this paper.

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples UNIMO-2: End-to-End Unified Vision-Language Grounded Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:24.340654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:32:24.340654Z digest=sha256:80874de256147e53fdd7210f81371f7c1c10935697381d35d61a581620b29b17

Observation e642db5e-37b7-4813-9d80-47e8aa636e2f · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey UNIMO-2: End-to-End Unified Vision-Language Grounded Learning

Reference 254

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:29.268960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:29.268960Z digest=sha256:decb4ccde602d223ebe2a5143ef9687100e9e794d760743e28cb4f6876a1efad

Observation 7b2b4650-5df9-4f6d-9f0b-f383f0ea5116 · inbound

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization cites this paper.

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization UNIMO-2: End-to-End Unified Vision-Language Grounded Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:54.755656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:54.755656Z digest=sha256:19ef0b31adaa387cd8f9640aaf6826c60535816d8f6dd1ec7a153a0ba40edce2

Observation 4200575c-d8cf-4f29-8547-a760be096b83 · inbound

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook cites this paper.

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook UNIMO-2: End-to-End Unified Vision-Language Grounded Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.569437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T10:40:56.726020Z digest=sha256:4443cd6de9214ae18d8758f4d5d89309cde9f13aa1e3f9af449ebecf62e92a9e