Pith. sign in

Paper Citation Record · LEDGER

Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2004.00849.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2004.00849 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.398783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:17:07.110860Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 546fb144-4606-4b62-9fba-c2f0036bb45a · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.656061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:40a176d6d00f8682b75de3ad9e84befc2fb4ee32ba73ca3b498566ecfb283397

Observation f8337825-ee2a-41f7-b9bf-622fe1072f80 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:06.416304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:45b37156054445f7ce02d1299b71d1105b2161e2a90af219cb79964d3da8f63e

Observation 41625290-b7e6-4102-898a-aaf729961f51 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:27:59.078329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:84b5878b7e524772e1c249b29320b2a4bbac3b23bcfabe23884fc70ec4f9cf59

Observation 17b0b189-121f-48ae-9ba4-4a09a08e92c7 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:25:59.812425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:c7edaa4c8033098ec30ab3e38d11287dd5021082796291ddb7ed367fd8c9408b

Observation c2b3d7aa-2bb6-42cd-bbac-903041669d8e · inbound

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms cites this paper.

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:54:50.103584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:54:50.103584Z digest=sha256:e6043d22eae7a9621606a3f5475af0008ecfe977d32496296f7025cd5b04b424

Observation 0e90948c-19b9-4e50-8e10-f5988d1ed44f · inbound

Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval cites this paper.

Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:05:57.824935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:05:57.824935Z digest=sha256:f77c5db972c8796c9356d4cc2f1a801dfba8aa254bee12fc16a0a5ff3307db4b

Observation ddbae891-510a-4397-85da-e265bd4f729e · inbound

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey cites this paper.

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:31.150679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:02:31.150679Z digest=sha256:4e02542bb13317134cf8288e9378f0104fdca732195dfebda904fb786fed037c

Observation f00f412f-09e4-47c4-ba4f-6a72d0084848 · inbound

MIMIC: Multimodal Islamophobic Meme Identification and Classification cites this paper.

MIMIC: Multimodal Islamophobic Meme Identification and Classification Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:10:10.345535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:10:10.345535Z digest=sha256:03e97c4ea8f143616addecf936d99aca976efb1fb0e93602c85ae412f3a28095

Observation 380d9bd6-798c-477f-baf9-a65940a3cca3 · inbound

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples cites this paper.

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:24.293617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:32:24.293617Z digest=sha256:259960b6f20fcd06a55e1c0e4b786756f191b9a813571529153423b09cc4b539

Observation 997aa0cb-9f17-408b-b1a0-f8a892d9fe35 · inbound

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval cites this paper.

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:32:44.256855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:32:44.256855Z digest=sha256:6f6a6ed59e57e505f234a3363a30e1bc4f215e58fa714a19ebd5f811440f7584

Observation df65f2e9-2856-4d1d-9b55-89350183965f · inbound

Foundations of GenIR cites this paper.

Foundations of GenIR Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:06:04.892482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:06:04.892482Z digest=sha256:073a49b1ef01012a4901a45276af6713ccd6a0f4d36f756f9de3902b59ac36f6

Observation 52eeb334-1c21-423b-918a-a19ec82e28f3 · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 245

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:29.229637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:29.229637Z digest=sha256:4eae56de93d7bcf4f55e7ad0f69ba16d5cebf6cd9c32a85c822cb456cd146a55

Observation b271ee9a-b7ac-4699-8be2-4a2b025c1f69 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:39:22.424251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:cf8d457198a185e4b9abf007b742bbce0bad02699f634e788e5e27973afcacbe

Observation 82dfdea0-1cc7-4bf2-8f76-0f3a7d6007d2 · inbound

Improving vision-language alignment with graph spiking hybrid Networks cites this paper.

Improving vision-language alignment with graph spiking hybrid Networks Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T21:30:14.667388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:30:14.667388Z digest=sha256:05dcf170208bd765a7f7b9ba1af9cfa6b8ac17971f5ac48edfbacc784f1b8f3d

Observation 278729c4-f92d-4e66-98dc-e91ce91b4616 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.739735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.739735Z digest=sha256:692a6a799f65ae156c2405e131327018cf1c63e8834f9f2a091d338863f4f876

Observation 276fe121-e048-4455-a1f0-4be44f6352e2 · inbound

GeoMM: On Geodesic Perspective for Multi-modal Learning cites this paper.

GeoMM: On Geodesic Perspective for Multi-modal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.561792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.561792Z digest=sha256:5a257625ed9efca6d9230a7a34616d25dff62df18e3c2ac9aed24a963a3358fc

Observation 23ea6d0e-235a-495f-bd8c-67dcabef7be7 · inbound

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer cites this paper.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:39.997161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:39.997161Z digest=sha256:6d8cd962c1570e1c4ecd59827a598d6c0239ba10ad82997da761193bdf05ecac

Observation 4957cf28-d34e-49d7-8250-0910a178c91b · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:50.658973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:50.658973Z digest=sha256:fee8718400fc8269399994c525cdd409f3211d48394f2e01ce13054bd628471c

Observation 47a14ad9-84c8-4574-90ca-4e2cff0b5ee6 · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.772130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.772130Z digest=sha256:ef413121ec7678dbc4eb403adc7d232053612793a6fa28127ac93542c1164ab6

Observation e62cc5bd-370a-4985-a8b9-d5f8d555a58b · inbound

Kinky vortons in the 2HDM cites this paper.

Kinky vortons in the 2HDM Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T20:29:13.300410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:29:13.300410Z digest=sha256:534ac7245da474ec748f68fb316c0882fbba2456648cc83981efaea09188b13f

Observation 05e19b0a-5d10-4397-bf3b-759b7091ba53 · inbound

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding cites this paper.

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.113515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T15:15:24.784685Z digest=sha256:e538e410b7a0c5d1c725cc95fdf4dffe3d4c87449548bac46e815d9906cb9de1

Observation 84c5ce49-402f-46c9-a776-285a972c0e64 · inbound

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding cites this paper.

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T00:09:18.398783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:09:18.398783Z digest=sha256:edac3f578421bac1de88b60a772bf9a780ae1f55f836c817e0ca8ef933bfb2c3