Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2004.00849.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:09:18.398783Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T15:17:07.110860Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 546fb144-4606-4b62-9fba-c2f0036bb45a · inbound
GIT: A Generative Image-to-text Transformer for Vision and Language Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8337825-ee2a-41f7-b9bf-622fe1072f80 · inbound
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 41625290-b7e6-4102-898a-aaf729961f51 · inbound
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 197
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 17b0b189-121f-48ae-9ba4-4a09a08e92c7 · inbound
Agent AI: Surveying the Horizons of Multimodal Interaction Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 287
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2b3d7aa-2bb6-42cd-bbac-903041669d8e · inbound
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e90948c-19b9-4e50-8e10-f5988d1ed44f · inbound
Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbae891-510a-4397-85da-e265bd4f729e · inbound
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 161
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00f412f-09e4-47c4-ba4f-6a72d0084848 · inbound
MIMIC: Multimodal Islamophobic Meme Identification and Classification Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380d9bd6-798c-477f-baf9-a65940a3cca3 · inbound
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997aa0cb-9f17-408b-b1a0-f8a892d9fe35 · inbound
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df65f2e9-2856-4d1d-9b55-89350183965f · inbound
Foundations of GenIR Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52eeb334-1c21-423b-918a-a19ec82e28f3 · inbound
Visual question answering: from early developments to recent advances -- a survey Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 245
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b271ee9a-b7ac-4699-8be2-4a2b025c1f69 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82dfdea0-1cc7-4bf2-8f76-0f3a7d6007d2 · inbound
Improving vision-language alignment with graph spiking hybrid Networks Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278729c4-f92d-4e66-98dc-e91ce91b4616 · inbound
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 179
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276fe121-e048-4455-a1f0-4be44f6352e2 · inbound
GeoMM: On Geodesic Perspective for Multi-modal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ea6d0e-235a-495f-bd8c-67dcabef7be7 · inbound
RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4957cf28-d34e-49d7-8250-0910a178c91b · inbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a14ad9-84c8-4574-90ca-4e2cff0b5ee6 · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62cc5bd-370a-4985-a8b9-d5f8d555a58b · inbound
Kinky vortons in the 2HDM Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e19b0a-5d10-4397-bf3b-759b7091ba53 · inbound
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84c5ce49-402f-46c9-a776-285a972c0e64 · inbound
Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.