Pith. sign in

Paper Citation Record · LEDGER

InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2003.13198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2003.13198 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:02:31.632210Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.933033Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a642598c-7b87-44b7-8bec-c36abd4048e0 · inbound

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey cites this paper.

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 263

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:31.632210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:02:31.632210Z digest=sha256:c6bd97b8019d95e01cfbfa88553e6b9f83548239e25f1b7cc97529e131f4388f

Observation 939adbb6-a456-46c7-992a-95cdb71130bd · inbound

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement cites this paper.

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:33:06.885125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:33:06.885125Z digest=sha256:03fc7965d8d84f7c047a99d8edab0adb62b25fa3caed965aa8030df9350d4159

Observation e8adc19f-c2b5-4ba0-8cf0-85c9d03affb5 · inbound

Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search cites this paper.

Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:55:24.071530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:55:24.071530Z digest=sha256:57e3d95b78a47235f0f9633bc3dc565f57f21a6a1fe9fac7aef60c841e7fe2ed

Observation e69e4afb-502b-4b3c-88bc-c5efe01eba00 · inbound

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains cites this paper.

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:20:27.609924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T23:19:11.439175Z digest=sha256:bd411aebd940db29b0ac6ef6748c6fd8adaad829c5fae2a91884d24bde6a3a76

Observation cea11255-1ed5-44a1-a8a8-507e810d7400 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:43.934461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:27577e05aeaa888be289739acb46adc6659a32c212ad7e9df86b8a1a95cd61d4

Observation 0e5a0528-1149-40b1-853c-d2b2a835cc56 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining

Reference 146

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:ffcdafda297dcae83eed031cab354b7ccee164232387587e88088ad4ea33c406