Pith. sign in

Paper Citation Record · LEDGER

mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2205.12005.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.12005 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:28.217225Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:56.632867Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe4fe666-abea-4baf-a398-efc8c3d7cea8 · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.669979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:71ba996a4d533e096d8e1487ed7dcb426e8b0d7256e7cfd22c367b9477d8295f

Observation cbc7e6ea-4b69-49b7-9690-b57795862f74 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.283540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:5f72e9cee12daa76d58ff141394fcd56346968bf33f55106d01e54db1f564781

Observation 0547117b-2545-43d3-8918-a74655c7f44c · inbound

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment cites this paper.

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:43:03.449820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:43:03.310237Z digest=sha256:c6ca810fbc591ceb461b32f785ce2b463e6a1ae79a684b0488b2c234e1a740b3

Observation a22d27b4-0d92-4ae2-a3e9-4c5822de9a63 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.425941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:c238a083d1f217d1dfa661a2b1d669cca3981f9a2ecd5ec7f9d0dea7d4b737ef

Observation 23b3c1fa-12d1-4479-8187-630a5c314b33 · inbound

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models cites this paper.

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:16:43.715851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T15:15:19.523035Z digest=sha256:6424538844f950c839ae0b50d2991377146ba9810e1a078a974b0f42fe2f6784

Observation f4722c6c-9f1a-44ad-9e7b-9df5dcc26f57 · inbound

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? cites this paper.

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:28.217225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:28.217225Z digest=sha256:ae80e8b05ec080ee51b8e97c69de11eaf84d10adc725e830219749d635faa60a

Observation b0f23c11-628c-4107-8673-7a996920c4ab · inbound

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning cites this paper.

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:06.418154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:06.418154Z digest=sha256:ae99e3c0ee75b0697adf43dd633b8b279bf65328dc45bf5ea1d189a58f4c5189

Observation 0a0b8c8a-8ab5-424b-8abc-6111086bdfc5 · inbound

ReFrame: Rectification Framework for Image Explaining Architectures cites this paper.

ReFrame: Rectification Framework for Image Explaining Architectures mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:59.764651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:59.764651Z digest=sha256:5b53b5ce28070756309e83a3727b10541c0862a47b8a3f543ba923cb61dc1f3c

Observation 650fe551-ea2f-416f-8532-9aa34c0f1079 · inbound

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling cites this paper.

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.410901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.410901Z digest=sha256:5d2b57e9ce4b06c1a7e1d1987cf179eff2b1a45ad1900d1f99c379ae41560398

Observation 1c2d67bd-b0b0-4603-9b6d-329e2e9c975a · inbound

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection cites this paper.

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:46.361810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:46.361810Z digest=sha256:97c910d5d27db402c314e438a8cb78e7684d4b169275a9a53ff8d8f6d29e0f84

Observation 757067d8-2202-4a58-84c5-823ce07363e1 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:36.456455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:36.456455Z digest=sha256:8cba5ec9daad37a63025496956b0695646cb7de782ed266e03e7c7262e67e1c6

Observation c7014bc0-822a-4cee-8cad-1eca19200620 · inbound

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback cites this paper.

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.380981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:26:42.034972Z digest=sha256:872bde7b7df0e3b82e34a4bf690fe9d2b0552a877af7d4b3796fe48f5de25ca5

Observation eeb09fb3-51ba-4736-a08f-288dded6a40d · inbound

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning cites this paper.

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.818843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:46:22.891034Z digest=sha256:db5cb82b859f301ccf95be6137a28b7c459604d407a9f7f99b76b332d057c72b

Observation 55733916-8a99-4c15-871f-0e32c9a9a304 · inbound

VisChronos: Revolutionizing Image Captioning Through Real-Life Events cites this paper.

VisChronos: Revolutionizing Image Captioning Through Real-Life Events mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.634915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:34:38.597959Z digest=sha256:d0ea63e9f9b2b783d93702b2a56f6f8b4d50f1e83c2f058eacead9e397fc7450