Pith. sign in

Paper Citation Record · LEDGER

mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2205.12005.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.12005 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:28.217225Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:56.632867Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe4fe666-abea-4baf-a398-efc8c3d7cea8 · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.669979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:2273f4b757f9801b83cb5d3400e33e292a6aba5540b573763c517767a6d6f17c

Observation cbc7e6ea-4b69-49b7-9690-b57795862f74 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.283540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:fd01cdc5398a8df6dd7ace164fb3f8dd7495f7004134c11afa7f94f32d52f3cf

Observation 0547117b-2545-43d3-8918-a74655c7f44c · inbound

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment cites this paper.

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:43:03.449820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:43:03.310237Z digest=sha256:e6351f8a2933a56e55d39afcf5d4e7b39fe679ce7a40da76731df3a59cebcd39

Observation a22d27b4-0d92-4ae2-a3e9-4c5822de9a63 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.425941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:ae002e344736b6575a3685239c00799642fd2b95867f720e00f127534aabc30f

Observation 23b3c1fa-12d1-4479-8187-630a5c314b33 · inbound

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models cites this paper.

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:16:43.715851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T15:15:19.523035Z digest=sha256:0a8c607e57f426f50c403fec9ef8f15bacde7b6623d3d6be1af5b44fa6908324

Observation f4722c6c-9f1a-44ad-9e7b-9df5dcc26f57 · inbound

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? cites this paper.

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:28.217225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:28.217225Z digest=sha256:ae80e8b05ec080ee51b8e97c69de11eaf84d10adc725e830219749d635faa60a

Observation b0f23c11-628c-4107-8673-7a996920c4ab · inbound

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning cites this paper.

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:06.418154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:06.418154Z digest=sha256:ae99e3c0ee75b0697adf43dd633b8b279bf65328dc45bf5ea1d189a58f4c5189

Observation 0a0b8c8a-8ab5-424b-8abc-6111086bdfc5 · inbound

ReFrame: Rectification Framework for Image Explaining Architectures cites this paper.

ReFrame: Rectification Framework for Image Explaining Architectures mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:59.764651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:59.764651Z digest=sha256:5b53b5ce28070756309e83a3727b10541c0862a47b8a3f543ba923cb61dc1f3c

Observation 650fe551-ea2f-416f-8532-9aa34c0f1079 · inbound

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling cites this paper.

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:19.410901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:19.410901Z digest=sha256:5d2b57e9ce4b06c1a7e1d1987cf179eff2b1a45ad1900d1f99c379ae41560398

Observation 1c2d67bd-b0b0-4603-9b6d-329e2e9c975a · inbound

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection cites this paper.

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:46.361810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:46.361810Z digest=sha256:97c910d5d27db402c314e438a8cb78e7684d4b169275a9a53ff8d8f6d29e0f84

Observation 757067d8-2202-4a58-84c5-823ce07363e1 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:36.456455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:36.456455Z digest=sha256:8cba5ec9daad37a63025496956b0695646cb7de782ed266e03e7c7262e67e1c6

Observation c7014bc0-822a-4cee-8cad-1eca19200620 · inbound

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback cites this paper.

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.380981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:26:42.034972Z digest=sha256:a715610f7eb57dc2e1ba8a62b4769b423b5a22903a991bb428fea8ffb512e9e4

Observation eeb09fb3-51ba-4736-a08f-288dded6a40d · inbound

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning cites this paper.

AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.818843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:46:22.891034Z digest=sha256:6edfe3fefc528e11831899fb1bd3a582d5bc266fb71d7895eb896b10579ad55e

Observation 55733916-8a99-4c15-871f-0e32c9a9a304 · inbound

VisChronos: Revolutionizing Image Captioning Through Real-Life Events cites this paper.

VisChronos: Revolutionizing Image Captioning Through Real-Life Events mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:29:56.634915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:34:38.597959Z digest=sha256:d9df6adbed0bf3d26e42b6f69ca0aed8ee1c36163409a2ce76161aa310bdf008