Pith. sign in

Paper Citation Record · LEDGER

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2311.03354.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.03354 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:07:33.048750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.091209Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab28b99f-416d-4fc3-9526-e3528ecc82bd · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.278672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:ae8f80e17ef64baefaa4de1f853b2489ad7c330d1a42e471337b9b0981c584e0

Observation bb030ae3-1284-4822-a24a-017b5584ad1a · inbound

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM cites this paper.

EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:07:33.048750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:07:33.048750Z digest=sha256:a4d8fa2f1ad5e2236097e61e32c35345c4531f98c0965f17c401b6d8f04a51a1

Observation cac99366-4373-490d-94e6-68415c833b96 · inbound

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor cites this paper.

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:24:58.539838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:24:58.539838Z digest=sha256:32fdf31668a928a89523a2cfcba699666bddd4fa8583a36683ca780dafc38a3a

Observation 70171532-8b38-4b87-800a-21addb893a3e · inbound

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models cites this paper.

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:17:06.194020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:17:06.194020Z digest=sha256:85b7a6cbbe467dba2d5abd5d37896a0452ec2980738d30f07b41129da962233b

Observation 07dd34a8-6821-4dd7-bb14-7233e708c526 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.093837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:2ea30665ebca40d22fef53a39d8e8fe7723d9258597b5c28316f61774238d932