Pith. sign in

Paper Citation Record · LEDGER

CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2401.02582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.02582 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:49:01.434232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T09:51:13.622166Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc5f59a9-aab2-4e61-aaa6-dc25118dc086 · inbound

A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs cites this paper.

A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:01.434232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:49:01.434232Z digest=sha256:af84c2271f2b911a32295de1e791c88b06668ec7b756cc6257e8fb70ffbdfa64

Observation d65a3995-e35e-4557-86fa-76ddc471beb9 · inbound

Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources cites this paper.

Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:22:30.735595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:22:30.735595Z digest=sha256:f55d95f89ea4b766ff30515993b9d84f006aab01c195e5cbf8c6f462bec302a6

Observation a4454b98-c9c2-4fcd-857e-251e059622f2 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.683745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.683745Z digest=sha256:24a0d40ed680edaf9e53575cec996d18121c3496339c57ca8fc27db96b13a801

Observation af53e087-66fa-4a22-80f5-43a94b385d8b · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.703843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:5a6faf15e29e51fa47674ead534ab064b1fda2bb8c373eec0610e28bce5600c6

Observation 8670db17-2935-4c00-b800-55c2984d1be3 · inbound

Can Multimodal Large Language Models Understand Spatial Relations? cites this paper.

Can Multimodal Large Language Models Understand Spatial Relations? CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:07.427455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:07.427455Z digest=sha256:db0d5a882cfb03f35f62c64361bc61cbc8306c4e7d311910920c6c65af2b7f2b

Observation 5793574b-1f18-435d-8eb9-a6f95939c91f · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.064735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.064735Z digest=sha256:2d18eaa13b781cea6cf520fb71af1bb790765c0cf0dfb240aed13c97960fbc5a

Observation db9e8f2c-82da-40c2-9525-65d048c539d0 · inbound

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge cites this paper.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:44.018505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:44.018505Z digest=sha256:9e568aa16dc43b457616c848b4e569d558676024fea9407b9d7d2e360cc0e6f6

Observation 90347a6c-eb5a-4ac5-9b91-3134c0d0cba6 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:36.765868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:36.765868Z digest=sha256:bd870e4bc565faf7e14277ca62b23c99df95f82ecda5367cd7fecf203e36373e

Observation 3d26c41d-b290-45bc-adf5-eab104e8dd3b · inbound

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior cites this paper.

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:55.887586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:55.887586Z digest=sha256:29babd44f9b61f5331f608a04de40cdcaeca2a3ab06e8ba770174ed518711475

Observation 6e96639b-de49-4f13-b5ff-0d16025061b7 · inbound

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models cites this paper.

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:51:13.625048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T09:48:39.943352Z digest=sha256:63e79441cc0130e1c079ba88fbad73b8c82a34f0575771e6ea1c852006c3b0cc

Observation cd670699-7307-452d-85fd-260683163256 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.075026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:43.075026Z digest=sha256:d5ee6ae590f855f5540c43c3fb02af3532204303706d2253b3c16c7ddd99992d

Observation f0f39b06-f2e2-487d-9d91-f9b1e1246ea0 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:08.021152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:08.021152Z digest=sha256:68e2b90a06bb39cea9d34be7f99fab316900074e0cec6cd65cc1c6f625cfbb00