Pith. sign in

Paper Citation Record · LEDGER

See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2301.05226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.05226 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:33.004657Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T22:42:13.506064Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f804bb78-8fa3-4b9b-918e-eff6e51e1459 · inbound

Multimodal Chain-of-Thought Reasoning in Language Models cites this paper.

Multimodal Chain-of-Thought Reasoning in Language Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.534776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:e46a7f221f71ab4a1df49ac028c90de5ee4727afe07c75dcb736756b21467a63

Observation 1e7d941f-6c12-4b9e-90a4-30d6526632be · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.574723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.574723Z digest=sha256:8b269200e681cfda81791b6f5e1df001bfe805ed98979a42d1af98f4352e203a

Observation fdbd6c0a-404a-4908-9f9f-0152b88d2277 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.682387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:43358a31577b3329c8d8f958e1cbeb9f3a38f6a07023b6fddbda8de8763db3be

Observation 0cc0e3a7-0202-4161-86e2-959667e1bc7e · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.509384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:b652507508f867db5f39acd4625808face26c58763c961eb2b3db2e452399f7f

Observation ef3516b1-70a7-4167-a7cb-03482615c18c · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.168772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.168772Z digest=sha256:0408415da2c5c8a5bbd1a62ae70079a6ea05000ed712a8e2cbe98cbc38f944ff

Observation 431d7a65-b4a8-4db9-8a57-5d0e8b9fee88 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.800954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:55.800954Z digest=sha256:afeee4abba7820ec4b1818744275f7c7af698a5e8114f4cb715b38f3947bf511

Observation 1500536c-e013-453b-9790-3876a40af7fa · inbound

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach cites this paper.

Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:33.004657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:24:33.004657Z digest=sha256:4793188750c85abb9725a16d253ddba6d741dcb6b9d0a90a17935f691c746a8c

Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.881851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.881851Z digest=sha256:4eb1dd6a26918571647c22898c428fcd73dea5cc882abd2b0f6ea2bf553afc01

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · inbound

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning cites this paper.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:6734ce7eac25d64894a6bae5f78d131c01f9876cd6679c899e8aa4c2c5adce4c

Observation e53b242c-3402-4e09-9564-67467c7bba45 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:02:52.625798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:d090a71bbb81f5d97a7590ca81013c1ca6dc8463150847489497fd0a0b060913

Observation da99e7ad-f238-43cd-86f2-4d729a8d488d · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:15.141748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:15.141748Z digest=sha256:bceaadfbd607af240aebeb1c934798fb71facc325afa04acbb371f5ceb168eb4

Observation 535c2719-7ef5-4792-bad4-758da2518608 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:01.504726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:01.504726Z digest=sha256:e0b4fe4c5b43373448281974fd1e64cbb8f28ff3eacb0a44ae3f7423fe71b054