Pith. sign in

Paper Citation Record · LEDGER

See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2301.05226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.05226 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:11:54.574723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T22:42:13.506064Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f804bb78-8fa3-4b9b-918e-eff6e51e1459 · inbound

Multimodal Chain-of-Thought Reasoning in Language Models cites this paper.

Multimodal Chain-of-Thought Reasoning in Language Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.534776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:ade2e483601a56481a6aa8106c634c2ed8e76bd799636570de15bcad8a290637

Observation 1e7d941f-6c12-4b9e-90a4-30d6526632be · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.574723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.574723Z digest=sha256:9024c107f1c1d4f94ea3f70171b9b33cfbe97eec0eda80836c0c8135c227b7c2

Observation fdbd6c0a-404a-4908-9f9f-0152b88d2277 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.682387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:aecd294c4bfb0f44877cc636be5c74be399d67a2e1c5bb7639bcaa30efba692f

Observation 0cc0e3a7-0202-4161-86e2-959667e1bc7e · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.509384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:4292fa6e59fb27d25aa05e662a460d743665b9c521b4bc8cb8395941e18d575a

Observation 431d7a65-b4a8-4db9-8a57-5d0e8b9fee88 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.800954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:55.800954Z digest=sha256:097f6536797de3ea751fc01d6aa3160afa22e76c98874919b1958e4c44a17756

Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.881851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.881851Z digest=sha256:a0e789caaae01fe4fc1a18693ad3978068ae6ab74b08ad9ef8bb4dd8e6f11cf8

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · inbound

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning cites this paper.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:8e408b480cc43b6c256906683fd23db3711ec5e77d8dfe573851b79a1c6e59f5

Observation e53b242c-3402-4e09-9564-67467c7bba45 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:02:52.625798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:0c8e2e117b95ee8499ee52776e5a52b79664affb1cf9ab9f6d9a1e935bfdca04

Observation da99e7ad-f238-43cd-86f2-4d729a8d488d · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:15.141748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:15.141748Z digest=sha256:6abf08137e09b5570b168b4ac5f7bfd270c4fbf4c4c254f76542052de2f43189

Observation 535c2719-7ef5-4792-bad4-758da2518608 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:01.504726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:01.504726Z digest=sha256:b43f6cd0f78800793314c2314ef6a184e9a70c1f5a4d98d6d0202f3606614815