Pith. sign in

Paper Citation Record · LEDGER

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.11571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11571 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:36.844173Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:54:33.563266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0abaabf4-c43e-48f3-8d8c-ed11ae2f1b6c · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallucination of Multimodal Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.587700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.587700Z digest=sha256:a1f93ed03901ad44ad85688c557e1a0f5e1fd4cc9d57068ee090d6f6d3d877cd

Observation de2de84e-4f9d-4c43-bdca-87a18ce0ff6b · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.660752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.660752Z digest=sha256:a91c358d4c12538dfd91f04a6ff9335bd024f291759fb00ea64e2e1445043d97

Observation 7637cb9b-dc85-4776-8ff2-9efccd7b9cd2 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.703713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.703713Z digest=sha256:3a433317b59847043abd7789fc8fa9330b21e03425f6ea6870c212c2d36a47c7

Observation adb9539c-bbe9-4d46-bc61-7e061c22ee24 · outbound

This paper cites MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.778408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.778408Z digest=sha256:ca8ceda80ceb76b5b3de5e4ccfa1cf4dcca5f590725f53bc58c6d129eb8bb8e1

Observation de2f71a6-6d3e-42c6-bc3f-951120f45647 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.820419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.820419Z digest=sha256:566443f7644e2fe63e27b204599a8c137124f07ed42f7ece09292448d3d05b86

Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.881851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.881851Z digest=sha256:8973ac6a424e9f7650ade4a9b8fa3ddfcbb2aa01d09766859a0b648dcae6b7f1

Observation bbef4c86-23a9-4ba7-b7a3-cc1c5e026a43 · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.956173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.956173Z digest=sha256:a05a13bad8976b35f61c361f1fbddcd94c62047065ad3c63cbe4787db4f8df6a

Observation 90ecb6a8-ce9f-4510-9908-270fc34ff20f · outbound

This paper cites Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.061606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.061606Z digest=sha256:c9cd864c14b0a38cf5b0e0a5ce75acce00351c9104cf48d37479dfba52f25bfd

Observation 3a649fa7-a297-4640-82a1-1d40ae8fa644 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.177092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.177092Z digest=sha256:282aa732fa124c1dd4410322863580b9c09c3dc79fb9d8bd7deae89a6e0677c9

Observation a7ef88ad-7914-4c91-a6e4-0f7584eb693b · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.289278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.289278Z digest=sha256:c0e3719fb5aa25ee7d725d8a97964e969c2b584152604651da1f890be4ea5b59

Observation aa453ad4-0ad9-4b46-8019-78d046e0a629 · outbound

This paper cites Seed1.5-VL Technical Report.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Seed1.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.364515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.364515Z digest=sha256:9462de09707813476b8b4bdc8b374cf7b0624ac93d553794ffbe2e49d0b7e64b

Observation d0a20dea-c999-4157-9586-0d4a6ec543f5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.480638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.480638Z digest=sha256:c363d3ddb662a607e2eba4accdd76da0b1bcf34cb8cb1dd1a698214c9dab16d9

Observation 73d92c24-9872-40ad-a9a1-082c9a2b7a66 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.555893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.555893Z digest=sha256:078ca9bcc57c9faebfe022344c66e1e3ad7679d42e6c2b3a98f236a9413fb8fa

Observation 0ff9ea26-f393-4506-9922-25216861501e · outbound

This paper cites CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.651474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.651474Z digest=sha256:4147286a9fa508305d3e3cbda350d500bdbb3f0bb18c028581dbe741fca56787

Observation 57369070-914b-467a-9d5e-4de52cf965ce · outbound

This paper cites GPT-4o System Card.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.738113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.738113Z digest=sha256:977d2c0796a9a4d02bdb5a11c3bafbe37c307f92495457cb3a9ddbf5fb3880c3

Observation 80a94db2-df0c-498d-a628-6311cf99e159 · outbound

This paper cites OpenAI o1 System Card.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.813269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.813269Z digest=sha256:7a7c2ba32601a6a0c0e2d3243234d5161400be4414fd8da9e959d276f79d0cf7

Observation 35009f1b-316b-4e07-a07c-a50ae55b903a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.918115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.918115Z digest=sha256:fcacba61df548be27b4954b326208179d2b649bfeb1d84cdcd38503e5fbbc96a

Observation 9876737f-06d9-4d91-8603-06bd20891a76 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.024804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.024804Z digest=sha256:f35d7e5414d5290ba5982b2422a6e0fb3c50a02e74200fd3d55e53ba9890fe25

Observation 9a2ee76d-092a-4c18-a14f-371f0b13bfe2 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.115300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.115300Z digest=sha256:8c1d54f0b8b2711114e4ab9a46388c08fe7a78387a2a38ee3812450f2a98038b

Observation b759e0a0-aa7d-4e6b-bd61-b01086fb387b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:42.342580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:33.208153Z digest=sha256:7820004d92fbc45c526e84263d367585e7cfa0b61c52ddf0046db98093128568

Observation 70c288f9-a8eb-4fdc-b88e-b5b704dabdd5 · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.318843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.318843Z digest=sha256:96e89d23bf44e4c55a4f0879b3224dd6a4ad3cf721dcda71aa2a48ff69b760f1

Observation a3e949cd-5a64-4d0f-8c61-d0672c991c69 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.434821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.434821Z digest=sha256:1f53d6753a75816fdaff78605124407bab3ce311e9f2aacc045af91a614b8657

Observation 5f578a40-6dee-4156-92a5-2592e763fa5b · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.560796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.560796Z digest=sha256:8a0653c295236901b2577df06e97f8df14504b73d36542458b31296c72a488e2

Observation c6091be3-7e48-42a5-8e95-8f6c51e6a581 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.622232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.622232Z digest=sha256:48977846ce6a8598d60f3f153076c5166598c11ed7e4cdbe61ee9f1679bc0eee

Observation 9024d70c-e031-4c04-90c2-80499c3da584 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.722511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.722511Z digest=sha256:6019a14324ce353f47d7dac61d4a1bb39e9dd98d4b1c958c310bc467ea2d4e40

Observation 0c4133bd-d41a-42d7-8080-f873a69e81af · outbound

This paper cites Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.839250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.839250Z digest=sha256:ee21ce38a1feddcee0bbe2fe41250c3cecb94d389642a01d0c8b8852ecebec05

Observation d38dc04e-cd60-4b4d-9a3f-f55c27fbb8e2 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Measuring multimodal mathematical reasoning with math-vision dataset

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:42.037028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:33.917505Z digest=sha256:2ebae0e813a62bd4d60bed2801f1d67fd3e70ff12e2dbf6a21e347fe79da86a4

Observation 2edb0dae-c5a2-4436-b1b9-2868fed6b484 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.024686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.024686Z digest=sha256:5864245b986a1c28f0b0e01c7e72a749037b6653eb596ff0764a07ccf5dc478e

Observation 0793f6ff-3665-447c-b54b-3f85c43aa745 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Chain-of-thought prompting elicits reasoning in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.166022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.166022Z digest=sha256:fa691eb967ad4ddc83eef67cb9c2c22787c942501c9e2ae8e299e5e5d6275424

Observation 2642020f-54ce-45c5-b412-27e43624b0ba · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.276313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.276313Z digest=sha256:72746f45d876c78e6287a95728e97b7ffaf1915d92a824cd60f8173606a5f984

Observation 1165e788-0faf-4ce7-b2d3-b886ed25850a · outbound

This paper cites Valley2: Exploring Multimodal Models with Scalable Vision-Language Design.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.362773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.362773Z digest=sha256:5290de03ec2419ed8340cda05193c276253700308e48765ed848a35b7be4ba66

Observation b10a50c3-9317-44c1-8011-5dd93ad91905 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.464471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.464471Z digest=sha256:93157c4eab721ef9c643848369c230d7d9374c19084e7343311d20b7814ad161

Observation 64848bed-8331-43f6-a7f8-50025735980f · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.579200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.579200Z digest=sha256:1ad18a9a1f07cabe5923e3534c3cb40c3381460778060483e2b6102a81069c68

Observation f848019f-2edc-47ec-a455-f1c34aded2ed · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.668498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.668498Z digest=sha256:278d925be3cec42cbdd1f6c1f25c59538d881f52cfbea8d00971125626176fb3

Observation ab18994c-87b4-42b1-a5fe-574d3fcc4437 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.753888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.753888Z digest=sha256:74732203cfa31d74c2698aee8cd2e72b539807d66884bb70d8cb11fd7541c756

Observation 907799d7-17df-4531-9005-ab4ee4715382 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.797292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.797292Z digest=sha256:0e42419f2c6c2b159c62f4a21192646f301b641442833b2c78b558a1caa85d9f

Observation 3cbee560-f79c-4c3e-850a-a752a3295f19 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.864725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.864725Z digest=sha256:00cea1f1ee858820a5cb597ac9332edc49bd3b0038b267767e4c0e540f9fbed4

Observation 20999872-3733-4c93-b6a4-1859a808f77e · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.794677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:34.953774Z digest=sha256:01c880e40344be4d8fb135d9351330b62b28d5c722a3103ee6a50f4201a3070d

Observation 5f681d69-358f-4b4a-bfa6-f689294ea924 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.652440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.041699Z digest=sha256:729f41f6b659c110515a61031d58f4d758ad404d09173f498873ba94ed1d5ae1

Observation 98dcbe5b-8481-4df7-a5c6-8d29ffa9a212 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.516516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.156688Z digest=sha256:873c1c5741ab6a66bad5d7f3c5fbc67d5e418925f6ba76d0cf704efd927cba0b

Observation f842e7a6-5627-4b8f-b46b-b6a0001c49c9 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.348976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.228622Z digest=sha256:1925a90765f80d88f276c209e52c05b21bd577cabb4e86527ed5b92da7adf8c8

Observation 31826df4-9b47-4a54-976c-26c18e42758d · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.133657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.271391Z digest=sha256:e4d1c6a738938312fec9b60321201153c2e06436eb2f7de6ecf4efd38aadb77a

Observation aeb8bcfc-5433-4087-b4dc-e468a5675797 · outbound

This paper cites The screen visible in the image is small and not a touch screen.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The screen visible in the image is small and not a touch screen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.932629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.356982Z digest=sha256:cd51441ef9720bd898734e8bba211ff1693328b78554438e1bdc379cc3984fc5

Observation 3176c547-2f17-4fcb-bf8b-c8dfe69166b4 · outbound

This paper cites This option is speculative and cannot be confirmed from the image alone.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is speculative and cannot be confirmed from the image alone

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.667712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.443459Z digest=sha256:f3a2681bd46cfe0143c35f7657f238ca7dfc615c6f2b3829a3bd02cdaf21db47

Observation 85f92ca6-d4c5-419e-9a24-4d2b0a23fd56 · outbound

This paper cites This option is likely correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.401958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.521699Z digest=sha256:84404f97000c16fa5b1d4ef585b243850d9f30340cb8cad1ecf7a725ed2eb981

Observation 223a4946-abc7-4129-9639-9c05df960629 · outbound

This paper cites The image does not show any indication of this technology.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The image does not show any indication of this technology

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.194944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.582676Z digest=sha256:658ec8dab6365e401a2ea8bf7683efb89b512e0e01690fb5fc3303306b18eaa0

Observation 9ef69d52-90a2-4fa2-aef9-a1de3c5d9348 · outbound

This paper cites <vcues_2>The screen visible in the image is small and not a touch screen</vcues_2>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is small and not a touch screen</vcues_2>

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.964892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.669343Z digest=sha256:07dd26efd2a1bb28215e41936faaed1520a11c412bde81e0e32f8313f2217fa1

Observation ca1fb7ff-5e97-4513-9931-84e7dfb958d2 · outbound

This paper cites </vcues_3>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? </vcues_3>

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.782624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.723483Z digest=sha256:d73e6779adf56c593ceedfad02b526c766d43e63913d32a1e11ef9806c6743da

Observation 03bc4ae7-ed92-4abf-b66d-ce84abb40fbf · outbound

This paper cites This option is likely correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.570953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.756741Z digest=sha256:3506235cabd85341ccead31df44d2a736acdd93a27fafb24b6267e7911a6c8ae

Observation 93890703-e101-4190-9a39-d083b9789e5e · outbound

This paper cites <vcues_5>The image does not show any indication of this technology</vcues_5>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_5>The image does not show any indication of this technology</vcues_5>

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.399955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.811204Z digest=sha256:0e3a06cf07adfeac5cc88469903c802f368e5c5c1fc8851f2784094e11546484

Observation 86ef2ce4-a475-454c-a7bc-9a658381ebf6 · outbound

This paper cites <vcues_2>The screen visible in the image is very small and not a touch screen</vcues_2>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is very small and not a touch screen</vcues_2>

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.293227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.883434Z digest=sha256:4dbe9aee2cb6ff6b0f5f19aa1f91fa8fc0b564455e1855f682c2cacbbc1cf7e7

Observation a5ad5350-509d-4b95-a5f8-da23aa4ca87d · outbound

This paper cites Therefore, this option is incorrect.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Therefore, this option is incorrect

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.113386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:35.976409Z digest=sha256:b67e530e549f8741716e4845743fbdbdfc51005944cce61a8f6446a6dd633609

Observation 47154189-fc59-433a-811a-67dce248275b · outbound

This paper cites This option is correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is correct

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.971452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.012004Z digest=sha256:daab7da9d4172c9d4b75fb88dfe0d99eb9188d27907ab0890b6cbec34b035b4e

Observation c073b7ea-f187-44a0-936e-fbcc93b2c171 · outbound

This paper cites new_option.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? new_option

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.796243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.086645Z digest=sha256:2ec5a5f98daf3ce0d7a750e8979047bb9cae8bd468375aa8a214c2ea3271d313

Observation c9f65da7-5d69-429c-b231-0e8f3e28b1f1 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:38.666183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.188559Z digest=sha256:ff63d150137a5d71ed9a9151c2ce69b4b3cecf70079f7e70ce50aa258aed3420

Observation b7d41924-d39f-4128-ac98-7de16039e3d0 · outbound

This paper cites The presence of dirt could be from the environment it has traveled through, such as dust, debris, or even road salt in some areas.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The presence of dirt could be from the environment it has traveled through, such as dust, debris, or even road salt in some areas

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.554796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.276358Z digest=sha256:05d4e6f2924cc5d376c95961691356f71e73f27d66092ebd6ca6e92ebb8b425c

Observation b4f9552f-97e2-435e-adae-3957608376dd · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:38.421824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.364831Z digest=sha256:0a1a226bc4f4a2dce846e94a9021823e8641d10e17f0380aa87c5714ec0165c4

Observation 72575578-c4d6-4e54-93ac-36c4e9d7c5c3 · outbound

This paper cites The train was caught in a rainstorm: While rain can cause dirt to accumulate, the image does not show signs of recent rain, such as wet surfaces or water streaks.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The train was caught in a rainstorm: While rain can cause dirt to accumulate, the image does not show signs of recent rain, such as wet surfaces or water streaks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.195822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.449471Z digest=sha256:e43b32c1e67e679d2462c7b7396299d03afbf41a157befe3fd9173421f20d1c7

Observation 655077ef-3a7b-479a-b7ff-2410614e6255 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.953934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.536880Z digest=sha256:1ed17e3a98eb6ccf8e9b5ba0896afd4b4a6402228925e34a84e31920e283937d

Observation f339e64e-a48b-4402-ab33-5d9271709600 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.681408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.603461Z digest=sha256:6d2b210aa1a5c437ebd31d5b57e626e9de1e66442da916a91a662aba31c9ea25

Observation d9a69d1d-bb67-4fcd-9848-5b7df5cd8ab2 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.501653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.704682Z digest=sha256:974c2ff3bea670f7ecf41e8cd89f21c8a3137f429376f4afdeb44bea9e08fac0

Observation 885b32cd-a831-483a-838a-0b01f99498e6 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.393645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.792454Z digest=sha256:c6490be562e2730f40d5137c8e56f8997ca59835e66aa960943439cd96e0a7e2

Observation 958fcb89-84a5-4a7d-8d98-5019d2c34456 · outbound

This paper cites Given these observations, the most reasonable inference is that the setting is a small residential room, likely a bedroom in an apartment or a small house.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Given these observations, the most reasonable inference is that the setting is a small residential room, likely a bedroom in an apartment or a small house

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:37.321432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:07:36.844173Z digest=sha256:9e7eb314c8d58848228403b492297493540802965a59455d722657365945e0ca

Pith citing papers

Observation 9f88081b-4638-42e2-99c6-e46901bbed90 · inbound

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models cites this paper.

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T09:54:33.563266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:54:33.563266Z digest=sha256:299fa646f44667271309b2924320b40f54e9f3c96552a1139fed05b53deeaa4d