Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T21:38:08.719207Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2607.26769.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T21:38:08.719207Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ceea200d-26cc-4835-bebe-0a5c63819f39 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc8c243-2b73-43c4-8c2d-9d94347d6c46 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? M3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9602b30-09bb-41bf-a683-ec629425ae12 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8bd2ea-cb7d-41ac-9776-b5e389f93bde · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Rbench-v: A primary assessment for visual reasoning models with multi-modal outputs, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed848c46-b96c-4906-a652-80b2149d984a · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ec7674-8f35-45b8-9aed-e43058c53d4a · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae03d3e-953e-4dc4-ae66-d125d35e78e8 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c97480-da36-43a6-8444-6bbe2ef7ac27 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? MME-CoT: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26edf29c-00db-4b42-9e5b-b9cba6e26a89 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Lawrence Zitnick, and Ross Girshick
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5cd433-5fda-4e54-8d95-dc95faaa9a25 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? DROID: A large-scale in-the-wild robot manipulation dataset, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ab71dd-167b-41ec-b3d5-80ef2b4fbedd · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Reliable thinking with images.arXiv preprint arXiv:2602.12916, 2026
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a9bba5-43ad-4e1f-8ccf-6cf4219b6cd4 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fc47bb5-820d-40b2-8b64-954a4f1816eb · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027a5dad-9f0d-431b-b61b-eabd304e06e6 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? TwiFF (think with future frames): A large-scale dataset for dynamic visual reasoning.arXivpreprintarXiv:2602.10675, 2026
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550a36d5-e98f-4dc5-a847-930aaaf19aa1 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.arXivpreprint arXiv:2603.21754, 2026
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a2c605-cbd1-491a-85eb-ce6a73f88ac0 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? On the faithfulness of visual thinking: Measurement and enhancement.arXivpreprintarXiv:2510.23482, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ea9b14-85c4-4574-bead-124ce19ac16d · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? MathVista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81730bcf-1afd-4da9-9602-61da1aae1d10 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Prism-bench: A benchmark of puzzle-based visual tasks with cot error detection, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25dadfcb-2293-408d-89ce-9cb6a266b4f7 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? V-thinker: Interactive thinking with images.arXiv preprint arXiv:2511.04460, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53241768-9313-4d39-8061-bc396a0728b6 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprintarXiv:2510.14958, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad27c28d-a75e-4e4b-aa7e-abc964918073 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272c90f6-1083-4a6b-9aab-ec9c628fac2f · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Le, and Denny Zhou
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29784332-b18d-4693-8f56-eae7f94016a1 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Vic-bench: Benchmarking visual-interleaved chain-of-thought capability in mllms with free-style intermediate state representations.arXivpreprint arXiv:2505.14404, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d11f08f-6272-47f8-8f0b-86454827cea6 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b906d49-9f1c-4e96-80b6-e2f12fa6e915 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e619ffa4-27fd-44aa-b403-6fdd003eaac2 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? MMMU: A massive multi-discipline multimodal understandingandreasoningbenchmarkforexpertAGI
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a47c920-6337-489b-9066-7eeca7643bad · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? VLABench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks, 2024
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073a3bd0-2d98-4fee-9a1d-aa0a4375f4a4 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Multimodal Chain-of-Thought Reasoning in Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c63436e-f1a4-4505-ad89-e5fbdde84415 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Thinking with images as continuous actions: Numerical visual chain-of-thought.arXivpreprint arXiv:2602.23959, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d13fa59-de63-489e-8463-8ecb04d257e9 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af436eb-a351-4837-8823-d4bd82cac763 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Whenvisualizingisthefirststeptoreasoning: Mira, a benchmark for visual chain-of-thought.arXiv preprintarXiv:2511.02779, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd613f3f-87f9-4367-8f0e-00eee46f6ad6 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? What, whether and how? unveiling process reward models for thinking with images reasoning.arXiv preprint arXiv:2602.08346, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea7a89c3-2523-493c-9dfe-df3b26c7a07f · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Never collapse multiple ideas
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f93aa1-d8f1-41ee-8153-9961409ea493 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? stress concentration at joint
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6d4d3a-c20a-41f5-9c25-f9c99f8e3a68 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Friction
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a5884a5-7bb6-4ae4-974b-f918cef2afab · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08dc0956-b00d-4d40-ad2d-eaf6172f2eb1 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73371f61-5bda-437e-8158-355095454732 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? The ‘action‘ array MUST contain exactly one item
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fb6c98-3648-44e3-9397-bed18c948cfd · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? shape":
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0163d187-5a0c-4a07-8c62-31d7b7c92983 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb0d507e-921a-42ac-b5d7-31b20d7a9d1f · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef1cb31-8a30-44f4-91c4-f4a841973bda · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09f6f84-41d0-4a1b-9f9e-aa5c50f8b0bb · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81ec226-fa95-4910-bd17-eb783fbd6bcf · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfabcd36-3d59-4f97-bbd4-0e7f9aa34685 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed798cd9-3c39-475c-9962-1063a0f46ccb · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? step": N,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3abdb69-2b02-42de-b2ab-62595c09227b · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d865021d-3a72-4605-b188-b4920efaa418 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? question
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ff0102-c181-4c95-81ca-5abdfb27fc93 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66922652-a1bd-4d0d-83ea-e71211668e61 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a31e8a4-eec4-4e35-93d2-87159b70832f · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0171318b-b2a9-4722-847a-39185a9ca77a · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19856ffa-3bbf-4bbc-926f-01f3aa63e7d2 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7595bf09-4d9a-4ca9-aaaa-80aa95e8860d · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? correct": true or false,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f89d73d-2acd-44b0-b3ef-64eb57072401 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? If there is no effective visual operation, use null
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e5dae0-b272-4eb8-8b2f-63125b8636c1 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? - 1: The visual action directly targets task-relevant evidence and is useful for solving the question
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1a0813-d11b-4d1d-8668-48540d221c93 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? - 1: The rendered visual state faithfully executes the intended action
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02bd525-1d62-4e88-b55a-e4607d86d433 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? key_step_id
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567dba69-5ae0-4abb-96f5-cd620d8946ba · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033c161a-91ed-4d0b-afdd-2c1b0f6104b8 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240a4b35-a912-49af-bca8-25fa75d08216 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463a13b3-6b47-4ae8-bf6b-5bdd88df3907 · outbound
See2Think: Do Multimodal Models Really Use Intermediate Visual States? This construction covers high-action/failed-answer cases, rendering failures, partial renders, and weak action selection across all audited models and environments
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.