Pith. sign in

Paper Citation Record · LEDGER

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2607.26769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26769 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-30T21:38:08.719207Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ceea200d-26cc-4835-bebe-0a5c63819f39 · outbound

This paper cites Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.533422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.533422Z digest=sha256:6acd82076f5890ad112186a55d4fe48110ae70c3accee9305f4342b399c57cf8

Observation abc8c243-2b73-43c4-8c2d-9d94347d6c46 · outbound

This paper cites M3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? M3CoT: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.537659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.537659Z digest=sha256:581ac72de5a3e4b00d4d58c3dc31ed7336bde146e835cbbfd8a4871ca7cd5d80

Observation a9602b30-09bb-41bf-a683-ec629425ae12 · outbound

This paper cites Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.541273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.541273Z digest=sha256:59ebef4afc1991e6f8b58cb6e81260016e478a858ba8efd9b385d3ff0238b8c8

Observation 7d8bd2ea-cb7d-41ac-9776-b5e389f93bde · outbound

This paper cites Rbench-v: A primary assessment for visual reasoning models with multi-modal outputs, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Rbench-v: A primary assessment for visual reasoning models with multi-modal outputs, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.544761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.544761Z digest=sha256:76fc80f0302935114ec733ae7ce3c572d2866b1ccab1f299854f8dbab6616e20

Observation ed848c46-b96c-4906-a652-80b2149d984a · outbound

This paper cites Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.548069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.548069Z digest=sha256:cdb9494ee87a880224cd97c28d37466c1c94d0c46ee9d5b8b81e18eb3dc5f01e

Observation 19ec7674-8f35-45b8-9aed-e43058c53d4a · outbound

This paper cites TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.553205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.553205Z digest=sha256:491fdc774cf09f5bd71f4b7a466997c2544fbc7683608ffebb81fc69c48f6e77

Observation dae03d3e-953e-4dc4-ae66-d125d35e78e8 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.556986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.556986Z digest=sha256:7e77c9e54b8127ea344d909c1ce667d5c1392ff298bf7ab7f395972879df3748

Observation 45c97480-da36-43a6-8444-6bbe2ef7ac27 · outbound

This paper cites MME-CoT: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? MME-CoT: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.560274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.560274Z digest=sha256:f98243efd37d7bf921d84627bf9e7f4cce6dd844f960d7af8d9e82e3b46ce746

Observation 26edf29c-00db-4b42-9e5b-b9cba6e26a89 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Lawrence Zitnick, and Ross Girshick

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.563237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.563237Z digest=sha256:49523826f06b1ad9a43004a5b64e9ac4532d4bbfa364c840c18acd92ae615c0d

Observation 8b5cd433-5fda-4e54-8d95-dc95faaa9a25 · outbound

This paper cites DROID: A large-scale in-the-wild robot manipulation dataset, 2024.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? DROID: A large-scale in-the-wild robot manipulation dataset, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.566343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.566343Z digest=sha256:9471cd0d5af886f26cb51b9344e44a6882f9b19283ff77dc6cdd6d91dd42b80e

Observation d4ab71dd-167b-41ec-b3d5-80ef2b4fbedd · outbound

This paper cites Reliable thinking with images.arXiv preprint arXiv:2602.12916, 2026.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Reliable thinking with images.arXiv preprint arXiv:2602.12916, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.570083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.570083Z digest=sha256:aa9809ee39c749913ae91791ef7c22864c8abb4ffbf624f26e10be145c81ae24

Observation 67a9bba5-43ad-4e1f-8ccf-6cf4219b6cd4 · outbound

This paper cites S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.572980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.572980Z digest=sha256:bb9462afc423df35c4d87a05af04c5a6e08da8f83b136f5ccba8ec132402d93f

Observation 6fc47bb5-820d-40b2-8b64-954a4f1816eb · outbound

This paper cites Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.576140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.576140Z digest=sha256:dbb28810e2a39cfccf91fc418b21769c58b8241fdd836c5be6339032aa234fca

Observation 027a5dad-9f0d-431b-b61b-eabd304e06e6 · outbound

This paper cites TwiFF (think with future frames): A large-scale dataset for dynamic visual reasoning.arXivpreprintarXiv:2602.10675, 2026.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? TwiFF (think with future frames): A large-scale dataset for dynamic visual reasoning.arXivpreprintarXiv:2602.10675, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.578969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.578969Z digest=sha256:5cc25002840118664f3f46d7090ab8d52d95a95af38f4311d81ed4702732f9b8

Observation 550a36d5-e98f-4dc5-a847-930aaaf19aa1 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.arXivpreprint arXiv:2603.21754, 2026.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.arXivpreprint arXiv:2603.21754, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.581883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.581883Z digest=sha256:9d87dd68090facc08c73d266cb3c6ad347c808e465b3ef034bd69804dc377232

Observation 34a2c605-cbd1-491a-85eb-ce6a73f88ac0 · outbound

This paper cites On the faithfulness of visual thinking: Measurement and enhancement.arXivpreprintarXiv:2510.23482, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? On the faithfulness of visual thinking: Measurement and enhancement.arXivpreprintarXiv:2510.23482, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.584808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.584808Z digest=sha256:b6ee9d6795f2c4603a6596442c16ba3bc0388af28decb58d28358bd0db304608

Observation 64ea9b14-85c4-4574-bead-124ce19ac16d · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? MathVista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.587689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.587689Z digest=sha256:a7f59fe59a377d9b157461fcaeb048fb2b23487b6d8efef58d325a7eef4a9af8

Observation 81730bcf-1afd-4da9-9602-61da1aae1d10 · outbound

This paper cites Prism-bench: A benchmark of puzzle-based visual tasks with cot error detection, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Prism-bench: A benchmark of puzzle-based visual tasks with cot error detection, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.590463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.590463Z digest=sha256:d9d2cb0c8e519201776c8fb687a0e39da7ee560b2ddf2369541a0237222b9f72

Observation 25dadfcb-2293-408d-89ce-9cb6a266b4f7 · outbound

This paper cites V-thinker: Interactive thinking with images.arXiv preprint arXiv:2511.04460, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? V-thinker: Interactive thinking with images.arXiv preprint arXiv:2511.04460, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.593537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.593537Z digest=sha256:a91a3667f08b5abb9b0655626083f2d2070e5d2362ca43a0311388e41812e0e0

Observation 53241768-9313-4d39-8061-bc396a0728b6 · outbound

This paper cites Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprintarXiv:2510.14958, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprintarXiv:2510.14958, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.596386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.596386Z digest=sha256:d1ec371888fa46e0992b1c8f45dc8436bdac0a001656989a26176ecc22218690

Observation ad27c28d-a75e-4e4b-aa7e-abc964918073 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.599382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.599382Z digest=sha256:3fd4e520a674488b1a4d978433fd992194ac18511cce040b4b4e8ef446181389

Observation 272c90f6-1083-4a6b-9aab-ec9c628fac2f · outbound

This paper cites Le, and Denny Zhou.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Le, and Denny Zhou

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.602477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.602477Z digest=sha256:e1e9a6eac26c61fcaf5a10b0e8d09ae310b923defe506bdb66e8c987db48ec65

Observation 29784332-b18d-4693-8f56-eae7f94016a1 · outbound

This paper cites Vic-bench: Benchmarking visual-interleaved chain-of-thought capability in mllms with free-style intermediate state representations.arXivpreprint arXiv:2505.14404, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Vic-bench: Benchmarking visual-interleaved chain-of-thought capability in mllms with free-style intermediate state representations.arXivpreprint arXiv:2505.14404, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.605298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.605298Z digest=sha256:82cbd97b2b0941f90c858d646549a7c5ab33bd85cb1eb3d5bb65c3992a6aa549

Observation 0d11f08f-6272-47f8-8f0b-86454827cea6 · outbound

This paper cites How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.608340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.608340Z digest=sha256:c9b5b95ee4077d3d95e3aed6227325f2c4694dc982167289d6f85ae23b3781bf

Observation 9b906d49-9f1c-4e96-80b6-e2f12fa6e915 · outbound

This paper cites Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.611674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.611674Z digest=sha256:6b5e98b83aa18e73d135a2ea001c5f1d5b2a1e49336fc011a3ca5d2a38498c13

Observation e619ffa4-27fd-44aa-b403-6fdd003eaac2 · outbound

This paper cites MMMU: A massive multi-discipline multimodal understandingandreasoningbenchmarkforexpertAGI.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? MMMU: A massive multi-discipline multimodal understandingandreasoningbenchmarkforexpertAGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.614768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.614768Z digest=sha256:5c313794b71407572fbb88f3f157650a822133f0d7bf0611e77e5b28f24180df

Observation 7a47c920-6337-489b-9066-7eeca7643bad · outbound

This paper cites VLABench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks, 2024.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? VLABench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.617971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.617971Z digest=sha256:0022c785495dc1c8c8bf04228ec39269e553f2b781abcaf17ddaa3d129ebb11c

Observation 073a3bd0-2d98-4fee-9a1d-aa0a4375f4a4 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Multimodal Chain-of-Thought Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.620949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.620949Z digest=sha256:5bb4117d9e467df42d2207c1dada4a500d7a35e78087422819b0e43a2e4854d4

Observation 1c63436e-f1a4-4505-ad89-e5fbdde84415 · outbound

This paper cites Thinking with images as continuous actions: Numerical visual chain-of-thought.arXivpreprint arXiv:2602.23959, 2026.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Thinking with images as continuous actions: Numerical visual chain-of-thought.arXivpreprint arXiv:2602.23959, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.624102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.624102Z digest=sha256:78cb73138c85cfc495d45a11a8a03be58a950ef48dd8c32e8ea3a1e4b8c8d3cf

Observation 4d13fa59-de63-489e-8463-8ecb04d257e9 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.627112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.627112Z digest=sha256:50f2a92b9987a6816d8106951c8958f309ca1f8492944de6bac03d4b49fd835f

Observation 3af436eb-a351-4837-8823-d4bd82cac763 · outbound

This paper cites Whenvisualizingisthefirststeptoreasoning: Mira, a benchmark for visual chain-of-thought.arXiv preprintarXiv:2511.02779, 2025.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Whenvisualizingisthefirststeptoreasoning: Mira, a benchmark for visual chain-of-thought.arXiv preprintarXiv:2511.02779, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.630224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.630224Z digest=sha256:2c0a0905a3c2dcbe89bfb270819732e51ac6d51d1f395375450adb903ea0c989

Observation fd613f3f-87f9-4367-8f0e-00eee46f6ad6 · outbound

This paper cites What, whether and how? unveiling process reward models for thinking with images reasoning.arXiv preprint arXiv:2602.08346, 2026.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? What, whether and how? unveiling process reward models for thinking with images reasoning.arXiv preprint arXiv:2602.08346, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.633014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.633014Z digest=sha256:6850f294ab96d6f9b4770a4492dc014601b7f69fcc4d286484d001d8a5f40c95

Observation ea7a89c3-2523-493c-9dfe-df3b26c7a07f · outbound

This paper cites Never collapse multiple ideas.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Never collapse multiple ideas

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.636423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.636423Z digest=sha256:a4451c60ce238dbedca2b0559fcb522f571894da483e7d780f17045f834c451c

Observation f3f93aa1-d8f1-41ee-8153-9961409ea493 · outbound

This paper cites stress concentration at joint.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? stress concentration at joint

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.639335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.639335Z digest=sha256:47ed5558653737f1640dc567391f7948a8d070a0385bf692b1bff8d853fcc204

Observation ef6d4d3a-c20a-41f5-9c25-f9c99f8e3a68 · outbound

This paper cites Friction.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Friction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.642487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.642487Z digest=sha256:e33f2eecfdf2269697781af2ba1f0dc74df745b906424c22cf872d4ff333524c

Observation 4a5884a5-7bb6-4ae4-974b-f918cef2afab · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.645604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.645604Z digest=sha256:43294d155d0b1ae294ea96f6718fc043b0650ed70332bfafea70dc9adbe076b3

Observation 08dc0956-b00d-4d40-ad2d-eaf6172f2eb1 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.648644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.648644Z digest=sha256:50f3a89fb317431fd00112929fc549ef11f23844e53525fc929588938e00562a

Observation 73371f61-5bda-437e-8158-355095454732 · outbound

This paper cites The ‘action‘ array MUST contain exactly one item.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? The ‘action‘ array MUST contain exactly one item

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.651535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.651535Z digest=sha256:c2f07392d2c4cdb3cd09bbc26a58a15de560e9ede914e64c030110e50dffb002

Observation c9fb6c98-3648-44e3-9397-bed18c948cfd · outbound

This paper cites shape":.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? shape":

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.654769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.654769Z digest=sha256:59a0e311baccfec8cb322f3b02a212f688470615c89f0e0bdc9ed6c434277a53

Observation 0163d187-5a0c-4a07-8c62-31d7b7c92983 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.657639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.657639Z digest=sha256:3be5f0f423c43a2cd089c918cc8336ff03a1f878acdf978564138694e2dcc224

Observation eb0d507e-921a-42ac-b5d7-31b20d7a9d1f · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.660614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.660614Z digest=sha256:e63eb8065e7f232d89ed887b0c330323acbb64b49a7649329c8c203257a776a5

Observation cef1cb31-8a30-44f4-91c4-f4a841973bda · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.663340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.663340Z digest=sha256:028fb30036d2320cfad952307c95478b030efa0a65d0727cd81694cee887cfd4

Observation f09f6f84-41d0-4a1b-9f9e-aa5c50f8b0bb · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.666191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.666191Z digest=sha256:f03029dc89d9d3308a04c5185353d65a31f86a0e8fe2f10a47e6daf45aeda2ea

Observation c81ec226-fa95-4910-bd17-eb783fbd6bcf · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.669064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.669064Z digest=sha256:26b5ac48a4862f6d18e0b5e5892047d03836f68bc11f093a73a8bd599f7780cc

Observation dfabcd36-3d59-4f97-bbd4-0e7f9aa34685 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.671819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.671819Z digest=sha256:08940570a39fe2b7cb76ce66e060c41a0b6cafdee95b9372a24ea2e31dfb940e

Observation ed798cd9-3c39-475c-9962-1063a0f46ccb · outbound

This paper cites step": N,.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? step": N,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.674624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.674624Z digest=sha256:c979317495a55acf4db93bfdfe30e04f16c7e37fd2a9f2a32edcda14c22f50bd

Observation b3abdb69-2b02-42de-b2ab-62595c09227b · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.677518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.677518Z digest=sha256:d7f8c3340cf4e7ebf227d16b49d2bea0ac010a6b9e835ce95a20cee89fea0eb7

Observation d865021d-3a72-4605-b188-b4920efaa418 · outbound

This paper cites question.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? question

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.680309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.680309Z digest=sha256:084791fd8419b7232913608c0204b21e36d5b5a3363943dc509d3ad4d4f222c1

Observation 48ff0102-c181-4c95-81ca-5abdfb27fc93 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.683015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.683015Z digest=sha256:32b959c80bb3069f3963d8305d18ab87f59aae37971d9476768218a90c2f79c5

Observation 66922652-a1bd-4d0d-83ea-e71211668e61 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.685679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.685679Z digest=sha256:31c015bc828a0b3efbee7c6b7e19a7363a408af976d69b5083a1641e4afc2f7c

Observation 4a31e8a4-eec4-4e35-93d2-87159b70832f · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.688692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.688692Z digest=sha256:aa1bbb9a17a01f279737833c124d23480a72a8de70a595bc2d37b2d388b0ae13

Observation 0171318b-b2a9-4722-847a-39185a9ca77a · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.691471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.691471Z digest=sha256:521a552c98e315a9f6f438803160816a8e23819e9b9587488ce77cb517d50176

Observation 19856ffa-3bbf-4bbc-926f-01f3aa63e7d2 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.694204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.694204Z digest=sha256:701fc9c4f0e10ad1012fb16294fa595280b9ad7411acc0fc6cd34485ef0fae77

Observation 7595bf09-4d9a-4ca9-aaaa-80aa95e8860d · outbound

This paper cites correct": true or false,.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? correct": true or false,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.696981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.696981Z digest=sha256:3f7087a9dd080b46720a9c3d52d13383873ae84bbf37a5a22281e8b8e15aed12

Observation 4f89d73d-2acd-44b0-b3ef-64eb57072401 · outbound

This paper cites If there is no effective visual operation, use null.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? If there is no effective visual operation, use null

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.699621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.699621Z digest=sha256:a3c057fa8321ac7a217867d4a32d782d0fe0d0c1d0a2f37b70bc51a950b17dd6

Observation a0e5dae0-b272-4eb8-8b2f-63125b8636c1 · outbound

This paper cites - 1: The visual action directly targets task-relevant evidence and is useful for solving the question.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? - 1: The visual action directly targets task-relevant evidence and is useful for solving the question

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.702230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.702230Z digest=sha256:caaa3954d47618b07b2a5b3e940b00c50bf438a3d473573d7f19e00557baf808

Observation 8a1a0813-d11b-4d1d-8668-48540d221c93 · outbound

This paper cites - 1: The rendered visual state faithfully executes the intended action.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? - 1: The rendered visual state faithfully executes the intended action

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.705049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.705049Z digest=sha256:a52eb27a4d80020e56db0a76a5fe1a5bf9388381594e1b218fbd699857ae70b0

Observation f02bd525-1d62-4e88-b55a-e4607d86d433 · outbound

This paper cites key_step_id.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? key_step_id

Reference 58

Resolution
malformed identifier
no resolver link, observed 2026-07-30T21:38:08.707859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.707859Z digest=sha256:e4deaf0fa71296daaa105859fd30411f4df24f51516231a98477fb8bd9365817

Observation 567dba69-5ae0-4abb-96f5-cd620d8946ba · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.711161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.711161Z digest=sha256:b7a1b568ebd39e24effc17832545a8b3713f27a7fa3cf156cc99a0f499e0dafd

Observation 033c161a-91ed-4d0b-afdd-2c1b0f6104b8 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.713952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.713952Z digest=sha256:0b00a38f9cea06805dcecc54c49a408accf4b6184aacea636d0fcb9a56b52f3c

Observation 240a4b35-a912-49af-bca8-25fa75d08216 · outbound

This paper cites an unresolved cited work.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.716494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.716494Z digest=sha256:7df792bdbbad1ce03a4dad606d1aa5272205a0409e1d6031bf5e7df847612469

Observation 463a13b3-6b47-4ae8-bf6b-5bdd88df3907 · outbound

This paper cites This construction covers high-action/failed-answer cases, rendering failures, partial renders, and weak action selection across all audited models and environments.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? This construction covers high-action/failed-answer cases, rendering failures, partial renders, and weak action selection across all audited models and environments

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.719207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.719207Z digest=sha256:03adefe9afcf2b05ed0301ae26504ba387fa771721a0b34d9697353eb6640e57

Pith citing papers

No inbound Pith citation observations are available.