Pith. sign in

Paper Citation Record · LEDGER

Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2406.09403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09403 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:37.543242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:06.763408Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8900b923-fec1-4134-bf1e-df2e0167bfa0 · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.850060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:526e6d7597c2e88843aa48dff6cf2e34c975271e3a52c02391943b368c872ced

Observation 3328d2d4-ff2d-417f-b578-9da61e9388d4 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:09:34.893504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:60cf0efaf9ed69d48a809e480a8e137d9b94c91b9f8aa2503adfbb26279d7d19

Observation 8306d360-a10e-4e24-b800-20a7536ac8ee · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.725866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:de5ca04c1f9bac20c4933c6ba76ea132eb5a359436a816fedb96e55d6094f26e

Observation fb008005-3db0-4d4a-9fcb-a8ab1f0688f3 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.413368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:967b2ab9fce4209b98dc6e7e375a01ce4f3dbb064f4a5aab2f0b180c78bd2edc

Observation 09fb45ae-c0ba-4d5d-98ed-9ffe2f8c4293 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.253558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:b580a559c2a033b757f92cb3ebd0979e64880773f6f38c8a80ded9bdddfe7159

Observation 799ab86b-46fc-4aa7-a343-b09fb88d3a34 · inbound

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization cites this paper.

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:37.543242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:37.543242Z digest=sha256:d678cf79b5ec25374dd70dc4a360795e1cd0f5f2deeb20a8654497fa5b5b80f9

Observation 449c37bd-508f-4e8b-aefc-cdf5f7300e83 · inbound

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning cites this paper.

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:17.684374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:17.684374Z digest=sha256:b5edc2672cd1877125d58b9ec7e54f42187f272f6472bff12ba199694f05e930

Observation 9ccf031d-5efd-4d4f-909b-35505b08caed · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.958914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.958914Z digest=sha256:7caf9ecbef853b15240b6bbf9afd62e4bc9ff1d13e642f8324e426df6d78a3ab

Observation 2fa008a5-8011-47d7-abef-8f08ead072a1 · inbound

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection cites this paper.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.478765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.478765Z digest=sha256:d09b60702ec4cb670f4c92a91c1355ac826360e46e48c1bd63d50c9ea1d702f8

Observation 6bc41505-3b21-460d-935c-54c92fd5fe5f · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:57.574309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:57.574309Z digest=sha256:50b9640ae1a230a3c179cb3e2211d6849b227f485facc66a94f1ffc89240e834

Observation a7706e48-a2e3-4ff6-bf07-642cdc4169d1 · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.817955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.817955Z digest=sha256:c0c67902ecd3a46f745020e239ba8b993f2a6251f1943d8e470a680dcc90e225

Observation 8c8afe21-35f2-4125-833e-060f72d877c9 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:51.959972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:0b8c337d0d1250829d7f24a658f8547934aff7a3b5777d3ab3986f83df8e049c

Observation ec006fd1-3b49-4d64-9d77-f134610fc033 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:49.091946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:49.091946Z digest=sha256:d8152855f03a394eac8e5f3908a0ee2b3d7b560cb611d13dd36678c18569876b

Observation 77bd8e6a-ba53-422d-8865-9b17c424115f · inbound

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems? cites this paper.

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems? Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:08:23.824053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:08:23.824053Z digest=sha256:0541aaa58bec7f646119710bf553858cf3d60689de719355f1697783b0dffae7

Observation e642d910-ee4e-4303-a84f-58f9b6d030a5 · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:47:15.175790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:d042490a64e5991ae06e8d2264a55f69d97fb94dfae3e0d0b9951705bcaa1e09

Observation 7ead8605-e0da-4ea8-a98c-dd2ddaf30320 · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.017947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:301a00f2a5b93684ef822b2a817271f4d9ba5b235c0d2bdb08e931ce2ff5b1f4

Observation b50ced97-7fe9-4bc9-afba-f046b0b71628 · inbound

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning cites this paper.

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:01.846583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:01.846583Z digest=sha256:1206ff7988a7e95f373c7613ebe200bbf7801312c8647df3a3e297ffc43f52b8

Observation 474ce44a-9320-46cf-baa5-8e0de07e3743 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.739981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.739981Z digest=sha256:dda08379b3d0ba3b04d3dfff43f2ce2d729a5d503bcaa9105c14bc9f2e05c548

Observation 66db5aa3-1ada-4b95-aeaa-f8a72ae5e371 · inbound

Beyond the Textual: Generating Coherent Visual Options for MCQs cites this paper.

Beyond the Textual: Generating Coherent Visual Options for MCQs Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:37.650215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:37.650215Z digest=sha256:8379e1f9977a19e49434c79ea54a01d4f2e8782b49c091b3e738db3a1cdea53b

Observation 024b8bf7-c1c0-4df1-b37b-ff80ff9e48b8 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.939778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.939778Z digest=sha256:fd58c871a472b8761a61df9e4ba4dc63965b5cd39fbde38976bcfc86169e7349

Observation 02c0cb71-5d05-4c7b-8462-0db2ef9e84f6 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.394788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:df6707cc4c285cc1b8abdbd76705ffc7b58bc1a02f91aba7f61c9a58f0a2acca

Observation 2626c9e7-bb7c-47fd-aba1-09bb216835f0 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.632101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:fd82a98dda2a8f091576b0e8fb7d5951076261796a4fafc5a5f655b483647f37

Observation 5318a328-3d3a-444a-8a0a-cdeebb3c6250 · inbound

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space cites this paper.

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:03:38.855828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:02:28.588225Z digest=sha256:15687863c1f55fff07db62755f960810da9f07c6023b8a4e6546ce727fd149f9

Observation 3a0f45fb-4e90-421f-8deb-b9013e7a7411 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:20.446939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:505ecb747314a52f47dbefd410a3ac185e6e4c7ca7e6da865838a7983afd6660

Observation c21318ac-c93c-4cd6-9883-fe74d5dd2cdb · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:48.407273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:963f84fc7ff5e30a360c00f59ff9b4d265c9e72b8ecf0025b57972504dc2a41b

Observation 7a20732f-bf09-41ac-806d-f3f8473ebc08 · inbound

Visual Reasoning through Tool-supervised Reinforcement Learning cites this paper.

Visual Reasoning through Tool-supervised Reinforcement Learning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:02.948682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:05:21.688216Z digest=sha256:d9657c7a8a14815db70add31263245d1788f25e93a1ba7f4b2bdbbeda2757c74

Observation 51e95d26-ad37-4a67-9494-451d57092c35 · inbound

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning cites this paper.

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:37.451329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T17:35:28.050906Z digest=sha256:d54e90990f5dc1b0dd835ae1eae52c912450b575492a26d54ecaec0e19d69e81

Observation e3be280d-8b2c-4293-bc23-66db75dd0d7a · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:28.416568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:58:21.232058Z digest=sha256:dfc3e9c53f631a69ea37bb40d24e24c444235cf383d7be5ed6b4700059a5af62

Observation bf9db31a-edea-4c24-8899-54588db98cf3 · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:50.847474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:23:14.695568Z digest=sha256:65b3001e4dc048ae0f5b0594e645df42a6834fdf81d2b26e1f127f5119dcaa76

Observation 81725268-a564-4a47-90d1-ff28399f9dd8 · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:34:05.257689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:33:24.285225Z digest=sha256:16945d34d2381302c709b0f9ef3ea5cab9b785d12763cf7b3c4665afdcf3d468

Observation e4ab9060-1d9c-474a-b4d9-70a023ddc4b1 · inbound

From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing cites this paper.

From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:24:56.099669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T03:20:26.350336Z digest=sha256:9cb144a3acb9d82ef7495cfe71ea57f23ffd8fa7ccb67d5dc57cd5794f18042a

Observation 1efc79e9-3578-42f3-bcba-246678e0a8fc · inbound

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction cites this paper.

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.507860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T06:05:45.046156Z digest=sha256:aea26fe23f38f8c874575ad6391f97b884c648a6ee6f108cda0336a022e70ff6

Observation 7fe46e34-862a-45e3-bb02-7d6a0ca2ebb3 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.584621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:9f1c88b1b7c4d02f503fb7e93d52d8820691cc2f3976498b38a5834f3014e83e

Observation 28089a4d-6922-42b1-ac12-c63d3975a84a · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.933236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:64c804e89a7588275ecb87e1dc996a62bf82cf0a95fb2dac4a8fde4cfdfc7716

Observation 31404f1e-d046-4a1d-8bd2-2854a8fc1228 · inbound

VESTA: Visual Exploration with Statistical Tool Agents cites this paper.

VESTA: Visual Exploration with Statistical Tool Agents Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.561642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:58:11.339217Z digest=sha256:051600de60c9836391dfe402c05804517a4d5d5171ff168d7d5c83229e4d52bc

Observation dada992f-503d-4648-9141-34cf93fcf1b3 · inbound

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers cites this paper.

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.200334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T18:59:51.362554Z digest=sha256:aa3540dd63f021644b541d0bac60261090eb319a7e382af7138601226361763f

Observation 8b94058f-7248-4cf1-b435-94e361fc733c · inbound

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models cites this paper.

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:29.328950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:56:05.047862Z digest=sha256:151538fa7c715320803ccd7f76f2616f0422248856f11dcb76c4483e3ac185a0

Observation 4685a236-406a-498e-bc56-925646062584 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.368872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:1145966f89a4fdd6a4ec328e0f3876a1bec6e93c2141fba5b45694db7740d2ec

Observation 7fe04b53-5aa9-47a0-8d4c-e05707d8f480 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.515808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:62dac8c3860dd321bb1b1650d0239015fbe5ec2095be5b2ea3ac4d3983b94ba8

Observation 4826be67-3467-438d-aac2-49412dd2fa53 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.586402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:dfa025ddaf579d9e8d7cdc0a547a903a1bb8ebfda0ee19bdf652c83d2f408a2e

Observation 726b1bbe-1a68-46bd-9c87-4d3bb4e28aa4 · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.253830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:a80c9a4b6589aead2fcb3469decaab567e124df792eb8aa7b077e354238209cc

Observation 2297cd65-98aa-4276-93c3-fbcc8832c40c · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.267994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:2bb0af95058a2a7751c421406f9c9ebadd0f83f5a5a14b5935486bab9ac5e14a

Observation 2f75a577-3ce0-4721-b756-35f7ad4e3911 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:06.765443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:73fdbeefcddafa5010600eb53097bbf50ac86b58ab6122ff1ae9075428ddf42e

Observation b548de94-82a2-40a3-a429-b8a17399cc04 · inbound

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning cites this paper.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T07:59:56.441098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T07:59:56.441098Z digest=sha256:fba571c6d22600e910fb54647503f6a59e382f691e42ca66081e674d78ee5598

Observation dae03d3e-953e-4dc4-ae66-d125d35e78e8 · inbound

See2Think: Do Multimodal Models Really Use Intermediate Visual States? cites this paper.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.556986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.556986Z digest=sha256:99336cb770bafb5572ba17ad61ced204d0d39d0f0782e54d7227065ed7bce650