Pith. sign in

Paper Citation Record · LEDGER

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

As of 13 August 2026, this Paper Citation Record lists 100 of 129 outbound references and 0 inbound Pith citation observations for arXiv:2608.09682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09682 v1

Coverage vector

measured 100 of 129 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:53:14.503030Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 129 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21d126e9-bc08-49be-ac7a-28c7872f3451 · outbound

This paper cites Latent reasoning with supervised thinking states.arXiv preprint arXiv:2602.08332, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Latent reasoning with supervised thinking states.arXiv preprint arXiv:2602.08332, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.922601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.922601Z digest=sha256:07c0c92514ecb1a4f9b1d36944199610d05e8b0f254ce51aa5b057fd5211b976

Observation 3b40b537-9d4c-4dae-87c7-72e0cbaead1d · outbound

This paper cites Acloserlookatbiasandchain-of-thought faithfulness of large (vision) language models.arXiv preprint arXiv:2505.23945, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Acloserlookatbiasandchain-of-thought faithfulness of large (vision) language models.arXiv preprint arXiv:2505.23945, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.930040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.930040Z digest=sha256:7646ac8dae6e321d439a75cde7963d49254e4cb46b6aa57d277ef15aae43971f

Observation 54c530f6-4381-45b3-8282-d55da1520d61 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.935715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.935715Z digest=sha256:f77b251a9c001adb81ae12589c54fb5b0703d1f3afd65713b6a294149a6b7ad9

Observation 4127ec8f-cb29-49ef-bf11-6a0e20ae4c03 · outbound

This paper cites MMStar: An evaluator-centric benchmark for multi-modal large language models, 2024.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning MMStar: An evaluator-centric benchmark for multi-modal large language models, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.941754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.941754Z digest=sha256:c4229a26ed4e8303e3723153779002e1a1307653750932ed9a6ed18dbae6017f

Observation e3a06ba7-1a84-4d28-9062-55bd6db2c470 · outbound

This paper cites Perception before reasoning: Two-stage reinforcement learning for visual reasoning.arXiv preprint arXiv:2509.13031, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Perception before reasoning: Two-stage reinforcement learning for visual reasoning.arXiv preprint arXiv:2509.13031, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.947142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.947142Z digest=sha256:24330bc4133cea26066111e5bbc6a42cc60efcd995c439fb1a4eaf18514c21ac

Observation c1d6048d-a793-41d2-8a99-295b545a062e · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.952698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.952698Z digest=sha256:c89111f880e1268f16fba1232c4e83d9f570fb1d74666636147b19b1846e9533

Observation 527f98d2-c979-4acc-b913-fbf7eec197af · outbound

This paper cites v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.959442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.959442Z digest=sha256:47b5cd90b0e121a0b17b648e9849b37918ce10d4f64038f7485d0b6b9ae9094f

Observation 4b2e1502-e09d-4db2-af0d-9f9e3250c7e9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.964796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.964796Z digest=sha256:0c323571ad217ca6905ec8b64f4f0b1a014fd05e07683a033e353ab8e20ede1f

Observation 509a382a-9523-41ba-98f9-4c79aa25a0ef · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.970205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.970205Z digest=sha256:26c0d37eec1c0eb13b8edfc91be5be7d825d1870d9ab379ecad7f62d3783d230

Observation 93107529-ff1e-4d89-8ff6-a0d9f5579f93 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.978385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.978385Z digest=sha256:3b4aa804eb1e992390aa2017fbc0a7e637d33d519c223646401aec1b7c9468bd

Observation 65d3ed44-1f15-4843-b40c-8316c5ac469c · outbound

This paper cites Revisiting the necessity of lengthy chain-of-thought in vision-centric reasoning.arXiv preprint arXiv:2511.22586, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Revisiting the necessity of lengthy chain-of-thought in vision-centric reasoning.arXiv preprint arXiv:2511.22586, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.984603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.984603Z digest=sha256:e9f34f7c5c6e61ad5b7f519183119f35845897a8565a38c845aabdab283a8843

Observation be44b9db-1f68-4022-8868-c22a5ce9043b · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning VLMEvalKit: An open-source toolkit for evaluating large multi-modality models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.990131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.990131Z digest=sha256:d2aad886afa07533e0be11f507cafd9ad28bcee885e6bd8a8a9a228c2cd37c6f

Observation 7931aa8e-29c5-4a9f-9949-92b81a812407 · outbound

This paper cites GRIT: Teaching MLLMs to Think with Images.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning GRIT: Teaching MLLMs to Think with Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.996266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.996266Z digest=sha256:39625d23789a4e27ae1e9fb19fc025f8326bc134f00c8f0f91e7a351c21e0ae2

Observation 9056dd3f-3141-49d4-8937-2bf9ac597373 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.002043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.002043Z digest=sha256:4689f2e562fe54f8fb08849aca7178e2a2aabdb6bd2dd15ea96449ec62d409be

Observation 195d31a5-11c3-492e-af18-b14ebf648ec6 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.008737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.008737Z digest=sha256:637ae3f44e924c586239a9426616aba7b577b48005a554c124a9f01793df9d1a

Observation 64d7c8d5-a9f5-446f-ab58-b427e8915449 · outbound

This paper cites Thinking with deltas: Incen- tivizingreinforcementlearningviadifferentialvisualreasoningpolicy.arXivpreprintarXiv:2601.06801, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thinking with deltas: Incen- tivizingreinforcementlearningviadifferentialvisualreasoningpolicy.arXivpreprintarXiv:2601.06801, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.016201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.016201Z digest=sha256:e985cfeefa237dd3b959cf5799cc3bf7ad545b10612c78d44c6a7e71447660ad

Observation 068e1c91-7dab-46d4-be3a-3a5502407ccd · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.021564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.021564Z digest=sha256:5a2fb619eee29c6197456e720ed207087a6548f81141752f326a138b94a0c223

Observation a332fe5b-83bd-4f84-a92e-1559c14595f7 · outbound

This paper cites GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.026958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.026958Z digest=sha256:0e313a8dae7c03a34061686ec24d91d4d5f0e217f6fd8ef92dcb3cf7c0f377bf

Observation 0456dc97-cbd6-472e-9331-b839f796af49 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Visual programming: Compositional visual reasoning without training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.032344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.032344Z digest=sha256:ed45a16a49d0eda9e7f5c436ca5e231c3fc29c72f10c7266a55148faef80b075

Observation e15c4967-d4e1-4555-9a7d-b39829d8404d · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.037446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.037446Z digest=sha256:b3b894ee8f4573c829570c9e227f852435836e1deeb4045c4813d8c26a355b06

Observation 661bcc98-115d-42d1-af70-47b2e98f38f0 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning DeepEyesV2: Toward Agentic Multimodal Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.046256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.046256Z digest=sha256:26ee64b14db1246b7f5fc0d567db89ad3f22d7ab0a30acdd0116a1f8129f88b3

Observation caeb8699-53c6-4525-ae64-b3f61ae401e4 · outbound

This paper cites Hollon, and Bryan Wang.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Hollon, and Bryan Wang

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.052496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.052496Z digest=sha256:915f20f4cd65158444bd7800864c6bbdae13b8cf5d2261429ce3f6a2b42a09cc

Observation 24edadf1-b96c-40e0-b450-1d1055e80bc6 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.059008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.059008Z digest=sha256:6f83b41f902611fe51003710ebd408dfd635536dbef1cc040d63fc98ffc22978

Observation 2688c6b2-61ba-4f3a-b127-b6d76bfd8826 · outbound

This paper cites VerlTool: Towards holistic agentic reinforcement learning with tool use.arXiv preprint arXiv:2509.01055, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning VerlTool: Towards holistic agentic reinforcement learning with tool use.arXiv preprint arXiv:2509.01055, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.065378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.065378Z digest=sha256:9d4469ce44934ef0cef7d069eeff35e81dd75308ff5009885a313efd052ab390

Observation 59f00b1d-851b-48ac-9c49-cf03c313ff85 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Kimi K2.5: Visual Agentic Intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.072026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.072026Z digest=sha256:c173fb099ad623e495169e7f39b88d8452264a864b943e9af47d04ae23efb290

Observation b105647b-87f3-4fb7-82bf-6f0446626b1f · outbound

This paper cites Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.077712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.077712Z digest=sha256:da5f13d211f4238c4d9ffae86c59ebc2b89081aaf1a506ab87e3750e046ba103

Observation 2897de1e-e8b0-45dd-9850-acec5ccdf6bc · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.083137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.083137Z digest=sha256:8740979c6534db887acfb8543f6c18d9d7593f1ca070171db90f22d89e087335

Observation 35cf9fd5-68e4-40d2-972e-a5679640d9ca · outbound

This paper cites Zebra-cot: A dataset for interleaved vision language reasoning.arXiv preprint arXiv:2507.16746, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Zebra-cot: A dataset for interleaved vision language reasoning.arXiv preprint arXiv:2507.16746, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.088153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.088153Z digest=sha256:321ec0b6e32b9da905e6e5d52a2c7858afa672cd272efcb27d8a3b505ee62f8c

Observation 76d54afa-1dd1-45c7-bd61-5613d7a3532b · outbound

This paper cites Tir-bench: A comprehensive benchmark for agentic thinking-with-images reasoning.arXiv preprint arXiv:2511.01833, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Tir-bench: A comprehensive benchmark for agentic thinking-with-images reasoning.arXiv preprint arXiv:2511.01833, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.092965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.092965Z digest=sha256:fa34dc85744e262ea3c7a91685a18b2cb29a362a2aca68d73154334948770a72

Observation 6155944d-2c98-4ebf-ac10-b83d52418bbb · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.097596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.097596Z digest=sha256:690d242f214825c0f435680d91590db21cd7be9508fce484b0ba76689ceae218

Observation 8c4c7087-0bf3-4c1c-ba1f-f214947de793 · outbound

This paper cites On the faithfulness of visual thinking: Measurement and enhancement.arXiv preprint arXiv:2510.23482, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning On the faithfulness of visual thinking: Measurement and enhancement.arXiv preprint arXiv:2510.23482, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.102763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.102763Z digest=sha256:7512f9e6c1e106d2f8fae126ef87dfb924e11cf435db0363de31f9525238a113

Observation fed3ace4-7f7a-4068-a766-df90f1678d4d · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.107674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.107674Z digest=sha256:bc255f8d99dda423ada6fa8586dafc8794e04899e435e710cf440eef57276b42

Observation 62e485d8-49de-4c3a-9605-ae3b6734b188 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Chameleon: Plug-and-play compositional reasoning with large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.113186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.113186Z digest=sha256:31b8cb7a3fff5cbe1f33f49cfb027bb70c8bcbedea906caacc83b91d482583cc

Observation 9004a401-c7d2-49cc-ba01-b07b1b464f24 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.118396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.118396Z digest=sha256:c3decd8ed1ea2ee1081f424113512ee0c34fa4cd91ade34c0afb239db3903ff4

Observation 87e7c981-a900-4a4f-9443-7f23e41f7676 · outbound

This paper cites What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.125300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.125300Z digest=sha256:08f6dc0ddd362ef7b6609f446a7f65138fdde4baa7539397b16ed1a5cdd5d02a

Observation 8c9b6ca7-8c46-4542-b286-5c2210ec5976 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.131378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.131378Z digest=sha256:880bed768e60c57580790890563b3196474a483cb549d24ed834ae415c766837

Observation d915d33d-2552-490f-b362-54b7e010fd32 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.137144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.137144Z digest=sha256:d133e7161fde52b7b427158a742a7fea9dac210cb1eec8674207a4164333fff9

Observation 38d5594e-04c8-45a8-9b98-56a44e4122d7 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.143566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.143566Z digest=sha256:b4ad4b26f3fe709de0e997c1b8f4d5bf4e6b4bfa0b0d62d752170cffeb1546bc

Observation ab05b758-cb33-47e7-92ea-c69c2c1dc7d4 · outbound

This paper cites GPT-4 Technical Report.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.149920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.149920Z digest=sha256:edaa0620822aa7191d0cd63c966dfe2b881fe9ccabae9d7ceb97ea691b90c628

Observation 29b9e7f1-eedf-448c-beb2-59ffb461b6e5 · outbound

This paper cites Thinking with images.https://openai.com/index/thinking-with-images/, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thinking with images.https://openai.com/index/thinking-with-images/, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.156357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.156357Z digest=sha256:442c2629cfd12fd42652cb6834e8501715c7d4bdcfc5922a879cee92325ba45b

Observation 7a9f1ba0-82a1-4a24-ae45-3cfb98252cde · outbound

This paper cites Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.161693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.161693Z digest=sha256:eb563b7a9c224c5a8b07363416c69f2938884865bb708f479805aca69f1f94bb

Observation c9a368e4-f37a-4ffd-b1d6-20ad3b5f427f · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.169132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.169132Z digest=sha256:8bce44e3a9d268addd13d5a188a83ec9120ac5fdfc82c714d662106672fe4eda

Observation d47f3c4d-f832-4c6b-a745-c38dbd5f852b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.175515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.175515Z digest=sha256:2980527c0b6697fa88db691861dd983ee8b2beafbddcb2da778c982bda7368ab

Observation 25ad21de-9c81-4709-b03a-9f0f39708c29 · outbound

This paper cites Qwen2.5-VL Technical Report.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Qwen2.5-VL Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.182177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.182177Z digest=sha256:cd87c0eef64a4ad53f8a2dba911a0b38174f6c12c90fa6f552b24ea9ba7f67a2

Observation 8c0aada3-d18d-4ebe-83af-b05f2c1bd462 · outbound

This paper cites Qwen3-VL Technical Report.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Qwen3-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.188571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.188571Z digest=sha256:a991ba00dc75797a1dcf18bd2c1ababe4ab22a1b04ae909e08d7666eaf919bb8

Observation fa21b7df-5558-4e33-8163-a62a77acc62f · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.194557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.194557Z digest=sha256:502fd7dded0fec445d2fd01e0d71fd6d1cf396c00f230f04c31eb037a3e6c183

Observation b553977f-9a03-4e10-9724-7816c539dd03 · outbound

This paper cites Grounded Reinforcement Learning for Visual Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Grounded Reinforcement Learning for Visual Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.200723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.200723Z digest=sha256:8cbce8dfa5faf1320bde986325195aae10c4bd3e4e24b7c7e0d15d2f64541e16

Observation 71aed7c4-6d94-4238-a052-da220eee155a · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Toolformer: Language models can teach themselves to use tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.206226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.206226Z digest=sha256:a52534fd6ba4217357a483049d13889ba1faf30479d2fbcf58bd829d2293adb8

Observation eb554f4a-9642-4522-ac1f-13538d74085c · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.211374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.211374Z digest=sha256:3ce015f0ca4c1b82aea2bce6983b3bd750b43e05cef2ae5228ce058034c3be42

Observation 36606029-2231-4dbc-87b7-538e2926826a · outbound

This paper cites HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.216761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.216761Z digest=sha256:20573ff6f8004a1587d3820532e406483ffe9d737cb893a4162725863ffa7efc

Observation d80610d9-2730-4bd8-bb41-1c6a246f93d8 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.221402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.221402Z digest=sha256:6180d6053e4e51b49d463da94c19c2d0b5ceea15334eddaeca1ed5d16100fa4f

Observation 7039400c-ff0f-4515-a00c-e38454fb4b39 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reflexion: Language agents with verbal reinforcement learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.226408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.226408Z digest=sha256:292930fc5f4df7e77db501be87922ccf885debb3ac11d5db021101540699a5e9

Observation bd7094a6-99ef-4afb-b609-325ea8259745 · outbound

This paper cites Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.231434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.231434Z digest=sha256:8ef6ed0778ba8a320b9115331bf1a3fd5f59a120182b317a5275517c5e0a15bd

Observation 2f895714-0304-4781-afac-8325cbee7e5e · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.237048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.237048Z digest=sha256:eecb1dae2e050d480b0f3cf7b3b1a18dcdbfc7598c379930dd0076e19024553d

Observation 6adf3fcd-a92f-4e84-9f6c-f35f8fdca13d · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.243692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.243692Z digest=sha256:fa7bfc9a8f5c109b3671f63ec60dd26d72a287425559add83ca91bf56f60dfa2

Observation 808134bd-c4aa-41ce-8ced-e241ff856bee · outbound

This paper cites When thinking hurts: Mitigating visual forgetting via frame repetition.arXiv preprint arXiv:2603.16256, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning When thinking hurts: Mitigating visual forgetting via frame repetition.arXiv preprint arXiv:2603.16256, 2026

Reference 56

Resolution
verified exact
raw_fallback, observed 2026-08-11T12:53:16.994849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:53:14.249540Z digest=sha256:8727662e7b38c4756dcfc6865fa1dcace614fb5d796f7f22138ec0cda8a481c3

Observation 33cfd61d-327a-47bf-831c-7f2f998d7ef9 · outbound

This paper cites FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:53:16.902776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:53:14.254925Z digest=sha256:3781842eac9b110bbebcaa65e0db44243263bd4ec805aebf61ab96e6ec064ab6

Observation a6faa050-844b-4cd5-9dc9-5a37b4bcd364 · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning ViperGPT: Visual inference via python execution for reasoning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.260904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.260904Z digest=sha256:738d672bdd9b59e175524ed878ebfc8dd65b76bc2c189fbacdf0e86f6bfccc3b

Observation c664eea2-5da5-491d-a904-af12e5d86968 · outbound

This paper cites CV-Bench: A computer vision benchmark for evaluating visual perception in multimodal language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning CV-Bench: A computer vision benchmark for evaluating visual perception in multimodal language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.266490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.266490Z digest=sha256:8e8a22b9b9d6f9a1510f388c701480fcbe3835ecc0a2621ba12121ffcbe1370d

Observation 984d86a0-d722-4cdd-a983-895defd786b0 · outbound

This paper cites an unresolved cited work.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.271924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.271924Z digest=sha256:f845ddea820fdb0db0643641c4bb6823dd01bd5ba88bc9e2b5dfc7ecdf104832

Observation cb4caa7c-52de-4d80-8dba-fa757ff24771 · outbound

This paper cites Journey before destination: Visual faithfulness in slow thinking.arXiv preprint arXiv:2512.12218, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Journey before destination: Visual faithfulness in slow thinking.arXiv preprint arXiv:2512.12218, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.277132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.277132Z digest=sha256:9720fed8825b5522435012fa7ffaccaa83f5749f44ecdcb531e18549b67cf993

Observation aa0df45c-08c5-43d0-8474-d7483c923177 · outbound

This paper cites GeoEyes: On-demand visual focusing for ultra-high-resolution remote sensing.arXiv preprint arXiv:2602.14201, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning GeoEyes: On-demand visual focusing for ultra-high-resolution remote sensing.arXiv preprint arXiv:2602.14201, 2026

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.282612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.282612Z digest=sha256:5f674f9ecbd21d747b4443e6f80f93e8e0899999064a28dc2047fb126d60f5aa

Observation 0fcc6dce-7fe1-493b-953e-1838a6248d78 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.289490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.289490Z digest=sha256:64a25d4ab1fbcfc6a1aa733a817272f5da0d40270098c95f208d83564c6eaf11

Observation 431c9e90-1e5c-4855-8d17-1096829a3f45 · outbound

This paper cites PLaT: Latent chain-of-thought as planning: Decoupling reasoning from verbalization.arXiv preprint arXiv:2601.21358, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning PLaT: Latent chain-of-thought as planning: Decoupling reasoning from verbalization.arXiv preprint arXiv:2601.21358, 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.295877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.295877Z digest=sha256:306baaeb0ec5641c9282aba5efef48bcf551450a0f66761629c432028951988e

Observation 13f7fe4e-f9e3-4b26-924d-51d354c8d659 · outbound

This paper cites VAGEN: Reinforcing world model reasoning for multi-turn VLM agents.arXiv preprint arXiv:2510.16907, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning VAGEN: Reinforcing world model reasoning for multi-turn VLM agents.arXiv preprint arXiv:2510.16907, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.301424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.301424Z digest=sha256:2d2afab7374fde87e7ee8a7244bf8e30fddf5243b64438f2d0c89fa7bc876c37

Observation 155234b5-f117-43e0-ac3f-685adee465c4 · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.307438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.307438Z digest=sha256:61fa6e358e19aedd7d9552145ef0b1c3efc2b21061aa3d0ad0e271216ca9be13

Observation ad4de20a-fa87-4e81-a3f1-935df8556f00 · outbound

This paper cites A practitioner’s guide to multi-turn agentic reinforcement learning.arXiv preprint arXiv:2510.01132, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning A practitioner’s guide to multi-turn agentic reinforcement learning.arXiv preprint arXiv:2510.01132, 2025

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.313080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.313080Z digest=sha256:109476534acf8e8b5f3ce00e5bfb368d4fee44c62dc8a6fd9a7a2874922a7690

Observation 7e8456bd-7ef2-4fdb-86cb-b3b5d1e618a7 · outbound

This paper cites Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.318402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.318402Z digest=sha256:83e2e28ef9f247bb04b2a9eaf274dc47f5d11e3d360daac59bdcc3619b0187b7

Observation f574102a-4bd5-4cf0-aaab-eeb6bc7aa745 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Self-consistency improves chain of thought reasoning in language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.324074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.324074Z digest=sha256:3b68c68f50ff05c30b0ec6bd3feb435400125bdb21d8989a4b3169a253a0196a

Observation 11876701-c04c-48cb-9602-4d5af4eb6f62 · outbound

This paper cites Simple o3: Towards Interleaved Vision-Language Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Simple o3: Towards Interleaved Vision-Language Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.330057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.330057Z digest=sha256:f9acb67b0b04cb19ff59eb31cd1493ddd5cc468808f0bf1f61d583121a581ec4

Observation 672af4c4-9a5f-40d2-abde-ad12cb180fd6 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.335916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.335916Z digest=sha256:1c85c63282a2f1d6f5a0625e09fac1c560334a141e3ff325ae6e8e341cf9ea64

Observation 58bc7e5f-585b-444a-b826-0f46909038fb · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.341547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.341547Z digest=sha256:d13b30bcd5a923ed12365f7b10b740034ed7235e01988f007d9c737b70163a1e

Observation f6834941-8cc1-470f-bd19-d04bb69d18b4 · outbound

This paper cites V-FAT: Benchmarking visual fidelity against text-bias.arXiv preprint arXiv:2601.04897, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning V-FAT: Benchmarking visual fidelity against text-bias.arXiv preprint arXiv:2601.04897, 2025

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.346906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.346906Z digest=sha256:5babb7b83a216b17cf3bf7f84db5f6a095f0648d60ad58c978613f2bace3026f

Observation 75cb0d3c-04e6-46b5-a8e7-454ba0800c7b · outbound

This paper cites Chi, Quoc V.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Chi, Quoc V

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.352064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.352064Z digest=sha256:5f36317f8bed09b610e7dc9fd2589b909343fa0de325301e30fb7baf705a771d

Observation 8614fd79-8787-43d5-83b9-21608da229f4 · outbound

This paper cites Zooming without zooming: Region- to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Zooming without zooming: Region- to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.357126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.357126Z digest=sha256:a5f8136d96e6cd75e99d6f3e92ea35c3b4ac4959948ab89f9043db4e926bb5f7

Observation 10131f61-7daf-46c1-82f6-f0f69697db60 · outbound

This paper cites Visual generation unlocks human-like reasoning through multimodal world models.arXiv preprint arXiv:2601.19834, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Visual generation unlocks human-like reasoning through multimodal world models.arXiv preprint arXiv:2601.19834, 2026

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.362216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.362216Z digest=sha256:c2078a9796b84549edd6fa471f29f2defc4d80c08c37ee79cc89917e39eea671

Observation 812332c9-4101-4395-ad7a-532e7f106164 · outbound

This paper cites VTool-R1: VLMs learn to think with images via reinforcement learning on multimodal tool use.arXiv preprint arXiv:2505.19255, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning VTool-R1: VLMs learn to think with images via reinforcement learning on multimodal tool use.arXiv preprint arXiv:2505.19255, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.367385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.367385Z digest=sha256:d14d9de4a88cfed7830bc7ad379252ee781489181d1d13c467525c6cd538ec81

Observation c1358137-74bd-4884-b10f-9687c7390116 · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.372654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.372654Z digest=sha256:638371521964fc493b813081d93d5ff91b5feb714d8f569a10b017fd4a499472

Observation 39fdb19d-0830-48e0-bb5b-af6e36d48e4a · outbound

This paper cites Tool-augmented policy optimization.arXiv preprint arXiv:2510.07038, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Tool-augmented policy optimization.arXiv preprint arXiv:2510.07038, 2025

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.378336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.378336Z digest=sha256:d8579ef1610eff413f5b8d43563167395eee8c391f0befb75c2709fcdfee30af

Observation d2e8fa83-702f-44e8-bdb0-0b664623ccfd · outbound

This paper cites Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.384423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.384423Z digest=sha256:316bd3f2ae1de9db5cd08667e7bc7172ba5f47487e3601cf66a54b10f1d5ea34

Observation e0bf9a99-100f-4ea1-aa7d-15c558030ec7 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.390585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.390585Z digest=sha256:5fb4a43206b94d859aa27db718bee37b6a9b6d35ee64d0e911fae1ed0972e329

Observation 68cb5366-89ba-4258-9f1a-99d967af8c4c · outbound

This paper cites Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.396078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.396078Z digest=sha256:be1f32f80b0aaffdb1922bd35e513ef0678657cc96c499b2c97d42e05146a59c

Observation 51316cec-d858-4612-af50-38963eb96922 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.403481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.403481Z digest=sha256:d789214c6381db452666f6d0ea5ca31f8087958bbc024837b76c6af47f49f711

Observation c22ace32-3b90-4012-b5f8-9e466331462e · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.409420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.409420Z digest=sha256:2d3ec705894554b210fe859218fa3fafafcab327deed22e77331ed177e7498b7

Observation 5e606f51-b262-42b5-8073-9e895eb919f4 · outbound

This paper cites VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.415257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.415257Z digest=sha256:b01e8d2f8f03223f4a63b69447689589ad6d4a9f835b8a47dbb449e6f1fe2699

Observation af5e4430-ad7c-4533-b666-a57b79c4825d · outbound

This paper cites Look-Back: Implicit Visual Re-focusing in MLLM Reasoning.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.422649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.422649Z digest=sha256:f7164757ea1027512456904979afccae9953a748e383e27bdd12f041a24c5b90

Observation c46a5d03-1cee-489d-98bd-52fb90f9d5b4 · outbound

This paper cites Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:53:15.665161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:53:14.428800Z digest=sha256:64abe2ef7d1a4bf31ba6bdf6945e61d746c5ba837a5b36b2bf7611f35eb93512

Observation 6dc2fa7b-8836-4961-aa4b-e74d4e6e85d4 · outbound

This paper cites Thinking with images via self-calling agent.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thinking with images via self-calling agent

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.434700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.434700Z digest=sha256:88c8e5caee88bcc67674714e7170e65bd8479f9153475da03880d7142a568adf

Observation ebbc4565-7328-45a9-9e9e-c80012b9984e · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.440324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.440324Z digest=sha256:1d48cfb7bb3b30d35989599d1f338254f323854e56b18606d70dbe2806d3fad5

Observation 6e03bf6e-d596-4191-833a-6f05b55e852a · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning React: Synergizing reasoning and acting in language models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.445645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.445645Z digest=sha256:9e69bbb46d50e1d0507aaaa181d439ba63caa2193b081a7dff8f878d54b55614

Observation bb86f08f-a4d6-4b04-937e-b32d0aaac4c4 · outbound

This paper cites an unresolved cited work.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.450944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.450944Z digest=sha256:def62670af92d67e6151f12787e3f13bea505012c8c52d7c2fc7c896feef5dca

Observation 93488707-e50d-4f90-8a7d-95b8d82a412d · outbound

This paper cites ProRL agent: Rollout-as-a-service for reinforcement learning training of multi-turn LLM agents.arXiv preprint arXiv:2603.18815, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning ProRL agent: Rollout-as-a-service for reinforcement learning training of multi-turn LLM agents.arXiv preprint arXiv:2603.18815, 2026

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.456716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.456716Z digest=sha256:76e2df5a0d39ec1488e54e1e7463174d1b7960768e2f3eeac28a0a39dc009bf2

Observation 1da67c1b-6c4a-472f-93ea-6f3d8f059e62 · outbound

This paper cites MM-CoT: A benchmark for probing visual chain-of-thought reasoning.arXiv preprint arXiv:2512.08228, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning MM-CoT: A benchmark for probing visual chain-of-thought reasoning.arXiv preprint arXiv:2512.08228, 2025

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.462467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.462467Z digest=sha256:acc84beb4c74805e65cb0a2dcfdd4971f633fc93d68359a0e531abf5fc7b25f8

Observation 1cde152f-f8de-45dd-8e0e-f00f4a94f582 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.468007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.468007Z digest=sha256:64d6047282205d5164bd93c30b044a1b0be2cb5ddd922ce88a9b1007f0363798

Observation ef2d5d0a-b1e9-435e-9982-f1221d152fad · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.474606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.474606Z digest=sha256:c91d10aa9b93b7e3fc1d4f85ffd5c0a00565f839a3d4e844549a08e126926953

Observation c12feb7b-3ba7-47e1-a0bc-5cb150c124f2 · outbound

This paper cites Thyme: Think Beyond Images.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Thyme: Think Beyond Images

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.480733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.480733Z digest=sha256:c518245233dd5056cc657600add3969dd1c986e7e2e2c07227f02e22270d72ed

Observation 4fb149dd-5d29-49f4-a574-6c1c48b07001 · outbound

This paper cites Skywork-r1v4: Toward agentic multimodal intelligence through interleaved thinking with images and deepresearch.arXiv preprint arXiv:2512.02395, 2025.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Skywork-r1v4: Toward agentic multimodal intelligence through interleaved thinking with images and deepresearch.arXiv preprint arXiv:2512.02395, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.486447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.486447Z digest=sha256:b04ff3f2dd65d4acbd42440ac69eabb3764befa184ea5bfbab1dfe9df248dc83

Observation 3a376dd0-d944-496d-b3da-cffe00b77a20 · outbound

This paper cites CM2: Reinforcement learning with checklist rewards for multi-turn and multi-step agentic tool use.arXiv preprint arXiv:2602.12268, 2026.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning CM2: Reinforcement learning with checklist rewards for multi-turn and multi-step agentic tool use.arXiv preprint arXiv:2602.12268, 2026

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.491968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.491968Z digest=sha256:2290b0a33e688cc645acaa9f9aa82f4690fc97b9f290375506ec05a649f8b29a

Observation abebf7fd-87a9-4b5f-ad26-07e75afc7002 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.497323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.497323Z digest=sha256:f66a175f499291943bf7d4ce86d08877dfca65cc808fdd4bd024737dd4da243b

Observation 754fb110-bb38-4730-85a6-4d1842ee219b · outbound

This paper cites On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.503030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.503030Z digest=sha256:2203e94f5446d5ae5277d43a5cd57db24252bc3ebd3ebde04856c99a4bb028ea

Pith citing papers

No inbound Pith citation observations are available.