Pith. sign in

Paper Citation Record · LEDGER

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 10 inbound Pith citation observations for arXiv:2505.20289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20289 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:40.120459Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:30:03.053598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.147388Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdd3af60-36f6-4ff6-9905-7332016ec311 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.679510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.679510Z digest=sha256:08c18e0b57f6924bf6a97077b6a89ebdcae46e99fadd19a205f680b70d2d0ddc

Observation d08f3d56-2dfc-49e1-80c0-075079462928 · outbound

This paper cites Language models are few-shot learners.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:43.052275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:35.772704Z digest=sha256:9651817bfbd5c3b381676931fe7a0a93ff01290fe922f536e96fc525ceb8c5f4

Observation ab323fc2-493c-4d18-b486-9882846af7be · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.839303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.839303Z digest=sha256:c57ef0fc92d0b6f25ec45eb7772b85d4dc95b940412fb3f1f037a7e18cb29166

Observation cdba0ffa-9c85-4a20-a6ba-90a550228e0e · outbound

This paper cites GPT-4 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.888788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.888788Z digest=sha256:4a7da1fb6c04fbcf8edb39bb8401094ae31f98040b00acd6c255174de5763592

Observation 688941ca-0a04-47d9-b69b-2e2700a5f615 · outbound

This paper cites Visual instruction tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.927742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.927742Z digest=sha256:d77322d56801f089f53b00ef049fb8b8c9179be4593fc90babb1e145a59bc5ee

Observation afde5757-ba96-4b33-b939-a1c33aab25b7 · outbound

This paper cites Qwen2.5 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.008311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.008311Z digest=sha256:79f353b9698e45fa2c1686356a83dad9997da818bb9674264852795c2c2c93a1

Observation e01f5360-5b6f-4a5b-b824-68d7ebcb87cc · outbound

This paper cites GPT-4o System Card.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.102797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.102797Z digest=sha256:89bb16f58623afe096ad24ed1fe3323ca5109a5a16b09a0763fd775c57375dfa

Observation 3f8024ca-e64d-44f5-8ce5-d7d481dfac4d · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.173060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.173060Z digest=sha256:4f468c37ea0720ad9c7ea461e78bd3fb437a3a298a2996d86759a0ff797d8310

Observation 8de9d6e9-170e-4728-98a1-1e4079043ccc · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.266705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.266705Z digest=sha256:61299c3d2937448eb653c071484b1bd00f89fd08a5aae36b3877fc9cc5d4ee77

Observation a73f13ec-d452-470f-9684-e0d8ba91a76f · outbound

This paper cites Pal: Program-aided language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pal: Program-aided language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.332759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.332759Z digest=sha256:433e00bec04f72f041d0dbd33a09bed8c4f34c694dc09cc01edbe92d0eed77c9

Observation 5e26eaad-20dc-4f94-a5a8-e12345f4ce40 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gorilla: Large language model connected with massive apis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.818668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:36.414302Z digest=sha256:c80f7067194154d42b2af76e9a4f1e8fb8623eb4e66c47254d8bf147fdbd9c9b

Observation 2fa008a5-8011-47d7-abef-8f08ead072a1 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.478765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.478765Z digest=sha256:d09b60702ec4cb670f4c92a91c1355ac826360e46e48c1bd63d50c9ea1d702f8

Observation 8875f002-9e09-4fbc-8494-38984220fb19 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual programming: Compositional visual reasoning without training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.556198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.556198Z digest=sha256:23e5293b973b0e6da87d754f8ae279b5683a61b18866aa8a6fb91ee5b6b4fa80

Observation 9de7870e-40fa-46a8-8b5c-b1aa33c00c86 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vipergpt: Visual inference via python execution for reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.614098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:36.625081Z digest=sha256:7cdaff5e1ede726caf6215297c177eebe53e087618979b5ff30473a4d7c18f90

Observation bff801ff-b139-4cb6-929f-eeb09cfb4686 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Toolformer: Language models can teach themselves to use tools

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.461624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:36.697970Z digest=sha256:5b4395211b5b8bbe33006660cbcfa157bb6ea746539612e92986f56b72d848db

Observation 6da95cc0-cc6c-4f67-96b8-8da502b1191f · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Llava-plus: Learning to use tools for creating multimodal agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.748773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.748773Z digest=sha256:4cb328b89ebcfcc4243c427cc7aa8a54e2da1856da1c9ceba89ccbd64527a824

Observation 98792aca-f97f-4bc4-baaa-cb14ae54f20e · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.843545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.843545Z digest=sha256:629104b7b0e02ac02cd4400e684ba358b3640a1bd5edd5de95a9b42447218f85

Observation 4dd80ee7-9102-4d5c-b7ef-a280e31fca07 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Reinforcement learning: An introduction, volume 1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.900888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.900888Z digest=sha256:79834db398b4ac1de8448a32346c718a853a95f02696bae460442affb2466a82

Observation 72fc4b0c-cf1c-40de-a9e1-66826631bbe0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.982009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.982009Z digest=sha256:12a61ce54332f498b3a0919befdb35b3d93c59423e7efe79d9b175679ed34be7

Observation 14f40af1-03f3-4515-ba06-7b0a00e206f9 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.062514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.062514Z digest=sha256:f678ef9e0bc09e284a500efe1f3d2659a8d4193cb7c0fdbca62133d26688992a

Observation 28f6d4d3-57c7-46a2-bf52-6987a05e02fe · outbound

This paper cites Internet-Augmented Dialogue Generation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internet-Augmented Dialogue Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.132107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.132107Z digest=sha256:bc682848c57ae9bfdf3921584020a10560cf5d076f783cb51222498353d0bcad

Observation 4aefe993-3f2b-4298-84de-a1e815114d44 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Training Verifiers to Solve Math Word Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.207212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.207212Z digest=sha256:ae48e1016247c0472922a28b32e9d58bbdd460c2990c0a698745a0d4fa8845cc

Observation 410dfcc5-d158-4ba6-8771-3b41ddc00ba4 · outbound

This paper cites Learning to reason with llms.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Learning to reason with llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.275098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:37.248131Z digest=sha256:5752efb57f6852e0ec38777b6fa8a28fb868373f356cf852114837a1274b764b

Observation ff8373f4-ace7-4f6b-a2cf-ffe6e4e5a260 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.341748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.341748Z digest=sha256:cec051070fbc9a9a5fac4a4c013f3df9addf058d26fa26d9aac498061c0eeb4b

Observation 100c0509-79ce-4388-bf35-80bd2e7076ea · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chain-of-thought prompting elicits reasoning in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.420694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.420694Z digest=sha256:ced68bc3b4790ad60ebcda13b996eb05848473ad8ead0fc2b351bf4357202ecd

Observation 4c15750b-4aaf-4190-9fda-8f7d5351c0ba · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.501506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.501506Z digest=sha256:62d9d5b1ca5faa1b1f282bfd0f9228b94b8eb38a640d710823356eba063d4104

Observation d494e456-33be-4295-947b-67488fb35dc9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.624601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.624601Z digest=sha256:04e0e5efe74ff01d64260b0d5be56e556e00a079e98c555764e10736d42e114f

Observation 7414b3f1-a4ed-4a47-96b6-565e97e5c9cb · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.694952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.694952Z digest=sha256:e12d4f13eec876232c4a973414fc3627a84df9a0e721cb9e64daab1978c135e7

Observation 301746a9-4f66-44b0-b061-cd7eb6fe2622 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Blink: Multimodal large language models can see but not perceive

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.771338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.771338Z digest=sha256:28b64437df810863c6f9034770e1ffa73950c04449261d784d318285ec28011a

Observation c5905332-3d3c-4336-9b61-542c0e738dad · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.840955Z digest=sha256:30ac4f195a709f231dbe928620066501c93891af025d9a98c9a463051205456c

Observation 79c0902b-3de5-453b-b578-d6bd921dc99a · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.936932Z digest=sha256:5b911f08072c4296bed102af700beeaf9163343fd3a5a6e60a4def45d996156f

Observation ed414bfd-688b-4641-afb4-c2a87b98951c · outbound

This paper cites Decoupled Weight Decay Regularization.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.012150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.012150Z digest=sha256:9ef51a27f83c42e4e77f8b86549aed0bfa26d3e769bcc03e6f677ae6290f209f

Observation 13414287-fbf4-4ebc-be64-4472e5db32cc · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.116421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.116421Z digest=sha256:82cc9c3d9f5205376b98e711fdd6ad42a8a95960509dfe2924af960fd4754761

Observation ad3c211d-e2e9-42da-81f4-cc249b8695fe · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.207103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.207103Z digest=sha256:6135eef8399dd663c67737284ba0381b660ba9265dcdec312bbf0a07a9f464b2

Observation 82d21851-fc21-4c4b-a93d-a59872686a91 · outbound

This paper cites ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.348879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.348879Z digest=sha256:ece56b4ebff0eb0d1bce3f5ff84846cd986d93854311ba7339031db985f89d54

Observation 3e2e839f-1f4f-446f-8e05-6288e7069970 · outbound

This paper cites The opencv library.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The opencv library

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.104637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:38.471275Z digest=sha256:79b686e7da2e857a20c405eadc20af2fe34cef90e28a5d6f5069cf4889d96e58

Observation 09391139-cc9f-4177-b139-9a2919c9105d · outbound

This paper cites Context-aware chart element detection.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Context-aware chart element detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.869761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:38.589707Z digest=sha256:45aedaf9933bfdf3c91ccde9a0c4e3088be79d4b9e3079801f87d3a22ad835df

Observation 9a651a89-6b09-4ff8-b9b2-645b107eff7b · outbound

This paper cites Chartocr: Data extraction from charts images via a deep hybrid framework.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chartocr: Data extraction from charts images via a deep hybrid framework

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.686448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:38.709688Z digest=sha256:6e886451a9b01517d5e3e50a7c7e960cd4571c397e748dff1a3560f2bd9d5f7c

Observation 51f9c9fe-6506-45db-9fe9-db05c6fa83c3 · outbound

This paper cites ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.807547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.807547Z digest=sha256:e629837a79756cc4d931ddcd44567f242be50c088006a9fb72f786e747badfb2

Observation 250dc581-5b92-4aa4-9bf2-6ca6872fe5f6 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.929113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.929113Z digest=sha256:67a700ddaf8a34f25ea1024df7ce8b8c1e4ee9e9d9b79ae0fe44032752ea47e2

Observation 1d2389d6-7575-4bad-bc9a-1c287df02f90 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.039095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.039095Z digest=sha256:34444aa1baadabce9b53786329eb5aa81a51a50673e1dde59b220250cf0fa6bc

Observation a9d18b27-1551-49af-bdd2-6886ad4bac68 · outbound

This paper cites Diagram formalization enhanced multi-modal geometry problem solver.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Diagram formalization enhanced multi-modal geometry problem solver

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.489164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:39.114084Z digest=sha256:129283f51cbae4c8c80c10f407b85fe802ebcc65bf93fcc97e63fdbb0316a68a

Observation aa07cd9b-8274-4328-8380-29586848cb9b · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.192083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.192083Z digest=sha256:34db56e2c0f94d888b91c22b5be8de8f0ea951c8c124d838420de47ccebbec9c

Observation 5c99f6ec-2a88-4c96-82ec-425ab61a8f4b · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.265256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.265256Z digest=sha256:dbf175feef2356d439f1ae566cfd316000fc622f18167cca0d72421e137ac84b

Observation 10cac71c-70f2-4a17-9fb6-b6f207e1c2a6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.319182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.319182Z digest=sha256:728e3726de63b59e6f2a0ff8f82a0e42705e83aaf1d8898aa05676863f35a93a

Observation 5ed9a8b4-64cd-4945-92a2-09d1c88075be · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.365404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.365404Z digest=sha256:7e9608588ba39844369771054cf4d40a16e1e2f5ec735c93da3bfd4ce54168b9

Observation dcc3e15b-a8b2-4a86-8408-0a0b0d4808b2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection PaliGemma: A versatile 3B VLM for transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.424086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.424086Z digest=sha256:d4e5f6a0aa0093d7cc716a941d77b91954d3c1f13609eef2b78c457e1664489a

Observation 3dfde684-e0ac-4615-8c21-75abe6408fa0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.471308Z digest=sha256:742e2ac22789e01b1d496d64805cbccaea09e40279247a71b20400a94f9802cf

Observation 9a6983b9-d7e6-45de-9416-b3084d58b5e0 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.333416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:39.533698Z digest=sha256:8214eae0babd289f806a6b177c43e0d9690121b3e2463313058a5a99fb43dd69

Observation 1f86a0a9-77b2-42c9-9c11-5294c7f4e863 · outbound

This paper cites The Llama 3 Herd of Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.588544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.588544Z digest=sha256:88d5386634c113ddcfc2d410278b6e0478a66cfe81ba9ea8540e5f5487751008

Observation 6ea81de3-cb5c-44bb-b0c2-ccfb6e8f4fbf · outbound

This paper cites Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.165486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:39.642125Z digest=sha256:fa621a8e5f3a859a264e29ddb6417be1b05e34a7db87108e219b32f0dc065fff

Observation 1b80ee2c-5315-4a66-8ea4-acccf288e58d · outbound

This paper cites Pixtral 12B.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pixtral 12B

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.719031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.719031Z digest=sha256:3ef7f069246d2c913b8ee3b6cef901c73edbbfee9b406ed485b85794f61d2b82

Observation f25f61f5-941e-4ba4-9855-539b3bfd3da8 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.764847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.764847Z digest=sha256:e47dc602591a8ffe45fe69272b3ebdd2ef0b7c150f776b8d31bd79aefc11a3e1

Observation 3006adf5-42af-4746-b61d-4e572fd8820a · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.820900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.820900Z digest=sha256:689fd5a9e737f8eded42ec113f3153149481a7b0a24ef22235c737942c66a0df

Observation 6cec24d0-f256-4c17-9cd9-8c44588cd93e · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.871961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.871961Z digest=sha256:523759d722bd2caf3264e549bc8ba2b3fc65e01b41d4de01b47ba83138f74cb2

Observation a2e7b4b2-2738-4452-a79d-b384cda0cb89 · outbound

This paper cites Improved baselines with visual instruction 15 tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Improved baselines with visual instruction 15 tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.966064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:39.941166Z digest=sha256:e4084316cd1e079a78be9e6bac44f8814fc9f83ba7bd3a50991d644e44fc6f2a

Observation e1a78340-085f-4938-a3ba-adeab8c7d730 · outbound

This paper cites xGen-MM (BLIP-3): A family of open large multimodal models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection xGen-MM (BLIP-3): A family of open large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.006031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.006031Z digest=sha256:10822f24b6c5d48c925b00e27b0eb71811011c802daf93fde01d87780d8d48c1

Observation 5361dfab-5fa2-4314-a908-cde932edbe19 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.054356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.054356Z digest=sha256:c39988a2c7ec0000ef58c5783c3f0064391adbb752b60e016641bc2bd1f05ecd

Observation 248a11fb-53c1-4729-a873-46e668f615f3 · outbound

This paper cites Vision language models are blind.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision language models are blind

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.721505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:59:40.120459Z digest=sha256:98360b2e8aa783b5a89d44f89934659db91b491cc3109f6db0a5c47d8b0510fc

Pith citing papers

Observation aa875bb9-c0b3-41b7-8c59-6dbbba4db11f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.261893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:428e7baeaff0e082b4b09cd4440f4e1a5e113f4e101b60b3984f5f7d4902d256

Observation cb67b269-a8a5-41a4-974f-2f01fb059d25 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:41:30.400550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:785386964ce8b2821173d3cf5c4cefc8ce6f6cc5cf312b4ec0a6ebd604f021eb

Observation 9b5f1228-cac0-48fb-9481-8f569a5422a9 · inbound

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents cites this paper.

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:10:49.569781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:10:49.569781Z digest=sha256:250bcaf3e8b387ab8c78e4efc73be11ab55dcc87ce93bf0d27f0c8dc91d6638e

Observation 36ae82ce-49a7-4938-abca-ade047ceada9 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.187051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:934c0404851f6c019a9608c7f6323b3f69c6882193b805f866ad1d4b462ecbc7

Observation e149e1ea-637f-4b6b-b131-7f05cef92eae · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.623176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:b176d40f7d9c38c507df521a1ea915096ce28e442023c737b9f4de961a7b64e2

Observation 11fce3b3-03fa-4695-a0bb-17051f699a0a · inbound

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization cites this paper.

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.702685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:19:10.346738Z digest=sha256:98eb4fafca7d6a51c550a7041f2f0765c02a2e7b7d813e51540241c989dc2fe0

Observation de041ff8-84d2-441e-9567-5a8772b38d6c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.695015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:29f7f2a1731cf4cd0e0f02b061360a4816d071b02862adeedb675d6e8e64b995

Observation f01cae5e-0002-47e1-b14a-147576225dbe · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.149238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:4d3ff4d000fccee47b9756c2702e4f5e4bd786952d62d7a1d724138355641a58

Observation 12c40016-1136-4111-a9a1-a24b220f201d · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.527737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.527737Z digest=sha256:b166e5c5f2601268c927c69629242f93a2e896e77bd9e5337d319247b78a9587

Observation 2ecdc0a2-7efb-4598-a93b-4def83e45495 · inbound

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use cites this paper.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.053598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.053598Z digest=sha256:52900a560eebb0c1f1c356fa76e32ce8e7b876c2c55a04c975fc75c9b32b5ee6