Pith. sign in

Paper Citation Record · LEDGER

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 11 inbound Pith citation observations for arXiv:2505.20289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20289 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:40.120459Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.734634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.147388Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdd3af60-36f6-4ff6-9905-7332016ec311 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.679510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.679510Z digest=sha256:cdfb8a142e5b025089bbee0abf5c24a2aa8ab979e80b0968ae4c255e700f2ab5

Observation d08f3d56-2dfc-49e1-80c0-075079462928 · outbound

This paper cites Language models are few-shot learners.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:43.052275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:35.772704Z digest=sha256:76ff8a6ac06757fa06d317139a18cdab9d8c0dd562589dacc27d52968c38ddce

Observation ab323fc2-493c-4d18-b486-9882846af7be · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.839303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.839303Z digest=sha256:f92d93948226db5299fb4c9a04a8ee7f1ac7e03fc2e7c362e62671a0a7af1bc0

Observation cdba0ffa-9c85-4a20-a6ba-90a550228e0e · outbound

This paper cites GPT-4 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.888788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.888788Z digest=sha256:40893d73b6e5ceb13ba27e2e7853d951cd87a2f8d14a21687e4c6997d410c98a

Observation 688941ca-0a04-47d9-b69b-2e2700a5f615 · outbound

This paper cites Visual instruction tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.927742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.927742Z digest=sha256:fbc3adaa33c120ed14704eb6ad0afb7bc3baabe005c192e5f29c895ecb50a3d1

Observation afde5757-ba96-4b33-b939-a1c33aab25b7 · outbound

This paper cites Qwen2.5 Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.008311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.008311Z digest=sha256:9dbfdd819599f60d15cb2b9fd5b13a7ffdfef9a4be82c3c8de0ab24fc62eb3f2

Observation e01f5360-5b6f-4a5b-b824-68d7ebcb87cc · outbound

This paper cites GPT-4o System Card.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.102797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.102797Z digest=sha256:197e3bfb2e76d6d5e77bbc9297f078521e5ccce057dae726c13441e12024d5b5

Observation 3f8024ca-e64d-44f5-8ce5-d7d481dfac4d · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.173060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.173060Z digest=sha256:27be6380c231b91301d7e77bdc1935c8a0a12f56751db29dc774b86395dbba34

Observation 8de9d6e9-170e-4728-98a1-1e4079043ccc · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.266705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.266705Z digest=sha256:2d2220fb8a7dfd4c0c98899cf63edcc23e481e0a3377b69df42c8dcc770bae1a

Observation a73f13ec-d452-470f-9684-e0d8ba91a76f · outbound

This paper cites Pal: Program-aided language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pal: Program-aided language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.332759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.332759Z digest=sha256:403a7a11c1758a595a14d3ad334b2dc1fe66dac9263571242af006088c4b58b1

Observation 5e26eaad-20dc-4f94-a5a8-e12345f4ce40 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gorilla: Large language model connected with massive apis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.818668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:36.414302Z digest=sha256:9d2ae8be91157bbaba7dbc3b5c4354df51c57d9263b3912ddb5edf7e406d46d6

Observation 2fa008a5-8011-47d7-abef-8f08ead072a1 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.478765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.478765Z digest=sha256:41e4bf7fcdfa93c0610018f1583d0c373ae3ccc22d35033df9e3409f44ce4b76

Observation 8875f002-9e09-4fbc-8494-38984220fb19 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual programming: Compositional visual reasoning without training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.556198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.556198Z digest=sha256:1be3ea410987c84473c1fdda9f9a40ded4b57c8f6c443008423d81a9043feb57

Observation 9de7870e-40fa-46a8-8b5c-b1aa33c00c86 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vipergpt: Visual inference via python execution for reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.614098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:36.625081Z digest=sha256:4fd95dedbb7d29f62e2f04e1f8c82f47c318360bfd6cd6abb01a7e659766de20

Observation bff801ff-b139-4cb6-929f-eeb09cfb4686 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Toolformer: Language models can teach themselves to use tools

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.461624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:36.697970Z digest=sha256:e0803a72b7733c8d3b9cd3d6760dbd5869bbbac9cacf4fde8a4dc642de080909

Observation 6da95cc0-cc6c-4f67-96b8-8da502b1191f · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Llava-plus: Learning to use tools for creating multimodal agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.748773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.748773Z digest=sha256:fa5871a5181f77dd03bcccc751fc760e965650d8238952aa938b9d0174e61716

Observation 98792aca-f97f-4bc4-baaa-cb14ae54f20e · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.843545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.843545Z digest=sha256:3942dc31dee04c5dfe37c96835db7d917d4ecc818e236c19ee4a50d8898a7a13

Observation 4dd80ee7-9102-4d5c-b7ef-a280e31fca07 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Reinforcement learning: An introduction, volume 1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.900888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.900888Z digest=sha256:6cf80ef94791ae9c85bf4660600220f062f8c5378e00c7fc9435a31aadf1f97c

Observation 72fc4b0c-cf1c-40de-a9e1-66826631bbe0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:36.982009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:36.982009Z digest=sha256:895f9fb17a42b5ff65754b8f6c7e9081f98d0449dce5a37d49c6b14a76c8db4a

Observation 14f40af1-03f3-4515-ba06-7b0a00e206f9 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.062514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.062514Z digest=sha256:9561257834f8fe162cb68d3dea4b36a81a28f8e16aff1a7c61ddc5f72bb6034a

Observation 28f6d4d3-57c7-46a2-bf52-6987a05e02fe · outbound

This paper cites Internet-Augmented Dialogue Generation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internet-Augmented Dialogue Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.132107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.132107Z digest=sha256:79fd50153c3aaa0378cc281069143c405defe9f5c2ee8a7d596bf5ff3cd6de42

Observation 4aefe993-3f2b-4298-84de-a1e815114d44 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Training Verifiers to Solve Math Word Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.207212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.207212Z digest=sha256:69c79297c4b08b8273ab2330b59c0461b4a2508173fc0efbf8cfe39981c60e7b

Observation 410dfcc5-d158-4ba6-8771-3b41ddc00ba4 · outbound

This paper cites Learning to reason with llms.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Learning to reason with llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.275098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:37.248131Z digest=sha256:d7fa45ee580a9f4979500350bcbc52c2ce460cff36b86e8676f47e644138f380

Observation ff8373f4-ace7-4f6b-a2cf-ffe6e4e5a260 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.341748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.341748Z digest=sha256:34617c0fb725a62373910cfe69803d8f2ff9d7f8fc58aab33095078bc8b48867

Observation 100c0509-79ce-4388-bf35-80bd2e7076ea · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chain-of-thought prompting elicits reasoning in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.420694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.420694Z digest=sha256:651b4d5096935b8590f09aa9409a97618b7dc35575f94fab7719c59b06beabba

Observation 4c15750b-4aaf-4190-9fda-8f7d5351c0ba · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.501506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.501506Z digest=sha256:22f03d8ee7b20c25f4381616a90771966577d5cc74b8db8dea958cf18cb11406

Observation d494e456-33be-4295-947b-67488fb35dc9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.624601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.624601Z digest=sha256:56a1a11f584dbb2a26b727b17283e1cb84dab442218300b937fb5e4a0c720a4f

Observation 7414b3f1-a4ed-4a47-96b6-565e97e5c9cb · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.694952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.694952Z digest=sha256:cb6f355c3cd3bb9c8ad2a0811a916d4336af75a86914ecb578dddb3aeb1e09ed

Observation 301746a9-4f66-44b0-b061-cd7eb6fe2622 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Blink: Multimodal large language models can see but not perceive

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.771338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.771338Z digest=sha256:6a3be025b61e6f3645ff259e0c9584808b1aaff7d45924246edb343a283bf625

Observation c5905332-3d3c-4336-9b61-542c0e738dad · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.840955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.840955Z digest=sha256:88842a6b34e129aedf5bef4df0538cf12721947ec534e80ac9c42d2d28572746

Observation 79c0902b-3de5-453b-b578-d6bd921dc99a · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:37.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:37.936932Z digest=sha256:8038b4048c04940f84c241ddc6dbd21d2eed001c940821eec0c31bf97ff53993

Observation ed414bfd-688b-4641-afb4-c2a87b98951c · outbound

This paper cites Decoupled Weight Decay Regularization.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.012150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.012150Z digest=sha256:c5955f94c95820ed36d74f53de184869f1442828aa5e14a2ec63e52c3a30449a

Observation 13414287-fbf4-4ebc-be64-4472e5db32cc · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.116421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.116421Z digest=sha256:7dedcb4c53e6a35375dd2087702eedb4ca682abb3d6fc50e31a7ce19af66efea

Observation ad3c211d-e2e9-42da-81f4-cc249b8695fe · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.207103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.207103Z digest=sha256:392eeee664a4c13ee1f6f2cf2508e9b8a44746f2a0baf0a0192c773a975a4fc3

Observation 82d21851-fc21-4c4b-a93d-a59872686a91 · outbound

This paper cites ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.348879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.348879Z digest=sha256:f0f90952432306779e8b6ab58d3b558af5610db73831b15176f77825d7cd6e01

Observation 3e2e839f-1f4f-446f-8e05-6288e7069970 · outbound

This paper cites The opencv library.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The opencv library

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:42.104637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:38.471275Z digest=sha256:7261c8b2f0ab178a048c9a8450fc01f1d760cf6a1009ceb76ab4e33296a345e3

Observation 09391139-cc9f-4177-b139-9a2919c9105d · outbound

This paper cites Context-aware chart element detection.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Context-aware chart element detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.869761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:38.589707Z digest=sha256:2453a32211d79fb30c61fa78ab2b175de944c90518fc1a5e33f02791089ecce4

Observation 9a651a89-6b09-4ff8-b9b2-645b107eff7b · outbound

This paper cites Chartocr: Data extraction from charts images via a deep hybrid framework.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Chartocr: Data extraction from charts images via a deep hybrid framework

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.686448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:38.709688Z digest=sha256:bad740188d4a566053aa9eaa96b5cadbcf91a8e55844d4e8e4229b7211371377

Observation 51f9c9fe-6506-45db-9fe9-db05c6fa83c3 · outbound

This paper cites ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.807547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.807547Z digest=sha256:ea95030b2a83e680f7ad50ba3632c0529d122de81c997a8a63981ac20969662c

Observation 250dc581-5b92-4aa4-9bf2-6ca6872fe5f6 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:38.929113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:38.929113Z digest=sha256:a4c40ed63f0a3e142e9a66b764ddba22e7dd1a6b98d6ffc7b7050519e2e869ad

Observation 1d2389d6-7575-4bad-bc9a-1c287df02f90 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Qwen2.5-VL Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.039095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.039095Z digest=sha256:9c354149fde83b117c7258e436701370c3c2d377a5d531e792a27aaf25909736

Observation a9d18b27-1551-49af-bdd2-6886ad4bac68 · outbound

This paper cites Diagram formalization enhanced multi-modal geometry problem solver.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Diagram formalization enhanced multi-modal geometry problem solver

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.489164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:39.114084Z digest=sha256:e93aa05a137fb37364e6ab1ccce1ee14bf7e6ce9efb6f662401c1767e7cbdd48

Observation aa07cd9b-8274-4328-8380-29586848cb9b · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.192083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.192083Z digest=sha256:1c714e53d26db1c6fbd0ad73c1c49c0124a80c5c5580faddae12914358fdb134

Observation 5c99f6ec-2a88-4c96-82ec-425ab61a8f4b · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.265256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.265256Z digest=sha256:cf0010e12e235b79b49200ad4104668df5a060833d5f54f40b410cc3cff847c8

Observation 10cac71c-70f2-4a17-9fb6-b6f207e1c2a6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.319182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.319182Z digest=sha256:8527b799c0b8336a5d0cc0779eb65725fb12a56677f14082f8033cc3de03f907

Observation 5ed9a8b4-64cd-4945-92a2-09d1c88075be · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.365404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.365404Z digest=sha256:47012644b5b8aee414e504147a090855e675bf7b83c5a94f7c71dbe2108e4d3c

Observation dcc3e15b-a8b2-4a86-8408-0a0b0d4808b2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection PaliGemma: A versatile 3B VLM for transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.424086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.424086Z digest=sha256:a79c8b97bb5e4dc97fb98b024224301bca1fa3432110fd683cde654f4b6192bf

Observation 3dfde684-e0ac-4615-8c21-75abe6408fa0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.471308Z digest=sha256:4da8893196bcdfbc9bc63bac76505143f87b9e1e22898d45645501eb068dea1a

Observation 9a6983b9-d7e6-45de-9416-b3084d58b5e0 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Internvl2: Better than the best—expanding performance boundaries of open- source multimodal models with the progressive scaling strategy, July 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.333416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:39.533698Z digest=sha256:df23834476e57c2dd7359a8774952758b58eb47a38a1e24a33aac4785a89f569

Observation 1f86a0a9-77b2-42c9-9c11-5294c7f4e863 · outbound

This paper cites The Llama 3 Herd of Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.588544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.588544Z digest=sha256:11e8730f7cbc0cb59c087eca868cbcc6b3ce88c56a56fe74fabb0b7d1b649193

Observation 6ea81de3-cb5c-44bb-b0c2-ccfb6e8f4fbf · outbound

This paper cites Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Cambrian-1: A fully open, vision- centric exploration of multimodal LLMs

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:41.165486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:39.642125Z digest=sha256:3509a7bd644cbc9b83b167bedadf399c790a11196bee947f17151b0409351e85

Observation 1b80ee2c-5315-4a66-8ea4-acccf288e58d · outbound

This paper cites Pixtral 12B.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Pixtral 12B

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.719031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.719031Z digest=sha256:0cde5dd2a33fed25a36d26d5c01a9feccd9dca8403561abd85555d7586a43368

Observation f25f61f5-941e-4ba4-9855-539b3bfd3da8 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.764847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.764847Z digest=sha256:59d03e4a2c561a289b6a9bd474cbda811e96cb87d4adb6c58d1327e525a9079e

Observation 3006adf5-42af-4746-b61d-4e572fd8820a · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.820900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.820900Z digest=sha256:3b568e99f53eb96f37189e6fdf873a33a88cac0ee07ec7e405e395ccb189dc56

Observation 6cec24d0-f256-4c17-9cd9-8c44588cd93e · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:39.871961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:39.871961Z digest=sha256:da4563d386543399a3fc2ea8df679c24b4a0364a7260a3a81a19ca5d33dcb5e1

Observation a2e7b4b2-2738-4452-a79d-b384cda0cb89 · outbound

This paper cites Improved baselines with visual instruction 15 tuning.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Improved baselines with visual instruction 15 tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.966064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:39.941166Z digest=sha256:fab6b3b5e5c6ae16bfa58e5ea4f77ffa1de112a51a799a397fa0212dcace5490

Observation e1a78340-085f-4938-a3ba-adeab8c7d730 · outbound

This paper cites xGen-MM (BLIP-3): A family of open large multimodal models.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection xGen-MM (BLIP-3): A family of open large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.006031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.006031Z digest=sha256:ddd5fd636d4bb537ad6c50474d7f511ee87346397a5aeec5d22304f35174d097

Observation 5361dfab-5fa2-4314-a908-cde932edbe19 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:40.054356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:40.054356Z digest=sha256:22edbc23588130507586a198368451f8f6b064518a367d8f40710d40bc52a55a

Observation 248a11fb-53c1-4729-a873-46e668f615f3 · outbound

This paper cites Vision language models are blind.

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection Vision language models are blind

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:59:40.721505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:59:40.120459Z digest=sha256:48197ebfce2082e49c140be42bd736cab0bace2b2064c4a3b544fc1b473a2dca

Pith citing papers

Observation 7df5dff1-ffe9-4fb5-9a2a-5e507ded0b9c · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.734634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.734634Z digest=sha256:99980d79b9de19d499484e164f1ed77a54473b63cacfc23c99951cd4bb5b33c2

Observation aa875bb9-c0b3-41b7-8c59-6dbbba4db11f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.261893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a95bb21277bcd2b27c7fbd1120dad46ed19ba8765e96575c31168119d01fb3d4

Observation cb67b269-a8a5-41a4-974f-2f01fb059d25 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:41:30.400550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:ada677e6769c80478c7c3bcdf40b9f2710686d1605bc442adc7656bb31bf630e

Observation 9b5f1228-cac0-48fb-9481-8f569a5422a9 · inbound

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents cites this paper.

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:10:49.569781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:10:49.569781Z digest=sha256:e9aaae1a9b38541db9fc81729c2c4427bd4369a3dfad7581ab1fb59eb7daaf54

Observation 36ae82ce-49a7-4938-abca-ade047ceada9 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.187051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:cc6b2b7eec7d5e665fc5fe0c75dcb652de138071c0cbaa42d5a7b67e6e9cdff8

Observation e149e1ea-637f-4b6b-b131-7f05cef92eae · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.623176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:e88a334bc0b6f72dab7aeedfacd5c30fa772849d33a0f95d362400f117de1e34

Observation 11fce3b3-03fa-4695-a0bb-17051f699a0a · inbound

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization cites this paper.

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.702685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:19:10.346738Z digest=sha256:ae73e397ea877ca7440d8d2980f07e0aea60156dff2dbcf9bdc3846ebce5b5bd

Observation de041ff8-84d2-441e-9567-5a8772b38d6c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.695015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:991703c943bde09938974b13d2bf53d95a66a7c186b3a98a929d7d7c0bfc6403

Observation f01cae5e-0002-47e1-b14a-147576225dbe · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.149238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:ef062df3a5bbdf6ef46bee141c95fe9f64c21a3bcc148846003df01d71460611

Observation 12c40016-1136-4111-a9a1-a24b220f201d · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.527737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.527737Z digest=sha256:9e9971972a2af746fac6ef401121252a65dac9af953fa6ec5ebad055bfdfd380

Observation 2ecdc0a2-7efb-4598-a93b-4def83e45495 · inbound

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use cites this paper.

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:30:03.053598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:30:03.053598Z digest=sha256:26dedec7a82c8d6e54413141c6d8f05f06823169f20bcd615096d337b50e8f22