Pith. sign in

Paper Citation Record · LEDGER

Visual Agentic Reinforcement Fine-Tuning

As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 26 inbound Pith citation observations for arXiv:2505.14246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14246 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:33.223922Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:04:19.981518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eba71093-4b50-4a74-be73-6e6aa8a27e7a · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Agentic Reinforcement Fine-Tuning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.774074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.774074Z digest=sha256:49647d56969312c86b55c2154a23b8ebb350999308afeef25ee2b543a2f3eaa9

Observation 37f415e6-f94f-4cfc-84e7-0dc74aefcce8 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Visual Agentic Reinforcement Fine-Tuning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.832792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.832792Z digest=sha256:de1f3a34306f98da88243b1951274e425d671e8219617e72c363058c72d1a31b

Observation 1bb37353-e8cb-4ac9-86bc-d6b8d4ebf0d1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Visual Agentic Reinforcement Fine-Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.897449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.897449Z digest=sha256:7b5c9019d923055b5caeb5f41a608364556e98fa1bfef814dcc7c06ca882c6cf

Observation 4ff072ad-0201-483e-acca-62a01329f947 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

Visual Agentic Reinforcement Fine-Tuning Rico: A mobile app dataset for building data-driven design applications

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.968524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:25.968524Z digest=sha256:bb4b49de9c64c2013973d07f73541a3cbbea074db77bb002692866ef6bf56a3c

Observation 6a7e0fa7-6b81-4ec4-b3ee-2b234a43114f · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Visual Agentic Reinforcement Fine-Tuning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.032221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.032221Z digest=sha256:7583ca06687f7b868d12b2711b6b735b64af9e85ca9143cf6ebef900136395e8

Observation 56f9d6cb-eb9e-4cb4-b5de-f05489c2026f · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

Visual Agentic Reinforcement Fine-Tuning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.090406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.090406Z digest=sha256:6e729a997ddb3be1e69c2fff7c41b758163cf124dc0ab51b97074728ba9d99f4

Observation a6bd3552-31f0-4518-b16d-d01d8f677910 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Visual Agentic Reinforcement Fine-Tuning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.151443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.151443Z digest=sha256:772c449a36ff2a8ec2918594bf6edb53fac4aac3de5540249394af5c18bf3b07

Observation 98273d02-832c-49ac-a4d8-b0b721538a11 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Visual Agentic Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.221489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.221489Z digest=sha256:8c0e2347e6eb79ce32891d09e4091d34b8a3315965d6661f6014a7e7f6f6a325

Observation 9a671e0c-9e19-4301-8e17-06234cd5b2c1 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Visual Agentic Reinforcement Fine-Tuning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.224356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.224356Z digest=sha256:64919b856b45daeb4d32919eefe444a68761eac29d1645457a58dea9a3879e2f

Observation dd2e3d2b-80fb-4e34-b9f7-f70395ea796c · outbound

This paper cites GPT-4o System Card.

Visual Agentic Reinforcement Fine-Tuning GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.252123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.252123Z digest=sha256:b1d770d22068494f6ab8f009da3be055d1e9c58b0ba8a09699ede8edc8ee8b10

Observation c5eef7de-a9df-42d2-ac17-2866a3206a6f · outbound

This paper cites OpenAI o1 System Card.

Visual Agentic Reinforcement Fine-Tuning OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.389486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.389486Z digest=sha256:3de383a71cfdd1003ab47f9dc2e6d4c20904b34b8cee4d552c38d8f792b8609a

Observation 1798631c-9641-43b4-913d-6c8f4b0fbb2f · outbound

This paper cites Funsd: A dataset for form understand- ing in noisy scanned documents.

Visual Agentic Reinforcement Fine-Tuning Funsd: A dataset for form understand- ing in noisy scanned documents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:35.636112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:26.593418Z digest=sha256:0278f0073807fcf489f184006dd916de6454e03c4b6d20273fe905de6ee43ccf

Observation 2d898de4-7829-41e0-a13d-1c5972db8f41 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Visual Agentic Reinforcement Fine-Tuning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.854529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.854529Z digest=sha256:1c227ed4539e262ef6368b34f390495743f64f90e8c21081a7b771b5e9655b99

Observation 03f2efd7-faa7-46c1-8941-40ac41b53eaa · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

Visual Agentic Reinforcement Fine-Tuning Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:35.419369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:26.990009Z digest=sha256:3357f3c81b7cc4430e4726d21692b8ee69c569a7f456c12fd0d16394a6b449e5

Observation ec2396d2-1747-4864-a0aa-8cf966d3235c · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Visual Agentic Reinforcement Fine-Tuning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.137332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.137332Z digest=sha256:49bb7bad2e80a31a275e66136e97d78ddac160619bc13b802c364878e78bc8d9

Observation 9f8120d5-6b86-41c8-9671-7af117e0f4ff · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Agentic Reinforcement Fine-Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.314659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.314659Z digest=sha256:8969bcb8fc06e2afc55f29a010b005fff98ba4e840eee80628b2984617799a95

Observation 12cec2a7-abdf-4484-a537-a7d49576ec7a · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Visual Agentic Reinforcement Fine-Tuning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.491865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.491865Z digest=sha256:a0393317d74799e6a3357a21e4e4a3e933d97dd8c995bf54f7a84d4d6818e520

Observation 81bd1770-9f80-471c-8d42-e52cec76a3c7 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Visual Agentic Reinforcement Fine-Tuning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.569008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.569008Z digest=sha256:8931340f5f3f8be947621f72d065b2204c2b8f4320079b3ef28625bc30e21edf

Observation 2b789c3e-1a87-4ba0-9b5e-9ad6d8fe20ba · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Visual Agentic Reinforcement Fine-Tuning WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.647588Z digest=sha256:429a57cb3118f130dd5ff3e8f7ca6027cded42dd59aebf2510c8c8da81e0ae17

Observation 242dc923-cbeb-4a5f-93d5-edf32159a6d5 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Visual Agentic Reinforcement Fine-Tuning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.826306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.826306Z digest=sha256:72096da0cac7eb7ec3000be96bbf9121e985740aca2c71af0f8023599f272579

Observation c66baf67-60d2-40f6-a412-24e70996a822 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Visual Agentic Reinforcement Fine-Tuning MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.012611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.012611Z digest=sha256:b2b8742c9ba07733bd7f1330b6f944f7830cf92824cad43c8cdcec00090f2630

Observation 2a16924f-b1dc-4240-b031-a515a10c3bdc · outbound

This paper cites DeepSeek-V3 Technical Report.

Visual Agentic Reinforcement Fine-Tuning DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.137316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.137316Z digest=sha256:f7896206e0a57e1f0999703fb0713a9c9297363f6ff5a17eada3fba34be72022

Observation 655ed0a7-707f-4fc9-9dcb-7c24f1f07623 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Visual Agentic Reinforcement Fine-Tuning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.278760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.278760Z digest=sha256:39a5aa7094a05c827dc428bb436daecef44b03b09ba0730f42c991aa6eb1b344

Observation 85a39c18-1c3e-4050-910b-f8a60a7772db · outbound

This paper cites Improved baselines with visual instruction tuning.

Visual Agentic Reinforcement Fine-Tuning Improved baselines with visual instruction tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.501618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.501618Z digest=sha256:aac48aa64816e410c6c1bff93253ee7427b8610f0aef6d897fec494fcca403a8

Observation 4c6df9bf-e8c5-4462-9379-efda43819b88 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Visual Agentic Reinforcement Fine-Tuning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.688036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.688036Z digest=sha256:219c887853085abc306b4dbc3c0dbd61ebe927525f11787431d1d7326f045db9

Observation 15e7b348-2b4a-49ad-843b-126a7805723d · outbound

This paper cites MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models.

Visual Agentic Reinforcement Fine-Tuning MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.862738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.862738Z digest=sha256:6e9b4d3578fb337e53735973fcc0cc061813f373f4079ec925b29e16f5d78237

Observation 1f32866b-9bd4-4725-8eee-ecdd2497a2a1 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

Visual Agentic Reinforcement Fine-Tuning Towards end-to-end unified scene text detection and layout analysis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:35.148871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:28.982324Z digest=sha256:a5e7d9a73ec8144338472e8be551343d63e6572dc53b8343e903af14b8bcd1ad

Observation f83f97ba-dff6-4823-ac40-8ae1fdb0b1d1 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Visual Agentic Reinforcement Fine-Tuning Docvqa: A dataset for vqa on document images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.180856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.180856Z digest=sha256:86db7bf50eb47b1f75a6d8c0669694f7cd73ddc46592e66606eea4575ee05923

Observation 1d5d6c53-b62d-4137-8bec-5b29142bfd59 · outbound

This paper cites Openai o3 and o4-mini system card.

Visual Agentic Reinforcement Fine-Tuning Openai o3 and o4-mini system card

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.960007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:29.364570Z digest=sha256:bb244fe51641db37bb4ed17303022ab2f112c47cb411128ffe84494173966f25

Observation 130f2762-e4b3-426d-83ae-f092012e2668 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Visual Agentic Reinforcement Fine-Tuning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.490479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.490479Z digest=sha256:0c0f0b7840a14ffcd9a81b947a6ec2d330dacabd09d59e4be9797a9e887baddf

Observation fa4980e5-7038-4d89-a2dc-b53568a92cef · outbound

This paper cites Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024.

Visual Agentic Reinforcement Fine-Tuning Gorilla: Large language model connected with massive apis.Advances in Neural Information Processing Systems, 37:126544–126565, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.594694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.594694Z digest=sha256:e1afc630399f711f48adc7081f9bde074397d59bf3be713b10dd348eee07113e

Observation 45b2f3a3-8c62-4438-b0d9-28ac5ba784e9 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Visual Agentic Reinforcement Fine-Tuning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.745789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.745789Z digest=sha256:19bd5e9539e029fd1f86bc4d345ae1d419b89fab2ba0295e5a30d8118bd21da6

Observation 28131c6f-787c-48c0-993b-cf351e59a231 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Visual Agentic Reinforcement Fine-Tuning ToolRL: Reward is All Tool Learning Needs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.872210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.872210Z digest=sha256:833a33eb6994571156853e02dc058f1096ce4ded67f11ba12bff396b9b3c3993

Observation 552cfcd0-d140-4f3c-927d-4e9fc3ff7cff · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Visual Agentic Reinforcement Fine-Tuning Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.987599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:29.987599Z digest=sha256:b0f4ad9404a7f8b2e9761e03f28d4ecddd6769bd235a72fc6b69408a1a980047

Observation d762b09c-88dd-4453-afd5-5fa0a8d0630a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Visual Agentic Reinforcement Fine-Tuning Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.145090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:30.145090Z digest=sha256:905f0db34b35945962f413dfe7e71bdfd836ff029d45796609370194eb5c5138

Observation 66484908-b300-4c1e-b938-d291a55cede8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Visual Agentic Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.233899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:30.233899Z digest=sha256:7c01869cd4369bf1ff4699083364770d41b9a02cc257eb63572ef790bd38e433

Observation c15a9e0d-c152-409d-b51a-d801dbb5b727 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Visual Agentic Reinforcement Fine-Tuning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.403929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:30.403929Z digest=sha256:88ebe817ffa6031fbfc158677ffa7e9ab08f03acc107856f0511900fe10865e8

Observation 5fe74b72-402f-4ed7-a260-a48087a6a253 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Visual Agentic Reinforcement Fine-Tuning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.574751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:30.574751Z digest=sha256:c3a05be3fcafce877ff4a96da97f984b4b48f1f98fb0ad98038731f3d43943f4

Observation b9b41e95-0d47-481f-b6e2-c4ee79f0e13e · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

Visual Agentic Reinforcement Fine-Tuning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.760995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:30.760995Z digest=sha256:2e9edafa36e960ef4c1314e56cd8a4ca3d640a5801926a17cdf5326344bab9be

Observation 660aa28c-4602-46f5-8cd9-c0fb6858fbf3 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Visual Agentic Reinforcement Fine-Tuning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.079097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.079097Z digest=sha256:5f27e39802555db25e9b8425e45cc4731eb8069d53c32b597b1b091dba7192bf

Observation 7991defb-546e-4abb-b9e3-28a9c4270a6c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Agentic Reinforcement Fine-Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.186825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.186825Z digest=sha256:11877904ec975df85628a09fdb0bbc6e2bcec67eaff3de181686a777e4ff00b0

Observation 2ce2bc79-6281-4e92-841f-cf235de462ee · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Visual Agentic Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.288162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.288162Z digest=sha256:cb080ff987ca8cafa446be439b11a6afacc5965700952f335da602b304dfb71f

Observation a40dff6a-fe12-4059-849e-2e610f03f36c · outbound

This paper cites Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554, 2022.

Visual Agentic Reinforcement Fine-Tuning Musique: Multihop questions via single-hop question composition.Transactions of the Association for Computational Linguistics, 10:539–554, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.477496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.477496Z digest=sha256:57873921f4ff310e5bc2e41daabfc85f42173b9e737f7cb02dfdc5b8952732af

Observation 9c1c7d64-fd79-40a6-a2bd-b1b8d8d2bf3f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Visual Agentic Reinforcement Fine-Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.628464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.628464Z digest=sha256:ba4d0c9aea8889527e09c8f1a6bbe51264d1d8302d6c239f776b2a7cfaff71d4

Observation 4fa59e86-7134-406c-a9d6-16d7f0b619b9 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Visual Agentic Reinforcement Fine-Tuning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.866291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.866291Z digest=sha256:d48d222ce479aacd454dcecf53d6b0f21f32914eec70819a891c67a8d270b606

Observation 8599fc09-f5f1-400f-9b79-f8c25aa03f61 · outbound

This paper cites Detecting texts of arbitrary orientations in natural images.

Visual Agentic Reinforcement Fine-Tuning Detecting texts of arbitrary orientations in natural images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.815225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:32.062592Z digest=sha256:94f8d41d41ec129bf226716604b4c143f40e68010b3f0d62333d2f998f5d56dc

Observation 6c40227d-ec06-4d68-8082-0a6ea4305e67 · outbound

This paper cites RlHF-V: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Visual Agentic Reinforcement Fine-Tuning RlHF-V: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.659736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:32.246782Z digest=sha256:d947de7449c3aab8622a8ab6779448d96a928d2bb62ff63457111d3765ddd59f

Observation 2e6cd8f5-b10e-4e77-ab84-ae3dcf275d60 · outbound

This paper cites RLAIF-V: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

Visual Agentic Reinforcement Fine-Tuning RLAIF-V: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.400015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.400015Z digest=sha256:18741bb6e4bd19e1050febed13440b490f36236c7c61f402ca1e4e9a507fa7b4

Observation af863026-8903-46e6-adf7-f264d3c4b139 · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Visual Agentic Reinforcement Fine-Tuning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.512442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.512442Z digest=sha256:116745c1b4d506b923ada3621ace222c77a44f67d70252b9e57f77e66998f64a

Observation bce2c218-ac62-4c3b-9be2-da4876f22f2b · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Visual Agentic Reinforcement Fine-Tuning InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.658392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.658392Z digest=sha256:b4a40a65646fd587be75a94b96e0e2a6540e2c66484c55ae53c9ad090a14c818

Observation f62d10de-d64c-42de-b2d7-77681a765934 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Visual Agentic Reinforcement Fine-Tuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.736707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.736707Z digest=sha256:4d7114de43bb4a3399500f9c0b40d7649fdf4712d8b3b2b0e16ea1d32ad99b17

Observation e1735659-6f1a-490c-ad60-14d9df918b7a · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Visual Agentic Reinforcement Fine-Tuning Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.811587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.811587Z digest=sha256:7e3677c1850d520a22a9eb18b778e449f9648f55d3c15f973f4a3a4584f706dd

Observation 7201fc3b-4375-46ef-8839-c0f51fdc0c73 · outbound

This paper cites an unresolved cited work.

Visual Agentic Reinforcement Fine-Tuning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.441995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:32.887357Z digest=sha256:a2c4f65fc8a8e107eee36fbe4d1bdc71086658d35c41ae09817e8f80a3841f7a

Observation 782efb4f-3f10-48d7-af85-c042c9b8637b · outbound

This paper cites the image.

Visual Agentic Reinforcement Fine-Tuning the image

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.253285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:32.983752Z digest=sha256:7123dd8ffc043731618f3badff54e97e9e78ae1609c023630510af356cf1796f

Observation 94189a73-5fb7-42d9-b570-b5eb87aa8083 · outbound

This paper cites an unresolved cited work.

Visual Agentic Reinforcement Fine-Tuning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:34.082769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:33.046243Z digest=sha256:93eb85a8e77d41fab7220bf84d2fd7cc722afefc757af145ac24e3d2cdb443bb

Observation 7ebc7938-f7eb-42c3-86aa-fb1ab16214ef · outbound

This paper cites an unresolved cited work.

Visual Agentic Reinforcement Fine-Tuning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:33.892630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:33.132888Z digest=sha256:fb6182fc3a8e4ee0f276168e08261ce51a446f1b03c451cdd150954c7013b275

Observation b4ca1758-b303-4889-ac1a-bcfa784645eb · outbound

This paper cites Can You Find Vermeer's Milkmaid at the Rijksmusem?.

Visual Agentic Reinforcement Fine-Tuning Can You Find Vermeer's Milkmaid at the Rijksmusem?

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.740334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:42:33.223922Z digest=sha256:b9801ad918070b059cffb5b300c16e90be14356eeeb400d8a0fa75e8315dcc11

Pith citing papers

Observation c725d446-5861-48db-920f-891c115909de · inbound

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning cites this paper.

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning Visual Agentic Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:04:19.981518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:04:19.981518Z digest=sha256:318d6935a69570056055b793476e0c13b72eda3775cb399dd2ff28908778e360

Observation cb7f07b1-1c5e-410c-aafa-bf50b790308d · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback Visual Agentic Reinforcement Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:31.664819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:31.664819Z digest=sha256:117a1243100612578a71a87180d03492f040a42d6b8d4572d2b9e959b335704f

Observation 9a1e24f4-4d32-4e6b-a632-3a1e3806b222 · inbound

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation cites this paper.

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation Visual Agentic Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T22:51:58.986937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:51:58.986937Z digest=sha256:96d7f9cac39d19c69d44439d93456d2ec333ac5e1f13909e8c0802b1e68765dd

Observation fc79799a-110e-4fd8-acd1-67023cab6372 · inbound

Omnidirectional Spatial Modeling from Correlated Panoramas cites this paper.

Omnidirectional Spatial Modeling from Correlated Panoramas Visual Agentic Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:37.737420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:37.737420Z digest=sha256:dfd62435c057240df00c0ce10ca9705f9b4bdbb16e63ad5fb396bd69fea411d6

Observation 033251b8-d0f4-498b-abd8-e5eb35679d80 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model Visual Agentic Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.373295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:8de1c557d1c98df69dd436ca194704721c231976c38ee4033de7f3b8d63a72ed

Observation af55a94a-d81b-49c0-9995-b46ac10f79fc · inbound

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection cites this paper.

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection Visual Agentic Reinforcement Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:41:17.900369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T21:39:46.031368Z digest=sha256:99bba248ab046574ab78e50172daa4e2a10f20f61a832bc3ee84ee66d77be484

Observation 2525a52f-b6e8-40ab-afff-3cafdedbf85a · inbound

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection cites this paper.

Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection Visual Agentic Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T15:38:29.010684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:38:29.010684Z digest=sha256:55e879b084c536e5e195ce95271b82f308e5e999c3d722e5196bbb36f00aeccb

Observation db3d9f0f-b51b-4689-8a98-9c83b448bdbe · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding Visual Agentic Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:08:12.704592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:6ae7ad6c217e4b7e80ad4bca266dc3f972cf8db2398ce98b4ab168044468e867

Observation f0b57e88-2868-4f54-b756-d5bfd82a6cf8 · inbound

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning cites this paper.

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:03.346715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:45:21.088097Z digest=sha256:aa1a305465a60650b85b119de3c95b138f2896a95179de1f13340b71809e1005

Observation 323e38f1-cf64-4873-8596-f0686381a9a1 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Visual Agentic Reinforcement Fine-Tuning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:10:26.341662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:e15eb599c1f6be29e80196128f67388aa9b9bbaf8c2a98656d02062b95c73de2

Observation d9d59b1f-9f42-4f44-a64c-754c24f77a8b · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Visual Agentic Reinforcement Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:22.349123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:22.349123Z digest=sha256:7f94b18b1b1c7a77d09eb98aecf1eeb85eda9a446104b791325edc01a3e0ebbd

Observation f541607c-fd75-4ef7-a8e9-acc00711daba · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Visual Agentic Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.151356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:e34189da679d3d93ca03d727bf3dcd86512df3f7fa935c10e430bad49709a3cc

Observation e2ea8c18-56ec-4367-b21d-c097e4887d49 · inbound

Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization cites this paper.

Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization Visual Agentic Reinforcement Fine-Tuning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:07.166393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T03:14:04.535856Z digest=sha256:9874fdf1f4764748e0b5d6496dd94d067e4e6b874606fd7d4342e4be179eb671

Observation dd30d3a7-b8fa-48d7-a4c5-4859b3c4b05f · inbound

Supermassive Black Hole Winds in X-rays: SUBWAYS IV. Tracing Radio Emission and Unveiling the Role of Winds cites this paper.

Supermassive Black Hole Winds in X-rays: SUBWAYS IV. Tracing Radio Emission and Unveiling the Role of Winds Visual Agentic Reinforcement Fine-Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T18:39:55.906516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:39:55.906516Z digest=sha256:09b60e625609d0349bb431f810a9c6ae3c389e13d383d07f6d66e28c9f3cc7a7

Observation 0538cd5a-9582-484b-99b4-99161069ed6e · inbound

S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images cites this paper.

S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images Visual Agentic Reinforcement Fine-Tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:34:07.576598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T22:26:03.080281Z digest=sha256:c7761ab5dc779d6a943835da446fd4df14ae52568bda2eedfdcd7399545ab5ec

Observation 6c77e42a-913e-4405-98e4-2bfc083b8fdb · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.959109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:5125ea31a45bdbceb0d30c708a3aaa235a3b90045cd099a8cdd1db096f254ee2

Observation 3a11aa03-4cac-476a-bbe5-0930825189ef · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:28.537742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:58:21.232058Z digest=sha256:6b3634e7dd1ecd90e714d4d0b5afe54dd4789786beed675c078d9ea935f896ce

Observation ebb3ff76-1be6-49dd-b1f3-dad5b93d9c85 · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:50.999336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T23:23:14.695568Z digest=sha256:81ef8b987a11d1f0c99bdc8b636fe032b55dcea99126c5c499e7db35569cf73f

Observation 334ef447-a13b-44cc-bf05-0d65216b2be4 · inbound

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning cites this paper.

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:34:05.177105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T08:33:24.285225Z digest=sha256:f1547b04bda52b0c6f72a058c381b8677ae91757068d1ce412a0d83d07f53eed

Observation bdd92be5-3bf5-4e97-be2e-a068244c8510 · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models Visual Agentic Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:48:53.501567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:02a1d5fa2ce1844112c09431c9ce43fdb9a2760386d84ffdbe4909c906bb189d

Observation 65fa70b1-12b0-4109-a82a-808d652b119b · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Visual Agentic Reinforcement Fine-Tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.344041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:405891d6e4198db42060ad4ad8cbb2470fb7b58be9319aaa366327f88b73dfb6

Observation dd367bf8-6560-4f72-be9f-de5dd70f0b21 · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents Visual Agentic Reinforcement Fine-Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.757158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:968637f26e0d06ee57fbd02a93dcbe2da8fa36b36ad260189330f03c5aefca9f

Observation b9c26ddd-41bb-49fc-b1fe-b9f03bffa7a2 · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Visual Agentic Reinforcement Fine-Tuning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:44.161925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:1e5593b6e1259fc2ec7699765baebaf447c401fd26eeed6cec454017f6bab02e

Observation 316a23a1-8827-4d56-ae2d-b7e8349b24dd · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Visual Agentic Reinforcement Fine-Tuning

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.099651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:137dae52525a52d6f7aec60e28158d91bb7c7aa8e1e8f407e7231b7a4a15f4f5

Observation 73392426-f97a-43ea-962c-b5d9e7e16716 · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Visual Agentic Reinforcement Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:49880dac99c2e9de04e00c91e57004b5763a7174f3cc531841c6076e031af67b

Observation 656c0542-960e-4728-9f4c-95f14bf59551 · inbound

InSight-doc: Agentic Visual Perception for Long-Document Understanding cites this paper.

InSight-doc: Agentic Visual Perception for Long-Document Understanding Visual Agentic Reinforcement Fine-Tuning

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:53.071368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:43:53.071368Z digest=sha256:334a83942e855580650d5a8a3d8423ab4d352589797382a6e67be27d5e833836