Pith. sign in

Paper Citation Record · LEDGER

VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.07523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07523 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:27:04.898305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.931811Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4dfd7126-316b-4ee5-8128-e5e9089aa97c · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 249

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.567891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:fd437a80df76215312202346795a746b6a442eae4bc5fd99172cd867c3333c4b

Observation 4b1ebc86-ba6f-4939-980c-8c88ca5f7ee0 · inbound

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning cites this paper.

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:02:41.363985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T11:02:41.335059Z digest=sha256:98a216313e525af54defe58d326d747fe334bd62fba4e91d333cda33d0aaa508

Observation 3e499aa5-a224-47c2-9505-ad6c9122e57a · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.898305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.898305Z digest=sha256:8c9430c273347f8cec6a36cd7cd20cffa6acf648919463c9dee1d1da1c721c33

Observation c9b6772d-36ec-45fe-ae17-068454f17a68 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.147145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:c826909c3366ee380e9e36b5b003ca096728d4d9ab24dad8f0cdbaa6cd7d6597

Observation 70a3c192-1f77-4e97-9cba-f790947f1247 · inbound

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception cites this paper.

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:04.467657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T05:18:57.925925Z digest=sha256:e1b44ef2411951dcb74a48ed66cb39d1738df87fe12a730cc1971f3684832d5a

Observation 4bc6dbf7-08f0-4a29-9c0b-c5531b1054b0 · inbound

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection cites this paper.

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:02.947045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T04:46:34.946640Z digest=sha256:f6f1a2e06424bbb18ff558d07690ddeb96d710911f69179077c796c5d3f397bd

Observation 42e182a2-3342-41fe-9686-0d1cf0474f6d · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.374154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:594c3510ca4424201d249302f5a86a54c43318fa0f32426f88a35b5ea2d2e3f5

Observation 3db26024-7264-4ea9-b29a-68da4f45e0a8 · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.792465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:0f51f022a7627fd28ff8a87c746528d5e29fdfe2e99495bfea994ab1d41a5ff7

Observation 81374035-eb40-4030-a6b4-b3102e2ff0aa · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:39.149779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:fd8946397c320b5bafeb1df73cf82599a02e6710fbd8410feb5dcae0143e7737

Observation d3d2e81d-c8b8-4840-84ac-c698e8848a0c · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 264

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.876971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:6aa45ae8bb5474f6dcb1c00c63415496c6be00cfc221e321be6b24e50dc92c5e

Observation e2f58707-c14a-4dbe-88d9-89caefa8f00a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.690051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:d0e2061224b831dd6561a311c7bbd93094267c673df3ffd4f42373efb9bd23f3

Observation 0cc1e7ea-ecd6-4cb0-899d-f2318a5dabe0 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.933520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:65c109cacf153f7ba0c60d6ed88c8d5248ec8b7e9749b4c20bb4f67eb745a39c