Pith. sign in

Paper Citation Record · LEDGER

VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.07523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07523 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:27:04.898305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.931811Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4dfd7126-316b-4ee5-8128-e5e9089aa97c · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 249

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.567891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:cbff009a865940d94ee58ccd4cb8f298bbc399066a61fb8ce329d5916b8fc876

Observation 4b1ebc86-ba6f-4939-980c-8c88ca5f7ee0 · inbound

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning cites this paper.

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:02:41.363985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:02:41.335059Z digest=sha256:4b27d9b5e508e9ecb2a1015f0a627b45414b4369e6a19022db0d6703d88793ee

Observation 3e499aa5-a224-47c2-9505-ad6c9122e57a · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.898305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.898305Z digest=sha256:2c5fea8f379a82ee6ae8333b09cc167a5fd62a36d803bf1276fcd44f72dd8307

Observation c9b6772d-36ec-45fe-ae17-068454f17a68 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.147145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:6229de33b73f9066ebecf6f759819641e89b81339388ce381758cd55cdc148a7

Observation 70a3c192-1f77-4e97-9cba-f790947f1247 · inbound

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception cites this paper.

CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:04.467657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:18:57.925925Z digest=sha256:65a89a763ce8ca85873a9c369956ae90cd93d28d709d6dd0caca1a0741011fb7

Observation 4bc6dbf7-08f0-4a29-9c0b-c5531b1054b0 · inbound

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection cites this paper.

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:02.947045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:46:34.946640Z digest=sha256:cf76b2ab4f46f0ffcaaa3e6dafb3578d9bd285840e6709a0faf9953522755204

Observation 42e182a2-3342-41fe-9686-0d1cf0474f6d · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.374154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:742e1e3aadafcfd7d201f636431dfd88b3267cb89fc641fedc58b471833dc982

Observation 3db26024-7264-4ea9-b29a-68da4f45e0a8 · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.792465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:7744d42bfced3135085005a36ccbbd1c439ec628fae5a64101d60956cbe5bf74

Observation 81374035-eb40-4030-a6b4-b3102e2ff0aa · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:39.149779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:ec6b4c42f9fbe461a2d24b184792b3b9ba25541c5c0c92d4237577ae5b0b3840

Observation d3d2e81d-c8b8-4840-84ac-c698e8848a0c · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 264

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.876971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:fc8cd03c38de05f5c26423094b0b27815ea46e25b9ea741fcaf4bca656c6f245

Observation e2f58707-c14a-4dbe-88d9-89caefa8f00a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.690051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:6077a46ce4f17966909b7670d058ceba31404314fcda2a4cdeea81beba624587

Observation 0cc1e7ea-ecd6-4cb0-899d-f2318a5dabe0 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.933520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:80867cf03f05863a579e6c54a52dfc5da94e6d3d10dc1243e2486e15212ca6af