Pith. sign in

Paper Citation Record · LEDGER

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04726 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.614624Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a482709-b5dd-4cb3-a056-9dfd3d727272 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.554226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.554226Z digest=sha256:8f10f51f582486559e9f0fc0e79895804bd3d3451d88164bfb5d679025fee0ab

Observation db58b1e2-504e-4452-9dae-b83961accf95 · outbound

This paper cites Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.563621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.563621Z digest=sha256:ddaf166f40b602aebbb064bf1ca57b5cf943d3c96619ca73a0b9214e75c7afc4

Observation 9133902b-5ab7-40d2-9911-fdede827cc13 · outbound

This paper cites Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.567634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.567634Z digest=sha256:623d1012732b32185e561071418f5afc9ad1f329590f9d20bc123a53aede185e

Observation 073a3ff1-8751-4f9b-9003-9930983f4ffa · outbound

This paper cites arXiv preprint arXiv:2601.21821.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2601.21821

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.572158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.572158Z digest=sha256:267b4fd640883628cf2ff8e615cf83d2147ba16b51197aeb7a4b71ec4b1d2eb7

Observation 5082ee0b-8a44-460d-9ebe-fc329006bb18 · outbound

This paper cites VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.576081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.576081Z digest=sha256:c3be34de4016476952f57e76647fe376a501e79458c29b0546d1c341f33fa26f

Observation e4571943-360e-4046-9080-7be6d136e523 · outbound

This paper cites We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.584703Z digest=sha256:c871b21a620f2b2001b459db72b572122e5a24138ad52c5b7a7170dd444c1980

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:7e5d0a2be6ac4b4596028ebe81bc66c9d69162359db889e8f77035d6d5ae5e5d

Observation fcdee749-e364-4976-8046-a77cfebff434 · outbound

This paper cites Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.602554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.602554Z digest=sha256:4540ea4105c1f3cd05412a5396a62260d4f1648a11e8708a66584d659275d021

Observation 62e59df4-17db-4dd9-b4fb-0c5a31f7e731 · outbound

This paper cites Group Sequence Policy Optimization.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.606491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.606491Z digest=sha256:c68bb12762cd0c5123bbeca4552246b567b5aa32e63625de17861c5438fbfdf2

Observation 393aa8f8-c4fc-49da-9c84-44ecef1a648a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.610661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.610661Z digest=sha256:25d12edb9cf3181128e71d6d9abd55107c3b3f26632fec003442137bf3489420

Observation fa9de672-90ac-483d-bee0-3e604dfe204a · outbound

This paper cites InInternational Conference on Learning Representations.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InInternational Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:02:44.323648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:02:43.614624Z digest=sha256:87098d81853e3d7b5e90623d68a73a7ab2716cfdf7f832e8f068d5e7fe3d30cc

Observation 1780b0e9-a7ba-4a10-923c-93daf34ea924 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.588950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.588950Z digest=sha256:4304fdd1e58808fbb0bfe467fdd45e3d7e749e9bd2d700a0672786cddb036667

Observation f4a02599-d57b-4a32-83b1-4d2d1094bfc0 · outbound

This paper cites Kimi-VL Technical Report.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Kimi-VL Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.558955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.558955Z digest=sha256:448124b624b85d74c01565ddb8c98ef6ef500c94296332ddfd9142f66a7d515e

Observation 7f726666-04ca-4bbe-834b-124876c1cfb1 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.580389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.580389Z digest=sha256:d48c5eb5736b5ff946a5c3da52eb5c22224a12c2f2f1fffb22861b17e18e373c

Observation 7b70000d-798d-4848-b40c-80934cbca12a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.593744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.593744Z digest=sha256:015c7e77999c3b2a967992fc4512aa9ebaf4abe590e15033569f27f2caefcab0

Observation 43834796-eb56-4b6b-85e4-d0ba82abc155 · outbound

This paper cites Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.545297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.545297Z digest=sha256:2c4b8985199cc1d9bc8105b0ea5b1787e767833e48e157c22a893a13ff3c522f

Observation 7e1ffec3-0e69-4bff-89f1-4d1b2d5f4022 · outbound

This paper cites arXiv preprint arXiv:2602.09483.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2602.09483

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.549873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.549873Z digest=sha256:c84799c512d0a7783cc1087a0dbdad57138049297bd145598a3a743c564e984f

Pith citing papers

No inbound Pith citation observations are available.