Pith. sign in

Paper Citation Record · LEDGER

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04726 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.614624Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a482709-b5dd-4cb3-a056-9dfd3d727272 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.554226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.554226Z digest=sha256:a0dc2549bc9dedd76374e47c59c33d5015e6fb8066b04212d026e64fdd4767e5

Observation db58b1e2-504e-4452-9dae-b83961accf95 · outbound

This paper cites Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.563621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.563621Z digest=sha256:cfa47c31ffa6024b2607ccb958f9d843eb88771660dd8ed0c4bcb812cc735974

Observation 9133902b-5ab7-40d2-9911-fdede827cc13 · outbound

This paper cites Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.567634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.567634Z digest=sha256:3ebb15d541be2f6ae9df64a5e0e155bed6968cc3492f0bee7cbfc4deb055f191

Observation 073a3ff1-8751-4f9b-9003-9930983f4ffa · outbound

This paper cites arXiv preprint arXiv:2601.21821.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2601.21821

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.572158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.572158Z digest=sha256:9f1c7cb39372886c2f0baf1dcef30d8ea4a54ad52006622a72cb520e27fc3384

Observation 5082ee0b-8a44-460d-9ebe-fc329006bb18 · outbound

This paper cites VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.576081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.576081Z digest=sha256:fa4806d18b3ae57467ac038bf7d29536c0760b80cbdda66be9db2b6b8483525b

Observation e4571943-360e-4046-9080-7be6d136e523 · outbound

This paper cites We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.584703Z digest=sha256:053f2d7cd4cc004563b84bd570a0caffbc9b44fd33868a8e1564fd7eebe47ad6

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:703e9384b493d660d54cb82f7a303ad16599233464f1f9972841417e9046bddb

Observation fcdee749-e364-4976-8046-a77cfebff434 · outbound

This paper cites Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.602554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.602554Z digest=sha256:4d4ea65f6bf96e0b41ee59b92b4a683c19922149c9677fe115bb20ecd7a72077

Observation 62e59df4-17db-4dd9-b4fb-0c5a31f7e731 · outbound

This paper cites Group Sequence Policy Optimization.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.606491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.606491Z digest=sha256:100b4b4357f6918b096d673014fe2482587c8859bbaa24ab027ef657d39e5937

Observation 393aa8f8-c4fc-49da-9c84-44ecef1a648a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.610661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.610661Z digest=sha256:7f53c6e92c19ccbf8543d37186bcd9f6716832387148bdcc87b70088dcac8c2b

Observation fa9de672-90ac-483d-bee0-3e604dfe204a · outbound

This paper cites InInternational Conference on Learning Representations.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InInternational Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:02:44.323648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:02:43.614624Z digest=sha256:f7608af33217723c617ed1d1e967c2def63ce25797cc71a5c78328bd7fd65b41

Observation 1780b0e9-a7ba-4a10-923c-93daf34ea924 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.588950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.588950Z digest=sha256:73056906ae186c0ea525d6a85ca7ec22b78332e067e86fa432e44fab332e478c

Observation f4a02599-d57b-4a32-83b1-4d2d1094bfc0 · outbound

This paper cites Kimi-VL Technical Report.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Kimi-VL Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.558955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.558955Z digest=sha256:f597b014574aed6f63a4fae635d452bf212d8ee93f75e332ab834a8185078715

Observation 7f726666-04ca-4bbe-834b-124876c1cfb1 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.580389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.580389Z digest=sha256:a7cb162d0b5ccfd322bac4e325267e849dfc9af74f970793c6460bc432e39f1b

Observation 7b70000d-798d-4848-b40c-80934cbca12a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.593744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.593744Z digest=sha256:34c28d0cfe36d769850cdd40f607b82b7875f4c516988c48eeff5a7e4573361c

Observation 43834796-eb56-4b6b-85e4-d0ba82abc155 · outbound

This paper cites Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.545297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.545297Z digest=sha256:8ea87874f56f1247efe880fe5c889025793c0e7b1edcfe507f98cb63166350cf

Observation 7e1ffec3-0e69-4bff-89f1-4d1b2d5f4022 · outbound

This paper cites arXiv preprint arXiv:2602.09483.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2602.09483

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.549873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.549873Z digest=sha256:c920f3cc7849a5e6faf7040d067274b3809e2337b7c2e262ffe5e489c4c9946d

Pith citing papers

No inbound Pith citation observations are available.