Pith. sign in

Paper Citation Record · LEDGER

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

As of 20 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.27902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27902 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:22:13.476296Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0816c7ec-d01c-49ea-8cfb-ddf9beee6657 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:08.635594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:08.635594Z digest=sha256:9530bc2ff32de558fbd8ed5334ff4dd6cc98696aaa1b2f19f3e3dd773cb7cb05

Observation 5a77cc0b-c772-4a29-a7f9-26dd1843bb3a · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:08.713601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:08.713601Z digest=sha256:f753bca1299bffb013982b993439553355ef78eb72d09f7225c52d14e2d04fa2

Observation a229d704-96b6-43c2-b753-376b5dcafc70 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:08.842461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:08.842461Z digest=sha256:f271f252dbcd779b72f86b93dfceb94a046387bedfdc24bd272d9ccca6d3e0e5

Observation fa1f0bf0-74c9-4c1f-8655-67520fdcdc7f · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.059622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.059622Z digest=sha256:f4114d978fd888f9a187143d92e53a825205118d80be7ff4faa2c538779c7ca6

Observation 8369cad0-077a-4848-95b2-d3782c02baa2 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.280058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.280058Z digest=sha256:717156e7a657e29cde7f8bde6838069be1bbfa61d7139b6101b918e18cc40823

Observation 23a99315-b9f8-4244-a5be-51064416e3fb · outbound

This paper cites PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.389422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.389422Z digest=sha256:d500a5ec43de2721749200821d93290abffa42639fa63733c8c3f74f0a1710d2

Observation 9cb41eaa-c793-44dd-8c3d-7902e66fb30b · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.543692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.543692Z digest=sha256:38a3a30b8a067491734d782117b59a587230e21751f19dce446afa7f2c584726

Observation 747e83ae-07e7-4aac-89f3-25c83b3a7ebd · outbound

This paper cites PUMA: Empowering Unified MLLM with Multi-granular Visual Generation.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.894169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.894169Z digest=sha256:d508887636f88aea815453d8212fcbbf5718bff1e7e070e2c4861dc886d0ba36

Observation 899b0b9e-048b-4f18-988a-473f7c7cb276 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-04T03:22:10.046145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.046145Z digest=sha256:c26a528fdabf03151ad19c307983856fa096612c1dc51e0a0b1908114cc2d58e

Observation 7bff85b2-51dd-4791-8893-504317f36c19 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.143840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.143840Z digest=sha256:f45edaaddd8fbf013a1aaf01064c7a51a862e0faccb80cdd47b27b402fe1836a

Observation 956929f4-1172-4943-991c-7eb6bc675c53 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.220936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.220936Z digest=sha256:129a361ef937542fd07e454a994815fae22528a6a7c7e36843e793b1f38e18de

Observation 38f03fc6-7b3b-4964-b095-48ca2859c123 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.373626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.373626Z digest=sha256:3b45832f782581d50931603f6fa80b3f1294be8e44ff50386cec80afd5c1c173

Observation a394859d-8836-4532-ae4a-967c79f609d2 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.538597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.538597Z digest=sha256:5049add91eb8ddfc3d51a24d8dda6dde6f658a8c7e79015cff2686ba573bda03

Observation 065f4784-ec0d-436c-bf5e-6df6887fa878 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.699451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.699451Z digest=sha256:cdc8f1182cbe59d57b8c9c503f98d2211770e6d5b0caf7c759e909f276625132

Observation 94567e98-9d72-4349-9d88-c28abc007cd5 · outbound

This paper cites 2026.TRL: Transformers Reinforcement Learning.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting 2026.TRL: Transformers Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:10.864868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:10.864868Z digest=sha256:33f8ac149f07adb62062d2bf48528984eb4efa8346c1205d556a22d9b1e15fbb

Observation 2c53b955-0b38-44ef-8b5e-3de8cc67ea0c · outbound

This paper cites Ghosh, Andrew D.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Ghosh, Andrew D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.027627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.027627Z digest=sha256:16265b32a5cf5b2bee2f3c93962d1f01bf4e24254b467fd933346cd0dc973502

Observation 3a5a217e-6803-48ea-a9ba-d8e7c45f285a · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.134453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.134453Z digest=sha256:484bd039cc826b1d4f937927cea9e077aec8dfd75205c5b45df22d1176ae910a

Observation 933c0ff1-b8ef-4674-bd5b-d633f85e6fd8 · outbound

This paper cites Jaakkola.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Jaakkola

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.250370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.250370Z digest=sha256:0a03f4706e86959f544e81fe29e0e4ce58a815e318de668c431a127a462aaf28

Observation 79aa3288-4732-4557-99ca-3925bac3ef6e · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.501927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.501927Z digest=sha256:4c89c4f154ba7f264830a89e36e74b06810dd4ca376591d345567cc2525f5d60

Observation 79f9e756-1d28-43cb-b9ae-7eef0ecc091f · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-04T03:22:11.616695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.616695Z digest=sha256:ea36cdb752488241c274d67b719ccb5656e9cb6936b38fca73e2f31a022ce747

Observation a1448a85-1c6b-485e-8eef-cb09c02c8327 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.728750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.728750Z digest=sha256:fb112fa4591a3b7c07b77aed0196cc62f6cd2fd870756188ef4f569150a97d48

Observation 5a995e49-2306-42d1-a581-f1b29c332eec · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.841840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.841840Z digest=sha256:f8e55a0965b4fff3a8e1554f34e97c91d9ce2b2f9c029eacf39ddbca43112d44

Observation d337fb5e-2d69-46fa-8ac6-f7ebe73e68a2 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:11.954241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:11.954241Z digest=sha256:dc500b72b201e0efd356514d5b468213758eb9b0737ca91ec3c4f702b12dcf68

Observation 21380f7e-bc39-4b87-8978-e52aca747018 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.068440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.068440Z digest=sha256:047e61a9079d3c748e7f325e8f0a464a800d2f247254dfd0ae793a21f4b8a2de

Observation bae44e8e-43a2-4654-aa1a-deba1ab87aec · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.168048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.168048Z digest=sha256:bb728bd5e227b79d6e87b879f2a4b3260ce684a5f86dd32615a47272d7c38e01

Observation 9d5cd095-d0e9-48e6-a1e3-b29ad38f1891 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.263797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.263797Z digest=sha256:df730505425ec1af33c519d94abf63f7d426c81c2b42e3892b3279a2a336ef67

Observation 42fa0ded-ec24-40af-8086-3e293c9637bb · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.427191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.427191Z digest=sha256:0f2308c6064d673d2825c5b6ce90e68a1fc4d68e504928b3143a0a3e15a7e601

Observation 85cf3a1e-6ba3-499d-b882-f56d83231148 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.537300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.537300Z digest=sha256:ed3b7f470ae5c67576e34516de292c0e0fd15923d15d3048caad21d4e469833c

Observation 10a34d74-fcd8-4d3b-85cf-b45a43142cbe · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.693033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.693033Z digest=sha256:16699c1896b7f328da1757d0dab7f312e29629df9b87ae9223580dd206e3cf38

Observation 83c3516d-1d32-430c-8e2c-5f2bfa47dc9f · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.786867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.786867Z digest=sha256:7711a96c11d6fd11925248d619cf58e13d5d570677aab9f71c7016d7f26927ef

Observation 4c492d79-1808-4bc5-abbe-c53d98a34e78 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:12.894951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:12.894951Z digest=sha256:695429ecd2d800d0976ade0b32a57cec39e950d817ff8a94f8ffef8595e78de2

Observation cbb9e6e1-fcd6-4cef-9143-45435a93bd73 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.000847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.000847Z digest=sha256:90275c1fafb6f7dcbcb88ac076549a263bdfa5fa5874b406c34eb2866ce64a73

Observation c0dff2dc-388e-4b7c-8144-0c8cdbe15379 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.017766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.017766Z digest=sha256:24a27409cb7f7647b3dbe2abd2111d6f4e34c98524992c11de1aeefc083c1d93

Observation 75d4acb4-98e5-49f7-bcb7-02e28050e88b · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.072027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.072027Z digest=sha256:6ee76354b073adbe677036edf3de0165f5db3285714492c1d02441a0e15021de

Observation b38b1434-8575-4419-a8a5-90cc621c1140 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.181141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.181141Z digest=sha256:d333fbfdd0c4dd45c0a9f44472bea392974ea15ba1affee5e15f27ccec21f15f

Observation 10b4d30e-c1fc-4823-a438-3c305649a415 · outbound

This paper cites Girshick, and Jian Sun.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Girshick, and Jian Sun

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.277872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.277872Z digest=sha256:3d567348248323dbc1f5fa7458b65207f76212127e0ef46a3979da53d8f3d02e

Observation c6bbd66f-604a-4e21-b07a-fff5f7b8c647 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.389549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.389549Z digest=sha256:5805af1021181be106e55418966f82c808aae85404e04c1450c4f357e04716e6

Observation 7b862218-86d7-4910-baed-d42515f6dcaa · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.403706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.403706Z digest=sha256:ea416dc31ffef5f7547f0be3e1cc9a4c0e659a162ab9a8d12040fd0259fba87a

Observation ac5c3528-5ec5-4cd3-89cd-1b78bb424377 · outbound

This paper cites Chen, Shuicheng Yan, Xulei Yang, and Xun Xu.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Chen, Shuicheng Yan, Xulei Yang, and Xun Xu

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.407724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.407724Z digest=sha256:6347e836dcc039be475fe54c41417e044b2457512f53b8543c6c2ac9ea4fb09f

Observation d1afe8d0-5bae-4443-bc4f-687789257ef8 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.412202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.412202Z digest=sha256:5ee9e74a72fc01107c53124cd7154717c51cf3f12e1a296b0784f85cfc26ebab

Observation 1fde44fd-7f07-4abf-bd90-5dfa852584f6 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Kimi K2.5: Visual Agentic Intelligence

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.415678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.415678Z digest=sha256:8407b46d73d168e9c30e32a18867907aa462a99d6e77f7b905572d323fbd9a43

Observation 33e4f02a-0f65-4081-9d92-dd2377f1b65b · outbound

This paper cites Qwen3-VL Technical Report.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Qwen3-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.419357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.419357Z digest=sha256:79dec52d191e550687632be0627928b2c9297e68f61bc22457ab3088e12738e1

Observation d70f1bb8-8e51-416c-92c9-de94debc53c6 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.423086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.423086Z digest=sha256:1e4b8378059341ca5ada3d29d42d6e8ea9c713692d6487494cba1d5b70a23527

Observation 8c73b1e9-9c59-4452-8339-0d06e63aa589 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.426661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.426661Z digest=sha256:bcbfff6088a8db8a4f90a9942606ea4b3266b5e807217f06a1d4db5dd1aa7689

Observation 60e668d0-c6bf-4620-bb4b-cba12f283345 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.433253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.433253Z digest=sha256:bf42dfb4504bf90428a33811432e284214c5f7b3747dfb72b18c7227d81e41b2

Observation b7548687-e613-43d7-adb0-815d124dbe61 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.440881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.440881Z digest=sha256:b7ed3500cf9b2cf92136fada4715d5f130e48aebff09c303ff059930b871fe38

Observation 8c313408-e7ba-46a0-94e9-5759e1870b06 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.445389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.445389Z digest=sha256:c772cbfd0fc558241314377a9f34120814bf27cdcd4d762de18d991b7431c9b8

Observation f1886b4b-a2c1-4fbb-8ee1-336dbfdb4932 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.449004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.449004Z digest=sha256:8b0742de5a60848f3aecece0bc41184511bf68a7d87fee0c833829b64e4d8e22

Observation 3d6fe268-aa9d-48df-95a5-7e93752eb3d3 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-04T03:22:13.452412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.452412Z digest=sha256:8dc9c9d1361092de57349c5d1e9656c984620389f458a149afa8ab61a3d78912

Observation c527fc35-e40e-4e7d-a161-a1d1fd544c97 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.455874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.455874Z digest=sha256:ca93434ccf4a37f5dfb514b6eaff02dd5365bef8059a1fb438180faa00e3e9ab

Observation 052d55a3-68b0-45a0-b674-45a41723dd03 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.459160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.459160Z digest=sha256:f60637e5e9d8def99a3d75a50523a142a52e25e160d8bbf3f371abd8d634da96

Observation 254765ab-daf6-475d-bb51-6a7b64df5fc3 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.462358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.462358Z digest=sha256:9aeeb5390a081f49e704cb96420c33d064880dda130880ede3715e0bc2a40c9d

Observation 75047152-3733-4929-a601-e7669280f211 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.465686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.465686Z digest=sha256:59b2a1950633b521071d3f82b676db010952f965c6fb947b40d406b0bf4f5cbd

Observation 23132b0d-f6e9-4a5a-83b7-3f619d0d40cd · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.469127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.469127Z digest=sha256:2eb1f13f608b02204b40527ff1800eef75e906ac73fe1c72fdf77b6902ebba42

Observation 05e5db74-957a-43d0-82da-a75d53475e70 · outbound

This paper cites an unresolved cited work.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.472887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.472887Z digest=sha256:5e9a7bd1fab2bef9b96ecc0765fbaa16edbc4c02c7a65bd5b4b2ce505e671ce7

Observation bbb46ea2-6612-40c0-aefb-4b35304dcfbc · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.476296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.476296Z digest=sha256:a8cc820ab96ecf40e8e48e166e971d4cf07d21db889c05b901661c21d07f9f57

Observation 5dad74da-3e23-4e3d-a438-b525d8e2062f · outbound

This paper cites doi:10.1609/AAAI.V34I07.6896.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting doi:10.1609/AAAI.V34I07.6896

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:13.436757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:13.436757Z digest=sha256:130f46017c29d40666ebfb35f5eee6d45e24a4acd2eeda450c85c8665ec5ab80

Observation 7e46ca2f-3e54-4aad-b826-50a5e1c32bd6 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.139190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.139190Z digest=sha256:f485b53779183b991e3a46b73512d20ec910bc3a308b603a742ecfa5da0bdcb0

Observation 9d31c732-9469-4766-8886-f54686eba420 · outbound

This paper cites arXiv:2510.14528 doi:10.48550/ARXIV.2510.14528.

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting arXiv:2510.14528 doi:10.48550/ARXIV.2510.14528

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T03:22:09.738175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:22:09.738175Z digest=sha256:076ad9d2ba6c9c964e9d82dcd3902b87de1b4bd8d74632268bd28f1e1daf1f5b

Pith citing papers

No inbound Pith citation observations are available.