Pith. sign in

Paper Citation Record · LEDGER

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2506.04999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04999 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:34:03.546459Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:51:08.895411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19b7fbcd-7a0e-47a9-91ed-980b563b0049 · outbound

This paper cites Word spotting and recognition with embedded attributes.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition with embedded attributes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.370790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.768761Z digest=sha256:3bb085dc802f180c262639e27a96d14000212628af85807a93d759a91e6c758b

Observation 8d125acd-11c2-4932-a0cd-de2d8ac1427c · outbound

This paper cites Integrating scene text and visual appearance for fine-grained image classification.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Integrating scene text and visual appearance for fine-grained image classification

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.355245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.851730Z digest=sha256:47436dd93ccdef7c00f14288a7e1a5aa9a0fa625e777a94bb0f583a51e51aa8b

Observation 725917d1-32b0-4003-9e5d-048aed65f324 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Composed image retrieval using contrastive learning and task-oriented clip-based features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.337860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.023632Z digest=sha256:1ca95f6d6609dad43117f8368563e4a8cee58c101daac4be9702af57073bbe29

Observation 6ab1ddc1-e6d6-446d-bdd9-94d793bc8a99 · outbound

This paper cites The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.322419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.221307Z digest=sha256:ed9c235bc5f76c32e4845665d92b273776e3e06a08aa59b1fc048ace0286b13d

Observation a9ecefec-40c7-452c-aad5-9d1010d67940 · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.306842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.383782Z digest=sha256:f3a4517a45874dbe31bf2698378f1ff2fcd7bee9c0a308002f67e023270d0126

Observation 67c6b5b4-6b3e-42ad-893f-1766c4ca0c3c · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.289434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.567461Z digest=sha256:8184a403acb73af0217d2e58774cdfd4fe9037abda4dec12636fe2c5b7abaca3

Observation 7326a63d-173c-49bd-aa02-dfb3177080ea · outbound

This paper cites K., Gomez, L., Karatzas, D., and Valveny, E.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts K., Gomez, L., Karatzas, D., and Valveny, E

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.271142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.750412Z digest=sha256:04a460bb7a9dbc74bd0e81ab27180128500c8836827becf9ed8d20ce4e9ff995

Observation d77fb228-da44-412e-94e4-46a9daa8bf59 · outbound

This paper cites Lsde: Levenshtein space deep embedding for query-by-string word spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Lsde: Levenshtein space deep embedding for query-by-string word spotting

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.254671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.944990Z digest=sha256:ce4441ab7fd9c07a4a602ae6a6438e4d303dbfe4a235520597c7e9b543d9afab

Observation 967183e2-ec4b-478b-b0b5-0e00341d3fc8 · outbound

This paper cites Single shot scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Single shot scene text retrieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.239639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.079254Z digest=sha256:33383c7c4177a5ef983797245e555237740141d4f3f60e54ea95b242a345683e

Observation 9c1f59e1-cfab-4ed4-b07b-eac0c396bfaa · outbound

This paper cites Synthetic data for text localisation in natural images.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Synthetic data for text localisation in natural images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.224142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.206032Z digest=sha256:521d3adae8e7dbe6ee1a315099fd0efb1654afbd8023a619af11f0ad7afb704b

Observation 7c64ac95-9dbd-4e08-8110-2a4d52a8a483 · outbound

This paper cites Bridging the gap between end-to-end and two-step text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging the gap between end-to-end and two-step text spotting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.206480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.368840Z digest=sha256:762d12405b08d92adc9095ef6eb389763a401d56feab6f1e7cf3fe9a9e7df8a6

Observation 55c12d19-33a6-49b6-895d-5f8a1cf31640 · outbound

This paper cites Reading text in the wild with convolutional neural networks.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Reading text in the wild with convolutional neural networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.052301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.502105Z digest=sha256:33a869009225ff10e7e58ec5e6110f30d6ee2f380979550ad755730a93263807

Observation b8e8a25f-ff91-4dea-a15b-86c46a14cfcd · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.037243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.685119Z digest=sha256:7850ce66c074d57b390611a01bb3442aa6e345a2bdc1e6f2e3cf88e72c43fddf

Observation 8ae8f700-b529-44ad-89e1-e9cc258ebd46 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.021425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.792242Z digest=sha256:048bfe844d1a511abc5ba5ae5da06be59228ec66bd25fea26ee7b5b14eacc528

Observation 09dc5171-96f9-4610-aa75-d5a0d8a44054 · outbound

This paper cites Mask textspotter v3: Segmentation proposal network for robust scene text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Mask textspotter v3: Segmentation proposal network for robust scene text spotting

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.003389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.915356Z digest=sha256:6e3bd17725a32678e4c78a446f92df033377398d57ff3264005820811356b15a

Observation 9bd96913-088a-4be8-9437-307444e8fa51 · outbound

This paper cites Parrot Captions Teach CLIP to Spot Text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Parrot Captions Teach CLIP to Spot Text

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:57.108940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:57.108940Z digest=sha256:6abb90308cc41ab32d6629a3b2bd1cfb0e01924e5d3e7bbb6609105bef8f912d

Observation a0e84eb5-63f5-4cce-abd7-ad5238686dfc · outbound

This paper cites Abcnet: Real-time scene text spotting with adaptive bezier-curve network.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Abcnet: Real-time scene text spotting with adaptive bezier-curve network

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.986619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.275665Z digest=sha256:4f663d05db117fe005ede083b907f7e565c196c71c2b87808793ab3c4ee36cc4

Observation 3394719f-ae53-4136-8686-c37337185e42 · outbound

This paper cites Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.968964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.405128Z digest=sha256:6266c28a13492732888aa10d86a5f8fea4fec5d81941191fb8d9888a6b9383e8

Observation a6f5e436-a131-431e-9444-31d7188beb4e · outbound

This paper cites Visual and semantic guided scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual and semantic guided scene text retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.952098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.566552Z digest=sha256:86304e95444f67f7a59b25b853bed0f9caca5fa50ddb11e2e0bcaeda5aec2474

Observation 1290ac95-3c8d-498c-8fba-e6136ab610f2 · outbound

This paper cites Arbitrary reading order scene text spotter with local semantics guidance.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Arbitrary reading order scene text spotter with local semantics guidance

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.935312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.698449Z digest=sha256:ae76248d077e28ab7cff7852a497c07fc165e81fbff2730d8c83dd607e2c06e5

Observation ca4dca92-f899-4c60-a392-a58de153c75d · outbound

This paper cites TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.917434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.885055Z digest=sha256:45b068e5f92675b7372dda5d6a58d78611e9bbecdebc824fc888664e32c58bf1

Observation 67d2aaff-0a84-42d9-97d3-b142b693894c · outbound

This paper cites F., Gomez, L., and Karatzas, D.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts F., Gomez, L., and Karatzas, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.896657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.018490Z digest=sha256:9d66695ab3815b26bcbea46874e96b48b4d3450f7d69694cda5c5a9a128d5c7c

Observation adf89f9b-e367-484c-a52c-d6b0c15db0ad · outbound

This paper cites Real-time lexicon-free scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Real-time lexicon-free scene text retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.879134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.137410Z digest=sha256:91ca2a4c6e739fecfaad57bef58d8362a41962f838e90e892b136e3eca1f2c62

Observation 9a27e253-a92b-4800-904f-084ab52987a0 · outbound

This paper cites Disentangling visual and written concepts in clip.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Disentangling visual and written concepts in clip

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.861704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.298068Z digest=sha256:8da0e2cdf94a9efe013a5b1b23d77294dcd1ea3806f13aaf7062c403e0e6f6b0

Observation 0fd555bc-2131-47b9-91f9-7effe32c3009 · outbound

This paper cites Word spotting and recognition via a joint deep embedding of image and text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition via a joint deep embedding of image and text

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.845036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.501734Z digest=sha256:8ed92ce1f3740d8c6318766cb2c6d8fae2e0de8a52ec45fbfdaf77686a024d0a

Observation 494d4877-9c19-4f90-8df1-26544445e99c · outbound

This paper cites Image retrieval using textual cues.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Image retrieval using textual cues

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.824192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.633790Z digest=sha256:17c1b1beb5a30b564617db8b7872e5d8b5dcb0e0d0faa9ffd931a062e7dd7b0b

Observation 43fdc69e-fd6a-441f-96d4-c51a7af8dabd · outbound

This paper cites Text perceptron: Towards end-to-end arbitrary-shaped text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text perceptron: Towards end-to-end arbitrary-shaped text spotting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.800334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.756760Z digest=sha256:d4030ddf5a860faccd151eac1bb5690b73f068d85f6b25cfe292b3564f7fddc6

Observation 33c759fe-ff90-40e3-899c-695a9e8c2622 · outbound

This paper cites SEED : Semantics enhanced encoder-decoder framework for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts SEED : Semantics enhanced encoder-decoder framework for scene text recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.777621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.932601Z digest=sha256:cb4a01cf6b119c6e575d516180cb870e4f2e5988921c46e82705aaa4753307d0

Observation 26669bf7-34ac-4dc0-9a7a-bb1f1fda7196 · outbound

This paper cites PIMNet : A parallel, iterative and mimicking network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts PIMNet : A parallel, iterative and mimicking network for scene text recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.757899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.113247Z digest=sha256:3a6a21c2cf9d4b19d8eebae93b5bf060fb9f3b1975d131b6f578d0e61ed1fbd1

Observation c812f105-6f83-4140-90a6-86a2dcbe04b7 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.739445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.311002Z digest=sha256:9feedf51e4234079b062e20040679bc9572bee394abbb18d2fe456941290efd0

Observation 29099d5c-5d3b-49fd-8f8e-2455b3566cfd · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.721235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.490623Z digest=sha256:fa9a4b4e26a447c413f335b25b5a4d0fe874abb3e14aca6374faa47570a6717d

Observation 3193a5a6-dac0-43c3-8005-6a8100dfc4d0 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.700159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.630256Z digest=sha256:476e90237213d27a10665037c12ec653dd5766dc4ab8b1a9827843b9bdf3b58c

Observation d5948cc0-dc9a-4651-a9cd-0024258b43a4 · outbound

This paper cites LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.677264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.737394Z digest=sha256:52cf2c373534a5b00f617c3a162aea7adc4617338f1de7d9a66b587b5a1988e4

Observation 36a0a6c3-4e20-452e-bf5d-7be37736e743 · outbound

This paper cites and Yang, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Yang, S

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.655824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.865196Z digest=sha256:c02c400d92c742216a98eea59152ee75909db04a4eee5ef13838596c1e5b3da2

Observation bb09becc-0361-4559-98c8-9d9091963ae1 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts What does clip know about a red circle? visual prompt engineering for vlms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.633680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.999588Z digest=sha256:4cd02ab0ed72392920e90202f716e93fa5f77ee545c221817cd7d3d799430acd

Observation 3b8b2a52-613f-4fa0-b765-416e8b945def · outbound

This paper cites Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.612274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.150960Z digest=sha256:ae929446695c2a03e7f1b541baba2c6271208af6190a515ce6ed738ba768cbc0

Observation 7582d782-bcc6-4193-a654-25ce888f7cbe · outbound

This paper cites Visual Text Processing: A Comprehensive Review and Unified Evaluation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual Text Processing: A Comprehensive Review and Unified Evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.263138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.263138Z digest=sha256:1f1b4f01f95be59ba816c379d441948339183cfc2c863641d5f9f3d6c000dcc9

Observation 8afb498f-98dd-4c1a-a222-54ea38e52ad3 · outbound

This paper cites Text siamese network for video textual keyframe detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text siamese network for video textual keyframe detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.592910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.423634Z digest=sha256:707baeb17738aaa5567a269edbd1e9ba8d3e43593747bbe9d554ba5e0c91a5a1

Observation 0b702319-8127-42fb-ab6a-b59835d7f895 · outbound

This paper cites ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.550841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.550841Z digest=sha256:02dbb966b8351d8cb32b7d36e1c6f6271386325fa302e0c76f5cfa43dd399e87

Observation 7bab5433-11c0-4b08-a4fa-649066da084a · outbound

This paper cites Alpha-clip: A clip model focusing on wherever you want.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Alpha-clip: A clip model focusing on wherever you want

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.567265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.739489Z digest=sha256:c5092f4755e13381748a288deaee49648f26d15ef8d0f48e67f0beda6f3f7693

Observation 6208d676-89e0-4553-b63e-49921c64e1d4 · outbound

This paper cites Scene text retrieval via joint text detection and similarity learning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Scene text retrieval via joint text detection and similarity learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.550844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.863099Z digest=sha256:dd7ee2d79004584f8b31255b512fce7dabdabf7c6e1da9ea4be7a8dad5c4352f

Observation 98b0e76a-9b63-4201-90b8-5e6fa7dcc40b · outbound

This paper cites and Belongie, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Belongie, S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.533936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.975641Z digest=sha256:3bf3dbf275809fd773219028047c7dd3a054f9a0db4f46e528d49d510ccf57a0

Observation 0230ef5f-da59-4ce7-8441-fa8aa021f513 · outbound

This paper cites TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.509147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.158728Z digest=sha256:85476b8527f808e7969435d02c1eb7fe18cc76e8a63b562141d65811c402b82a

Observation 840cf871-9a2c-4b19-9ba7-38f7781a32d6 · outbound

This paper cites TextBlock : Towards scene text spotting without fine-grained detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlock : Towards scene text spotting without fine-grained detection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.492576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.308373Z digest=sha256:944cce1994e36447b71a2cc158f899efdbf2b52fd41845220e2abacf4f72391c

Observation 05d1d8c1-31d6-442f-aadd-fb08bb350c8a · outbound

This paper cites Visual matching is enough for scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual matching is enough for scene text retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.475436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.467766Z digest=sha256:26e88a60b6163f825cc2ffd40d88a1ae53c0ddfc4df6461058e93c75731b84a9

Observation d868ad56-f596-4837-bcc9-a5ae60c5d046 · outbound

This paper cites and Brun, A.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Brun, A

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.452994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.645696Z digest=sha256:c3b5dd3f482e8a5e2bc0a6a2c3155baa80d8203da3b531335cb724a9984f3aa6

Observation 7787ef5a-3a50-4eac-af92-27cb0176e811 · outbound

This paper cites Bridging semantic gaps for language-supervised semantic segmentation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging semantic gaps for language-supervised semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.434978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.767073Z digest=sha256:b06c17c78f99f4b6170ef8a7740a675062be9c1bcfffe1420e2adf0fd7b15782

Observation f1967491-8899-4f2d-85a9-d16e43bfea0e · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:01.936414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:01.936414Z digest=sha256:0648fc19037a62f832e5adab64709e3d75f91fcdf3088a0496cab6a18ce76428

Observation d672d0c0-2f1d-46bb-81ac-1a78f2cbf2ed · outbound

This paper cites IPAD : Iterative, parallel, and diffusion-based network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts IPAD : Iterative, parallel, and diffusion-based network for scene text recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.415705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.139112Z digest=sha256:10ce289cc5e30980e5dcddc14c927215dcbb95b4f5981d73abf2d2248da0fa5d

Observation 329c6e7e-26df-4a66-af7f-658495b5cf18 · outbound

This paper cites Turning a clip model into a scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Turning a clip model into a scene text detector

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.393041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.295822Z digest=sha256:1439b1a6ed3f929d05c7980d99d7caacfdc416fbe274b7935c6e97f390e677d2

Observation ebecad2c-826a-4bac-b244-797a5e3837c7 · outbound

This paper cites Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.372619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.427022Z digest=sha256:60eba9015c24a3c5fa4fa9dbd3cd351aba7b770182144bc73f63ebf26b896a08

Observation 0b96d416-462b-4063-8ac6-9952abdee07c · outbound

This paper cites Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.345456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.623053Z digest=sha256:bbc78dd2aba3c9ec7538010e1c5a762d76048be84307d47a0b68ea600ec69c53

Observation 3cec4b21-caac-42c8-99a8-3e5dc85313e0 · outbound

This paper cites TextCtrl : Diffusion-based scene text editing with prior guidance control.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextCtrl : Diffusion-based scene text editing with prior guidance control

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.314377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.764417Z digest=sha256:24dfa4a01cd615a8893b387bb2dddd74e8189f5b323c42911047a66d39901eb6

Observation f22ec2ea-6ebf-4912-a5cb-199b77e2db53 · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Icdar 2019 robust reading challenge on reading chinese text on signboard

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.285972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.932231Z digest=sha256:5837ed6aa2dbeb73ae41e6d636063b3133e59c3c42ff8d2c7cf29052bc31e4c7

Observation 108619cb-9a22-4768-9e8b-9aa182cd25cf · outbound

This paper cites Linguistics-aware masked image modeling for self-supervised scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Linguistics-aware masked image modeling for self-supervised scene text recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.257084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.135547Z digest=sha256:4b95b98cb771aeef4172838ac1058b836bf695c81c3138488a581fbaeddcd993

Observation 2457c7fe-0a2f-45d4-bcff-032fe5305315 · outbound

This paper cites C., and Dai, B.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Dai, B

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.142177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.265589Z digest=sha256:155db17e03c0f2034913d55fdf699e5c9c245c09aa27de8470b7c31e0fb9c54c

Observation 7934d71d-86ab-4dcf-8d81-aed4847aa900 · outbound

This paper cites C., and Liu, Z.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Liu, Z

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:03.908191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.386877Z digest=sha256:a9688c027e3253df2d10ead0490e6e9a4e657b28b83c8e786328dccbcff3eab8

Observation ca56f343-47b7-4f98-b6f0-f9d6456a2e1a · outbound

This paper cites write newline.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:03.546459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:03.546459Z digest=sha256:208ed379c2d45bb2da89e694b5a178e760b466a99a7d59e98f575edef91a0427

Pith citing papers

Observation dad30a42-fba3-4c1c-b7a1-f2edd0a12c4e · inbound

MORE: A Multilingual Document Parsing Benchmark and Evaluation cites this paper.

MORE: A Multilingual Document Parsing Benchmark and Evaluation Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.895411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.895411Z digest=sha256:d9859641743dba9a89c3b6b7481abfae6d1e39d7b6e3ae16418d07260c662317