Pith. sign in

Paper Citation Record · LEDGER

InstructOCR: Instruction Boosting Scene Text Spotting

As of 20 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2412.15523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15523 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:24:16.101773Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed2e4b1-4105-4225-b75c-1c21a627f890 · outbound

This paper cites Context Perception Parallel Decoder for Scene Text Recognition.

InstructOCR: Instruction Boosting Scene Text Spotting Context Perception Parallel Decoder for Scene Text Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.029155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.029155Z digest=sha256:8860492ab57adefdceed927f5e73e159320981f139901417ce5423e471f32508

Observation 67e2161a-9d07-4967-993d-74c303dda212 · outbound

This paper cites InstructDiffusion: A Generalist Modeling Interface for Vision Tasks.

InstructOCR: Instruction Boosting Scene Text Spotting InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.033705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.033705Z digest=sha256:7707ce1c6c8ba4c3b2ed5a6d77b4904d992e79937dd1e5761526866f1d9d7915

Observation 7cbb866a-e573-4159-aeec-6e46a2ad7c9a · outbound

This paper cites OCR-free Document Understanding Transformer.

InstructOCR: Instruction Boosting Scene Text Spotting OCR-free Document Understanding Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.048960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.048960Z digest=sha256:59fdce79ea926f07d9b540c1e657da177a8c2f59ae2bfdf47d46c0320c4aa682

Observation 7c1003f8-8de2-47c5-8774-3b8ac79bff0f · outbound

This paper cites Segment Anything.

InstructOCR: Instruction Boosting Scene Text Spotting Segment Anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.054357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.054357Z digest=sha256:547992db8f938be2a529a502e61674d386d926519cddcfe2cc37a5f2be3c84c0

Observation 55638be9-ac1e-4fa9-b141-dde60e63ca80 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

InstructOCR: Instruction Boosting Scene Text Spotting Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.059498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.059498Z digest=sha256:547f9cf82505571b0472c941762c5079931d116269ebcba40720246a5567e992

Observation bcfc6fd0-f36c-4591-a414-ec1fbebe0002 · outbound

This paper cites SPTS v2: Single-Point Scene Text Spotting.

InstructOCR: Instruction Boosting Scene Text Spotting SPTS v2: Single-Point Scene Text Spotting

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:24:16.222183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.064952Z digest=sha256:00bb3d274de8f02f614e8ce9fe7fe9d51d4a224269276deec6fbbf46417151da

Observation ee1b0aa8-656e-4374-b20a-32bffae4b191 · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

InstructOCR: Instruction Boosting Scene Text Spotting KOSMOS-2.5: A Multimodal Literate Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.070213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.070213Z digest=sha256:4cdd2d663cf4dea7dde9e89b1a3688e2f471389b27ee63c08c64dc527a5ce089

Observation 08b6e3a0-caf1-43b4-82ba-31959c3f891f · outbound

This paper cites In 2017 14th IAPR international con- ference on document analysis and recognition (ICDAR), vol- ume 1, 1454–1459.

InstructOCR: Instruction Boosting Scene Text Spotting In 2017 14th IAPR international con- ference on document analysis and recognition (ICDAR), vol- ume 1, 1454–1459

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.365884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.075726Z digest=sha256:a7ee5cebc9f9d86fb5d1e39ef01382d1985229c158d8e06ba9dd90a5489e6d90

Observation 4dac2de8-9d73-4393-ae63-54579e5874c6 · outbound

This paper cites UPOCR: Towards Unified Pixel-Level OCR Interface.

InstructOCR: Instruction Boosting Scene Text Spotting UPOCR: Towards Unified Pixel-Level OCR Interface

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:24:16.185327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.088303Z digest=sha256:3c03dd1b2b691e2a00eebfbac922a377a97125db5de42ba2f9493caa8087897e

Observation c6d0c6cb-54eb-499f-adce-c11763938a7a · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

InstructOCR: Instruction Boosting Scene Text Spotting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.101773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.101773Z digest=sha256:e67850f6ee12b35987f8fa868a69ba7d00d21fb68529b3ed430444851852cc61

Observation f113a607-6fe4-4838-bc18-46fc8ff6dbc9 · outbound

This paper cites In 12th international conference on document analysis and recognition, 1484–1493.

InstructOCR: Instruction Boosting Scene Text Spotting In 12th international conference on document analysis and recognition, 1484–1493

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.039342Z digest=sha256:d3dfa845d153e72236455c5ec0ab1d26354f0cefdf2cf076852f1e94f60231c6

Observation 8f093086-639d-4626-b5c1-3082f28b7b1b · outbound

This paper cites In 13th international conference on document analysis and recognition, 1156–1160.

InstructOCR: Instruction Boosting Scene Text Spotting In 13th international conference on document analysis and recognition, 1156–1160

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.381022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.044359Z digest=sha256:3cd95ba158f8fef0169b348b5c8f8d6ee98fe4ed29937eff5faf8a2ccd421f8f

Observation 02c21729-2244-4df4-b7dc-09dfb6d184a5 · outbound

This paper cites In 2017 14th IAPR international conference on document anal- ysis and recognition (ICDAR), volume 1, 935–942.

InstructOCR: Instruction Boosting Scene Text Spotting In 2017 14th IAPR international conference on document anal- ysis and recognition (ICDAR), volume 1, 935–942

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.413833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T11:24:16.013602Z digest=sha256:f1209d707acf7dff5353aeb0c483ac8ec22c04fb8a900c47e468bf38ef7e2c1c

Observation d915012a-b1e2-44be-92fa-3a29e0daee91 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

InstructOCR: Instruction Boosting Scene Text Spotting BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.017859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.017859Z digest=sha256:48247dae66cd1945d6232e15176ee78ad0f5f105b58f3e8ea038adcbdf3a3211

Observation a5d838e3-9ea4-449b-9aac-ab771b998daa · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

InstructOCR: Instruction Boosting Scene Text Spotting An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.024051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.024051Z digest=sha256:5659333b7b12be85db11ac6528d8b4f68fc2c2ebb18ac1f8256f23fbda943e5a

Observation 3dca4bb9-7760-4cb3-8a55-75c53a62262a · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

InstructOCR: Instruction Boosting Scene Text Spotting Pix2seq: A Language Modeling Framework for Object Detection

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.009164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.009164Z digest=sha256:02ffe069136f56341d4fc30c9e24c25b3a68b626ff7b90fb06f7b366d7d9de30

Observation 7507f045-ef53-4b51-9d07-75b3bad64f5a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

InstructOCR: Instruction Boosting Scene Text Spotting GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.096991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.096991Z digest=sha256:3a932e62a162be2674c0eb45ca07fc418d3eda4a2a89685e019f8d4b2d9adf5d

Observation c522d627-393f-42a0-9000-e8c0a3b9b3d5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

InstructOCR: Instruction Boosting Scene Text Spotting Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.002128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.002128Z digest=sha256:690357ae1a01b79810efdf9876561c36c33a9462513ce361de934e6869e7d869

Observation 8c07e379-aa0b-41b6-927f-a3aafe304c2c · outbound

This paper cites OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition.

InstructOCR: Instruction Boosting Scene Text Spotting OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.092992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.092992Z digest=sha256:43e45545b35f7ed43aafda444745d149159a20fa60698639fca818ccb8a13245

Pith citing papers

No inbound Pith citation observations are available.