Pith. sign in

Paper Citation Record · LEDGER

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2606.07161.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07161 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:10:59.631543Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4785c829-436c-47bf-a32c-c663e4204170 · outbound

This paper cites Scene text recognition for text-based traffic signs,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Scene text recognition for text-based traffic signs,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:d7b240451d0d2c85ba7da888cad8d4e43de3369d1ce191ad8c81dbc425b17af5

Observation 19e53c88-7130-4d44-b649-901f53ce3217 · outbound

This paper cites Scene text detection and recognition: The deep learning era,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Scene text detection and recognition: The deep learning era,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:54d4e35ddfe2f48399a14dc377830b0ec125dc65885ba7bc93a613338b159b06

Observation d356a016-1b7f-49c4-b354-616507edf7e1 · outbound

This paper cites End- to-end video text spotting with transformer,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance End- to-end video text spotting with transformer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:3ba6514d2aca79996e3faceccc9f112644e0900a9cf54803727284e3d9e9044d

Observation 4c6b0cb7-1658-435b-bc86-93b695c810b0 · outbound

This paper cites GoMatching: A simple baseline for video text spotting via long and short term matching,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance GoMatching: A simple baseline for video text spotting via long and short term matching,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:0d06beeeec2c600b48e40c4d34b79aae496ad5e2e7c3536529487fdcfbb7c2f9

Observation 18fc7068-cbc4-4a3d-9fc8-b57e558128fc · outbound

This paper cites arXiv preprint arXiv:2505.22228 (2025) 2, 4.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance arXiv preprint arXiv:2505.22228 (2025) 2, 4

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.904890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:eb5413ac663326f4c32594b8db76c832922c2dd8dd1e25c66c024ef927fe6c2c

Observation 5ee2c62d-99ef-400a-b18d-477ae558001b · outbound

This paper cites DeepSolo: Let transformer decoder with explicit points solo for text spotting,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance DeepSolo: Let transformer decoder with explicit points solo for text spotting,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:8332d6507dc4c5b5b299c3762648a75c10aba2f2d75519a5dfa429db43654cfb

Observation 875f259e-2e01-49a2-a403-68242dff3481 · outbound

This paper cites Scene text recognition with permuted au- toregressive sequence models,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Scene text recognition with permuted au- toregressive sequence models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:b5ded6df095d6e1ffa8df7d0c6c602235e54d440e729c8b347d7e741930f0754

Observation e7be6838-e052-475c-a210-2bd3e9feadb9 · outbound

This paper cites SVTR: Scene text recognition with a single visual model,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance SVTR: Scene text recognition with a single visual model,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:056d946b631855c7930c75ec627d260d7f398b74460a1497a5c1ec1153673b35

Observation 3268f207-ab95-43d9-ad66-2b4279e8498a · outbound

This paper cites A bilingual, openworld video text dataset and end-to-end video text spotter with transformer,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance A bilingual, openworld video text dataset and end-to-end video text spotter with transformer,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:b695015895fbc660b029b7c85f7cabd4317b07a4ab52ffda992edfe4a2a7aa0f

Observation 16a51800-4092-42c2-9585-05ead374f9b9 · outbound

This paper cites DSText V2: A comprehensive video text spotting dataset for dense and small text,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance DSText V2: A comprehensive video text spotting dataset for dense and small text,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:103b939e139fc83fd058148153976f9452990ff90729de51a077ee73e45a4f26

Observation 09d9c7c3-2931-41e4-b4fb-49d4755050de · outbound

This paper cites ICDAR 2015 competition on robust reading,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance ICDAR 2015 competition on robust reading,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:62afd971a5c556e54cc67afb62125f3cb6111204e7f762659347db3653df48c2

Observation 1daa5f3e-c123-4420-b4ff-2c29f58abdf5 · outbound

This paper cites Textssr: Diffusion-based data synthesis for scene text recognition,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Textssr: Diffusion-based data synthesis for scene text recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:7c155180eae051224a09619990588b514daf6da563399519ece19218a68c2c89

Observation be95ccc1-ef56-4b27-aaa6-c0c2d5415d0e · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance LoRA: Low-rank adaptation of large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:8f899f6ac2854dee0f1888d58f85f47bc1c0fc857429a6462e3126365d20fd22

Observation a8030de4-c31e-436a-9411-1957279d8a41 · outbound

This paper cites Q-Adapter: Visual query adapter for extracting textually-related features in video captioning,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Q-Adapter: Visual query adapter for extracting textually-related features in video captioning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:96c3abb6532a32e8ec3ad7307561eb9f1198afeccdb9becf7422904b1a09a31d

Observation c8c97e7d-2e14-402d-92a4-433bac2d4d02 · outbound

This paper cites An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:33e84e19d307574a2fed744fca64705369af4e878271a70911ff9192d4106edd

Observation f107edfe-779e-4e77-9806-90a6419db82a · outbound

This paper cites Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:2beaca9e0f354cfc083b75da2e1cd2dc1fa5ad609cc9693f146d72073baa8b30

Observation b8e48c9a-361c-457f-8ded-69a73777212d · outbound

This paper cites SVIPTR: Fast and efficient scene text recognition with vision permutable extractor,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance SVIPTR: Fast and efficient scene text recognition with vision permutable extractor,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:c79471f751db22290ce679126115601ddec4aa1e6dde8f18cc457e7b4aa0f753

Observation 25a19c10-fb58-4c65-a31a-4097b6d5d9f2 · outbound

This paper cites DiffusionSTR: Diffusion model for scene text recognition,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance DiffusionSTR: Diffusion model for scene text recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:3eaba5917a49a8163d5ce3162f5189baddf77389dfcd530dc3816f7f8d19dd1d

Observation 073182ac-aad0-4487-ba45-7cf0960b7b64 · outbound

This paper cites Mask textspotter v3: Segmentation proposal network for robust scene text spotting,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Mask textspotter v3: Segmentation proposal network for robust scene text spotting,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:f428f19f7e5353b12faee79f3b838282da7789b5ff6b63b075853d67b939f27a

Observation 521e633b-b5c7-41d1-9297-b9a96d75ddde · outbound

This paper cites Swintextspotter: Scene text spotting via better synergy between text detection and text recognition,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Swintextspotter: Scene text spotting via better synergy between text detection and text recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:357072cef2dfefbd34541c11cf70851aed4607e6c992b96f65e94262045bee2d

Observation 86a69bbb-2687-49c3-a144-42872b0c8a7a · outbound

This paper cites FREE: A fast and robust end-to-end video text spotter,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance FREE: A fast and robust end-to-end video text spotter,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:0161b851941e9eaf5df1b0918fe1ed65f68ba09e6afca5a10dc48cfaf4b03e53

Observation f7f030a5-e2b9-4842-aff4-c6c65b6e0c25 · outbound

This paper cites CLIP4STR: A simple baseline for scene text recognition with pre-trained vision-language model,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance CLIP4STR: A simple baseline for scene text recognition with pre-trained vision-language model,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:754c7f13aa14c29d1728854fced9f51983122da5b3e7dd210324b11333eb55d2

Observation 9a6ca382-8500-4992-8488-4b43e1926d03 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Learning transferable visual models from natural language supervision,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:5ecec5971b11d3b73772f60c558d2fede5ed104e121879f1c46b82f4a3da28e4

Observation c138ae19-0c64-4c73-aa13-32e6bd6e211d · outbound

This paper cites Nougat: Neural optical understanding for academic documents,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Nougat: Neural optical understanding for academic documents,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:e62313b951495cc8063d500953d2259cec7ac1ba92d0212f3da10b6eb2bc8e5e

Observation 08f752af-44da-4e78-b4bb-8cc7265f3c39 · outbound

This paper cites Ovis2.5 Technical Report.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Ovis2.5 Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.899659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:3dc4e53c07d5a6aed72f76f5be4037560e0166d01cb3a0b996485954af3cdadc

Observation 4ef7eeea-271d-4e37-a679-afbf42a04b77 · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance VLMEvalKit: An open-source toolkit for evaluating large multi-modality models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:ff3727829201c6eb8d6b73f8cc8531734b51f2737122ce01e43f5be4ec5f082d

Observation 72cbb3ad-408b-4203-9a10-7943f7b91060 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.897437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:7d965ecc2bdc1be51e4dcc52caf5bbe7531735cc3567ee5ed44dacb7d8546613

Observation ff32975b-cc8d-43cb-a880-8f4f527abf64 · outbound

This paper cites Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:f15e87507b901973b701b19f27cce750324809e52f25360e5e328624d03847c7

Observation e8ff6cc0-5739-4626-a7f7-67c53bed7488 · outbound

This paper cites Mutating ordered $\tau$-rigid modules with applications to Nakayama algebras.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Mutating ordered $\tau$-rigid modules with applications to Nakayama algebras

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.902183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:b8ea754243d17908c96a2f39ef79fb58e5a5108a8ff644294d7a189427cf8a0e

Observation 187a44c9-ee85-4d9c-a253-6b3a3a16fcd3 · outbound

This paper cites RoadText-1K: Text detection & recognition dataset for driving videos,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance RoadText-1K: Text detection & recognition dataset for driving videos,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:7b839a2a29f04a9b41e99da82f4c63be9f7452dd33144189c8c0240d2e67b485

Observation 515c8b93-af60-4774-9868-22b7d4064f0a · outbound

This paper cites Evaluating multiple object tracking performance: the clear mot metrics,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Evaluating multiple object tracking performance: the clear mot metrics,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:65efa978cf71fd288c60629b50fdcb91feb07dc2385fad1d19eb64690d7317dd

Observation 637cf813-bf5c-4a46-bd8b-9c0d1324d26c · outbound

This paper cites Performance measures and a data set for multi-target, multi-camera tracking,.

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance Performance measures and a data set for multi-target, multi-camera tracking,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T22:10:59.631543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T22:10:59.631543Z digest=sha256:5b9d647dcc713b7cead964d212edc82c8e6bf54bed22e81d4e97e412ca8e4508

Pith citing papers

No inbound Pith citation observations are available.