Pith. sign in

Paper Citation Record · LEDGER

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2507.00454.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00454 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:08.038670Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:20:54.642591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:24:21.277633Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0634cba0-2421-4b87-88c0-09f2b4f9763e · outbound

This paper cites Hiptrack: Visual tracking with historical prompts.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Hiptrack: Visual tracking with historical prompts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.929103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.495167Z digest=sha256:454e696b794fe028a9a38ff86f28fec720e70b6bda6287f2ddd187e8777b8d82

Observation eabbfb60-8b03-41bb-91cc-f8cc805ca021 · outbound

This paper cites Robust object modeling for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Robust object modeling for visual tracking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.915407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.562427Z digest=sha256:bec0d7b6d77f5ad5a3a8d4a600c62597cc383fc4b5bd7f5fa1b7d16f8f3c3d4b

Observation b2115ddc-afda-49a4-93aa-6452e2a3f57e · outbound

This paper cites Ost: Refining text knowledge with optimal spatio-temporal descriptor for general video recognition.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Ost: Refining text knowledge with optimal spatio-temporal descriptor for general video recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.901038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.635785Z digest=sha256:1b2676bdedf32b98a78b6e66dbfa176eb9e83f7ed1d1daf256037597d81b691b

Observation d6c5bd46-d76c-4189-8199-7bc98634a03e · outbound

This paper cites Seqtrack: Sequence to sequence learning for visual ob- ject tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Seqtrack: Sequence to sequence learning for visual ob- ject tracking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.887068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.711114Z digest=sha256:11a4c42782bb4b725a46b6246c9173fa22618cf89f97e30834450f3fbc7fcf90

Observation a36d3e90-8321-4c59-9d61-30a9f50cbdd6 · outbound

This paper cites Mixformer: End-to-end tracking with iterative mixed atten- tion.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Mixformer: End-to-end tracking with iterative mixed atten- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.873201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.807241Z digest=sha256:a2726748aae614221f55afa2bd2b5908b79187e6ac744b950457f733cc4efea2

Observation b4c15925-262e-43c1-93f4-db01dc76931c · outbound

This paper cites MixFormerV2: Efficient Fully Transformer Tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales MixFormerV2: Efficient Fully Transformer Tracking

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:21:08.313273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:03.895699Z digest=sha256:85ea7c530ae0eb8a46fdfe34556b99cd57ae6e6864aa03cff58dbdf60735f075

Observation 2b70bf93-1333-489f-aee9-e2bb68d051e6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:04.010181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:04.010181Z digest=sha256:76dc90afda2a2369beb498f47e8a07e649b04ee2da324fa4fbfff6cd171fbdfa

Observation 2da48220-6dc6-4af6-9e3a-21ca566761d3 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.859467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.110836Z digest=sha256:d41495978147cfbf42022348295f572dfa6d53af36b06b25fecedc7ccaeef40e

Observation c932f6be-67eb-488a-8651-65a9c8cf0862 · outbound

This paper cites Siamese natural language tracker: Tracking by nat- ural language descriptions with siamese trackers.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Siamese natural language tracker: Tracking by nat- ural language descriptions with siamese trackers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.846051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.177793Z digest=sha256:adfbb0b455004676ac29ca18641ecd7e9723c0587c4b3b09dbd99350e80ae44a

Observation f43e7444-8140-4443-8336-05a45de9ae8a · outbound

This paper cites Generalized relation modeling for transformer tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Generalized relation modeling for transformer tracking

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.831765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.242868Z digest=sha256:8e4642f2be947e42e0de681543137846b1a78ce4574203acacb6c1f5a1544ebb

Observation a702d729-0758-4471-aacc-60cb5f5e99a8 · outbound

This paper cites Siamcar: Siamese fully convolutional classification and regression for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Siamcar: Siamese fully convolutional classification and regression for visual tracking

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.818032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.345312Z digest=sha256:eae1ba301e40b0869e6124bcaece7511cd9f51ed7b0538448eab4e469f884d88

Observation b178a3a2-e6b5-4675-8a29-79d9a77bd249 · outbound

This paper cites Divert more attention to vision-language tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Divert more attention to vision-language tracking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.804426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.494381Z digest=sha256:81d0177ad7668d65bb5b2569d132b7ad0091465103607bc0c626141d682bcf1b

Observation d1e1ea23-8165-4bd7-b684-830a03a2a0bc · outbound

This paper cites Masked autoencoders are scalable vision learners.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Masked autoencoders are scalable vision learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.790001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.590862Z digest=sha256:48672c514a28160924486d2d5565e9959934ceffa580016d9bf2c3478adcd7a0

Observation c21eb719-5b92-4c1f-ba85-e0dc9d398176 · outbound

This paper cites Target-aware tracking with long-term context at- tention.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Target-aware tracking with long-term context at- tention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.776016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.749674Z digest=sha256:6fed5f730a659d8b18a493d6533ca18a74863d7a25b2dbba543fe51936879c3e

Observation db567757-5c6d-4cbf-b351-a7e00a7646fe · outbound

This paper cites A multi-modal global in- stance tracking benchmark (mgit): better locating target in complex spatio-temporal and causal relationship.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales A multi-modal global in- stance tracking benchmark (mgit): better locating target in complex spatio-temporal and causal relationship

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.761688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.878689Z digest=sha256:3286d2bd9dd9f8c220b8f3193eae5e8d492b8e065632814670411867cf75a0ab

Observation 70a604bd-57af-43ad-af90-1cebdadefccb · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.748053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:04.990174Z digest=sha256:aad776713746313b631a2ee757cf0442a86a7a7e847a47b28e301749ccb4ef6a

Observation 1a1abc37-1e63-44cd-bd3c-b35a96fa3a58 · outbound

This paper cites Rtracker: Recoverable tracking via pn tree structured memory.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Rtracker: Recoverable tracking via pn tree structured memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.733707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.106071Z digest=sha256:c32eb2fd556627cd87e3e94c070d48e0e35cc26a21b9d89fe7dc03027abe8875

Observation 391f9ee8-b9d4-4ba3-b8d6-722e483c4b3f · outbound

This paper cites Towards sequence-level training for vi- sual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Towards sequence-level training for vi- sual tracking

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.718635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.230930Z digest=sha256:ef167c7fffd8da0f1d783c764f9c5518b29fd485b7c814bbbfa669e78354cc9b

Observation c26405f6-b8d3-403b-b6a2-eac472e374ce · outbound

This paper cites Citetracker: Correlating image and text for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Citetracker: Correlating image and text for visual tracking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.635446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.358470Z digest=sha256:1129320c89c6c0cc6a662c21e22d02dc8b7d806a5f2c3f4dd2c59da57a7a4c5e

Observation 886f1d5f-ed15-4466-93f3-ba4143f64452 · outbound

This paper cites Dtllm-vlt: Diverse text generation for visual language tracking based on llm.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Dtllm-vlt: Diverse text generation for visual language tracking based on llm

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:13.302631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.489098Z digest=sha256:6cae40e3c996ffa5f19ba1c908bfab487e9b5bbfb58102a7f6e8d58e3910dec9

Observation 4f11bd9e-500a-4cdd-8786-fb5beb1d4da9 · outbound

This paper cites Beyond MOT: Semantic Multi-Object Tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Beyond MOT: Semantic Multi-Object Tracking

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:21:08.163075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.618383Z digest=sha256:6beca9a7f1295e188f3fbd03a9070009c6bc381cee0b387c0ca89756184275cb

Observation 2b0a9cbc-00e2-4fba-bcc1-13c1fb6f4899 · outbound

This paper cites an unresolved cited work.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:21:13.057046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.728136Z digest=sha256:6ae9133da40e01362d877f70d64b45344ad50c5a1ca123d2361b1bf94bd21b3e

Observation 5011503f-7aa2-4a2d-8056-94d7647f901b · outbound

This paper cites Swintrack: A simple and strong baseline for trans- former tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Swintrack: A simple and strong baseline for trans- former tracking

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.777830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.856313Z digest=sha256:d11dbe994dd4cb4e8085ea08bcefe4eed140b7de24debf5f3e2c9040ce2793d1

Observation eaf084b2-d4ad-448d-9ec6-fbf7829d4b75 · outbound

This paper cites Tracking meets lora: Faster training, larger model, stronger performance.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Tracking meets lora: Faster training, larger model, stronger performance

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.550238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:05.995947Z digest=sha256:ffddd93aca27b4ee3486eeb734b4e6b709599bc94a847b8464c044fe2a2fdb9c

Observation 77c8ef6c-dc8b-43a3-bc8a-b950ded5fc8e · outbound

This paper cites Tracking by natural language specification with long short-term context decoupling.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Tracking by natural language specification with long short-term context decoupling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.343384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.105323Z digest=sha256:b289ed256fec88969de1e0b4cb106ac0b10c0f6a62f16ea1aab4d902c33bb626

Observation 4f3ab851-5150-4666-9679-f7bf9a13dc81 · outbound

This paper cites Unifying visual and vision-language tracking via contrastive learning.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unifying visual and vision-language tracking via contrastive learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:12.123095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.246551Z digest=sha256:674166ebeb46bd18ae8ab3bbed7cecea7d5625eb3426d866d71c7b3a6e326168

Observation 4afffe9b-0de5-45a7-a5a3-9e8e41b1db7b · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.866219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.412887Z digest=sha256:b38f00a7d52fea4e7607ab51dbf0a241a7ae076ba8495077993583333a43283c

Observation 11ce511f-a597-4d94-8a2d-43bd5d5a5c90 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Learning transferable visual models from natural language supervi- sion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:06.545414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:06.545414Z digest=sha256:8ee255513e6726d4e11887ae9f6d83f92f4a603549b5dd5e93ec8eff7c8e1745

Observation 2ccac787-eac7-48fc-b362-5d96d258a98a · outbound

This paper cites Context-aware integration of lan- guage and visual references for natural language tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Context-aware integration of lan- guage and visual references for natural language tracking

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.513979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.691730Z digest=sha256:227ef7bc6912e4eb2df9cd3c33ac8d1af9bdfb99c5c65692c6b72aaf8c329c4c

Observation 47e8c3f8-087a-489d-842b-f5551db649d3 · outbound

This paper cites Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:11.202632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.808290Z digest=sha256:70cc6dcde570d8632858f39bf06a752eb66163b08c00c1ad1f4156c99687b0d0

Observation 6905c2cb-a44a-4f2a-a8e6-199fbbf32e9a · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.992361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:06.922319Z digest=sha256:cce7a2a7dc88f02c247000009215ccd2a83ede43ae45534fcb87583d64c1be94

Observation 0440d920-f877-4a4a-995d-03cc064e50ca · outbound

This paper cites Autoregressive visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Autoregressive visual tracking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.682233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.042125Z digest=sha256:0fc45e1390898449432947710142194be21715b1960bcda6fce34a5d52729d39

Observation b970d63e-fa50-44fe-a6dd-600895141981 · outbound

This paper cites Improving visual grounding with multi-scale discrep- ancy information and centralized-transformer.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Improving visual grounding with multi-scale discrep- ancy information and centralized-transformer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:10.402450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.202804Z digest=sha256:8c1c5d4680092ef281dcbc654078ca273155da35a966f6616783acca69c05330

Observation 70f2c872-bded-448e-bab3-b82b5fd045f4 · outbound

This paper cites an unresolved cited work.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:21:10.155071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.347044Z digest=sha256:5e01e13f4a99fcdc032744b997ce0fa4b318ed874cb85ed666ebc24cd5d30aa0

Observation 0e722b8f-2db2-4702-a60c-816d6726111a · outbound

This paper cites Object track- ing benchmark.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Object track- ing benchmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.720456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.447212Z digest=sha256:8a30793896377f03b87f88b03eae180cb733a7b95eaa7537d89931fd2753430b

Observation 4ea25646-625c-4125-a4bb-cd649532ef37 · outbound

This paper cites Correlation-aware deep tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Correlation-aware deep tracking

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.516437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.587470Z digest=sha256:a8c88fec8e42b04d38a5666a36e5ff6735e199fa175205f021b9b5a23af13cfa

Observation 502c8dee-b8f0-4046-acd8-0faf4c29afeb · outbound

This paper cites Autore- gressive queries for adaptive tracking with spatio-temporal transformers.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Autore- gressive queries for adaptive tracking with spatio-temporal transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.363218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.678542Z digest=sha256:3f57482bbcb34d8470e11b0279778ad824c9c0a17f2cc5c9bcb48350c0bb637b

Observation 54f03118-3210-4a38-92eb-94896ca203ec · outbound

This paper cites Learning spatio-temporal transformer for vi- sual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Learning spatio-temporal transformer for vi- sual tracking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:09.150372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.754690Z digest=sha256:f2f3887ddfda7456c04d09b4baf0b2baa32d007704881d0fb64c2abd7252f1e9

Observation 33cdc7ef-33c0-4e99-984a-0690f2c19783 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.924328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.813342Z digest=sha256:3041369a34a95acbcc84157a8264c5b28ca777f94c5e46c9f6141cb2ea059c4e

Observation 09e8bdca-f269-459f-a03f-01a7f7d42358 · outbound

This paper cites Exploring the feature extraction and relation modeling for light-weight transformer tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Exploring the feature extraction and relation modeling for light-weight transformer tracking

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.809214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.876163Z digest=sha256:13b8ebd80c95a2ea2935905f0e1e914b7ab758484eaf647e7d99cec47051f493

Observation 536240b9-8b56-4e9f-b962-66ad43d28bc0 · outbound

This paper cites Ji, and Xianxian Li.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Ji, and Xianxian Li

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.669724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.946025Z digest=sha256:1a0ab99bccf29976c5aa226d46f0ee6b30068bfb66ee136c9b04b263e63ad332

Observation 329e24f2-28f4-4ac2-8d79-c3daba8fe15b · outbound

This paper cites Odtrack: Online dense temporal token learning for visual tracking.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Odtrack: Online dense temporal token learning for visual tracking

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.554193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:07.998946Z digest=sha256:a9aa8dfc4df4b332c2e55c34e493e08d871bc8bccbe0ee7e895bdb4b456f3ecc

Observation 587798bc-d420-4fb2-8e7a-8fb7a4e99f4c · outbound

This paper cites Joint visual grounding and tracking with natural language specifi- cation.

ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales Joint visual grounding and tracking with natural language specifi- cation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:21:08.433120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:21:08.038670Z digest=sha256:96531b4c662f705e067bec41aee73a395dd8c0dad6aa87881f4a58357415de10

Pith citing papers

Observation 9bd60c18-a9cc-4ce4-9dc2-387bbc900887 · inbound

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking cites this paper.

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.279204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T07:20:54.642591Z digest=sha256:564bbbe313c528ffa0880692fbbdb032065ab7d355e52621660e263163e313ae