Pith. sign in

Paper Citation Record · LEDGER

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

As of 18 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2507.19875.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19875 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:00:25.536019Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b060917-d2fc-4852-a4e9-0c929326980c · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Exploring Visual Prompts for Adapting Large-Scale Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.051633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.051633Z digest=sha256:6f8ae3ed14171defdda5e9cb878768ec6cc26333c06b85b03bc56e2a80777e11

Observation 2b6168c0-3fd5-451b-91ce-69a90b892be6 · outbound

This paper cites Qwen Technical Report.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.057677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.057677Z digest=sha256:2de8b63898cde2368bc8324d087023637a0d60a913d6f0dc901199780252c23e

Observation feac114b-a727-4bb9-ba01-7b223253147e · outbound

This paper cites ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.062724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.062724Z digest=sha256:4c7cb70a3c8b9e927e81ed35ab41e52963b7df32e422040e109db4873ae6b8e9

Observation 31ab4e52-3b86-4a21-9c3b-bc1e9f8aaef6 · outbound

This paper cites Visual objects in context.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Visual objects in context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.067661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.067661Z digest=sha256:ca4808d40b75dc47ca4d75704fd31909054ff2ec201205366f3b34b624da3686

Observation 6dd1badc-254e-41ee-a3ba-7f44c1ff97d2 · outbound

This paper cites Fully-convolutional siamese networks for object tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Fully-convolutional siamese networks for object tracking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.072541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.072541Z digest=sha256:e8ff8e489c51f2d5e3325ac489e5129bed6d53b7f033a8fcad069dfdc43a7176

Observation c458deef-7dc7-4f87-b3d3-1d08557fbc81 · outbound

This paper cites HIPTrack: Visual Tracking with Historical Prompts.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking HIPTrack: Visual Tracking with Historical Prompts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.077398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.077398Z digest=sha256:e371ce44458192f72034e8bba2fed5880d79ce6f824787e821757d01af9744ea

Observation d6664951-cb60-4632-94aa-56def1c74905 · outbound

This paper cites Hiptrack: Visual tracking with historical prompts.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Hiptrack: Visual tracking with historical prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.083037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.083037Z digest=sha256:bf83acf155a6c1945d806fc451bc77fbc10fe8d2c13c290ade25625e4ac384c8

Observation ea980acd-dab8-4366-8809-ccbad76b884d · outbound

This paper cites Robust object modeling for visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Robust object modeling for visual tracking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.087543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.087543Z digest=sha256:1edc89f0ae51f7d741482fbbae2c265596d3fe0a79014bef7483420327911e95

Observation 21d8f018-330a-4a16-9f34-107911ef917f · outbound

This paper cites The relative con- tribution of scene context and target features to visual search in scenes.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking The relative con- tribution of scene context and target features to visual search in scenes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.091932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.091932Z digest=sha256:fd83a17b59d4a4326819cba487449e3d579950ec29b233fc78a527128d2840d2

Observation 8fb728ed-985c-4e69-843e-ca843aeb2b6d · outbound

This paper cites Back- bone is all your need: A simplified architecture for visual object tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Back- bone is all your need: A simplified architecture for visual object tracking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.097020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.097020Z digest=sha256:f679128f1080344c68a61d213a5c2045fb2e52a91c081ca010465f42c31f6911

Observation af9e0079-0496-4bd5-9533-29d214d34c57 · outbound

This paper cites Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.101699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.101699Z digest=sha256:557a47cda71eb6cd27f95fd73de040d5541f20603f305a4288c09b1cfba1bd89

Observation 173655c0-8471-4ffb-b9c4-7fe625a453d2 · outbound

This paper cites Transformer tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Transformer tracking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.106886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.106886Z digest=sha256:9b682a396bacfaff3b428966a0dcf317f99c572e0a87e3869fb2660cf5fa29f5

Observation c3e2cd41-3d4d-4834-9634-5aab90328e39 · outbound

This paper cites Seqtrack: Sequence to sequence learning for visual ob- ject tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Seqtrack: Sequence to sequence learning for visual ob- ject tracking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.111506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.111506Z digest=sha256:33e4b1647e00ef165858dd8092e88fc8431568c9cc7ef475726296ca908f2c6c

Observation 5dda1452-603e-469b-9b69-278e31137e65 · outbound

This paper cites SUTrack: Towards Simple and Unified Single Object Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking SUTrack: Towards Simple and Unified Single Object Tracking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.115964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.115964Z digest=sha256:c228cd9b5ac10cafff2c4a5bc42a6ba3cd466e10c534b6cf2000064b23335508

Observation 43773dbe-f9ec-4cff-909b-deff0e94bff5 · outbound

This paper cites Siamese box adaptive network for visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Siamese box adaptive network for visual tracking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.120806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.120806Z digest=sha256:87cf293e59009c963ff90312aff5e6d9822510291916fb7e25ff23787a5a830d

Observation 68168ffe-117a-40a8-bb38-bf4d6d3509ce · outbound

This paper cites Mixformer: End-to-end tracking with iterative mixed atten- tion.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Mixformer: End-to-end tracking with iterative mixed atten- tion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.125439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.125439Z digest=sha256:d1d9df9d5dd6c34a57d3977be2a9d4b8cc060f112ff28d89bea8190001c57e88

Observation 977012a5-2c21-4eb4-a6f9-4e10441ca09e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.130259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.130259Z digest=sha256:a0e6ca70d4ef71f340c20f578ab7ecab446e05580d8cc7b020a0e5afb327b99c

Observation 6151760b-09d7-454a-9ac1-62b434f780cf · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.135466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.135466Z digest=sha256:fc16512cfdf711a02f8be1ec2999497394d3723b0f732c93d7af9d60f7f732cb

Observation b59bdd76-e86e-4ee8-b199-040eadc75684 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.139892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.139892Z digest=sha256:b34b1ccee3734404c4e12da7442ddef47954a709eac43d4ede3cf7f900d4f25f

Observation 9787d40e-5a6e-4c78-81ee-db889399770f · outbound

This paper cites Lasot: A high-quality large-scale single object tracking benchmark.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Lasot: A high-quality large-scale single object tracking benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.144678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.144678Z digest=sha256:389f9b768c7fe5bab1efc9c086c601fbd447041c5f9b9916ff189c86ea3fb3f5

Observation 22edd8b8-6ce5-4327-9e52-b260742f54cb · outbound

This paper cites Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.149355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.149355Z digest=sha256:70fe5d1a2f5f357b501675776dac7a671ec15995ac972aaeb59bd4b376bd2bd3

Observation 097562a7-14eb-456d-9e3d-68b66ceabb53 · outbound

This paper cites Real-time visual object tracking with natural lan- guage description.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Real-time visual object tracking with natural lan- guage description

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.154518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.154518Z digest=sha256:23bf6f8e3fb0ebbc9cd9e2dd121a40e9fa70eca8501acdc5616eda08c53d0197

Observation 7ba7bb2d-c6ac-44a7-a639-580704c5b823 · outbound

This paper cites Siamese natural language tracker: Tracking by natural lan- guage descriptions with siamese trackers.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Siamese natural language tracker: Tracking by natural lan- guage descriptions with siamese trackers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.159328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.159328Z digest=sha256:029ed518f2e7795348b8978c8d7970001aff9c9f663e675402405240a6d5129c

Observation 46e13a0a-1579-47ae-a1b5-a1dafd7713b8 · outbound

This paper cites Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:00:25.961926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.164615Z digest=sha256:80a945ba6b75ee018034dc52e6904370b56c2874713fd4e53591c75afca3e034

Observation 5d49d5c0-266c-46aa-9cb6-dab61f2a3a1c · outbound

This paper cites Memvlt: Vision- language tracking with adaptive memory-based prompts.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Memvlt: Vision- language tracking with adaptive memory-based prompts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.169975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.169975Z digest=sha256:c5b2562955cfe54f119ca20eef4a932a2b4499ce6128392c6ceec55d571cc7cb

Observation dfaa284f-62df-4dcf-88c4-6ad7470bb5a4 · outbound

This paper cites Narrlv: Towards a comprehensive narrative-centric evaluation for long video generation models.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Narrlv: Towards a comprehensive narrative-centric evaluation for long video generation models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.174980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.174980Z digest=sha256:b0246bdf8678b11d517dbcb2c35b50a394bc3cec43e4c99fada524417cf52dd2

Observation 6956bd4c-b23c-4673-b875-76a169d735b7 · outbound

This paper cites CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.180490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.180490Z digest=sha256:9fec6375ea05432f363ae2d80c45de33045cd125f155e843100d81c4bda965a4

Observation d89e1c7f-6819-40aa-aff1-97aad547f407 · outbound

This paper cites Aiatrack: Attention in attention for trans- former visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Aiatrack: Attention in attention for trans- former visual tracking

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.185727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.185727Z digest=sha256:275b5b207fe7914e6164fc1cc73d3588433a88c8fa27f37b688ba97505771029

Observation c33d4be3-f5c2-41f3-b167-71479d44de01 · outbound

This paper cites Generalized relation modeling for transformer tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Generalized relation modeling for transformer tracking

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.191737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.191737Z digest=sha256:a66d5bb36f3225346f25fab574c66aafe823ba6744a72e5695e33f55e247df40

Observation 93a04a3d-442f-4d53-ad55-f858ecfda2b4 · outbound

This paper cites Divert more attention to vision-language tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Divert more attention to vision-language tracking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.196227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.196227Z digest=sha256:98e8958437d9a9309531f604b292327da41f14ce88a005a2aa286ee5e2d7ad78

Observation fb14e636-5ffc-410d-b9d1-a77f337883d7 · outbound

This paper cites Learning Target-aware Representation for Visual Tracking via Informative Interactions.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Learning Target-aware Representation for Visual Tracking via Informative Interactions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.200880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.200880Z digest=sha256:85070787772ed697c0fc0791c7cae045c0355e4d617a8818e12bad734b58fb2d

Observation 98daaa08-b2b2-4287-b051-0c2c9679f50c · outbound

This paper cites Onetracker: Unifying visual object tracking with foundation models and efficient tuning.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Onetracker: Unifying visual object tracking with foundation models and efficient tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.205717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.205717Z digest=sha256:3edfe925e9c1a3f20b1b7e2c27cb2bc23037b2435bb64b8489a534757963a8af

Observation a46afc33-de4d-4909-8b30-324264d7ca63 · outbound

This paper cites A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.901630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.210408Z digest=sha256:2a08e3760cb50f64794dbed66cf65adf59b6095f43223860e79cda88bed4117b

Observation b4e30215-6014-4ec2-8208-1f21191831da · outbound

This paper cites Global instance tracking: Locating target more like humans.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Global instance tracking: Locating target more like humans

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.885828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.215048Z digest=sha256:77def6f924ceb496d00503f44184de644155f07e8681f1231270d5f255046600

Observation 8ce4568f-cf5a-4c12-9e6e-1b97b31b4db1 · outbound

This paper cites Sotverse: A user- defined task space of single object tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Sotverse: A user- defined task space of single object tracking

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.870425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.219467Z digest=sha256:d1254a7d9e96dfad6e595b0da4f7d298e335b102702600e4b86981e9e55bdc30

Observation 203bd55e-7cc5-47bf-b790-74b737a99935 · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.854863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.224012Z digest=sha256:7ae6961b8dafabec53c13b45e49ef67328b6b7ffad1837af26b81ee4729f6786

Observation ede8583e-037f-423a-b501-d7b53fdacb4f · outbound

This paper cites GPT-4o System Card.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.228732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.228732Z digest=sha256:bad0d31f6375436d06c199734cc92c78dbb2238e70ce71dc5a9db53b7595af54

Observation 2895fd57-2cb0-4fc0-aa7b-2b41d85103f9 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.233665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.233665Z digest=sha256:4423717fb49a80c7a2f42db3b2457419c21762ca917e68bb3398d078e80fcd27

Observation 107088b4-3ab2-4a67-a77a-c34de216eba4 · outbound

This paper cites Zoomtrack: Target-aware non-uniform resizing for efficient visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Zoomtrack: Target-aware non-uniform resizing for efficient visual tracking

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.839119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.238682Z digest=sha256:798f0223ad00ef477a09297aa8ffe0f4f11b398f112dea7251dc06d95a7393ea

Observation 78eb8ad7-c238-4d46-b3f6-3162a9256197 · outbound

This paper cites Multi- modal data fusion: an overview of methods, challenges, and prospects.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Multi- modal data fusion: an overview of methods, challenges, and prospects

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.822966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.243349Z digest=sha256:4713b3dba6ed82d62b5a3ca66c0ab8a14d53e9977dad76a43bb6f3459ef1b9c5

Observation 69bfe300-2ea0-496a-8c69-560bc7703344 · outbound

This paper cites Cornernet: Detecting objects as paired keypoints.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Cornernet: Detecting objects as paired keypoints

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.807538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.247946Z digest=sha256:3644bd2932b4d6c9c16b0fddbccfe4ebd1b5f6600428ac65313040113e8c3898

Observation fb452ae2-f0aa-4952-9742-fc8afed9e0b4 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.252699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.252699Z digest=sha256:9c7ea2b72b31ae1c593ec3ffd4b481d7b398f3d3d8e8edd937bab9ffa617f579

Observation bc7645a7-9796-4abd-9f0f-c3940de0187d · outbound

This paper cites SiamRPN++: Evolution of siamese visual tracking with very deep networks.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking SiamRPN++: Evolution of siamese visual tracking with very deep networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.792291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.257482Z digest=sha256:2a2ad8d0e106f1cbc05b6a1ff10c4b0cccfed45207abefd26b977f118fa968c1

Observation 4a85f1d4-547e-4ebe-bac4-e6fe7aa1de90 · outbound

This paper cites Dtllm-vlt: Diverse text generation for visual language tracking based on llm.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Dtllm-vlt: Diverse text generation for visual language tracking based on llm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.776586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.262034Z digest=sha256:45c84cf24b6a74200095cfb6606542bd7f0b8b2c01f294ab94e784dfd2858a48

Observation 9ad79807-54ae-4d3b-bb8e-c58ee2ca8bab · outbound

This paper cites DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.267463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.267463Z digest=sha256:1ff6cb3ca0b62d47ef92ff9d8ae4bd5822360a91a5811eaaea162f0b47790ea4

Observation c6b60fe2-130c-4dd0-8cfb-129f4d97ab6f · outbound

This paper cites How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.272754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.272754Z digest=sha256:4cf0da2d86a5b508091d3b0be6ee6e6e82ef8b5a2c4053fca95fd103b6ba1923

Observation 80bed04c-ce88-4da5-a30e-3d7518189832 · outbound

This paper cites Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.277681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.277681Z digest=sha256:a91677100c561d58a8cf0baae31c6d3f25422275a6a79962f08a9283172bb4e0

Observation eb33eb2f-f8c4-4151-a44f-ee6ca2229e8d · outbound

This paper cites Cross- modal target retrieval for tracking by natural language.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Cross- modal target retrieval for tracking by natural language

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.759192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.282504Z digest=sha256:9a1c5334ad5fdc679ed923f51ac1aa246ca60c98ab8f6508a502acca65627e0f

Observation f43e334b-00d2-45f0-a124-e26be12a0705 · outbound

This paper cites Tracking by natural language specification.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Tracking by natural language specification

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.741442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.286912Z digest=sha256:c69a4fa12a2caf73272cd6490d3785a52e2b3bfa1993e1f54a281844e05bbcf5

Observation a1c4ffdb-d657-4d1d-8c21-1fbbd62da4e5 · outbound

This paper cites Tracking meets lora: Faster training, larger model, stronger performance.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Tracking meets lora: Faster training, larger model, stronger performance

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.725689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.291623Z digest=sha256:0145af11ef8bf3ee10a4a9e8b1b92167266e4f9a18036aed87d6c3754804dc3b

Observation 6714be6a-de44-4d61-903c-ca2ce8887242 · outbound

This paper cites VMBench: A Benchmark for Perception-Aligned Video Motion Generation.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking VMBench: A Benchmark for Perception-Aligned Video Motion Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.296454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.296454Z digest=sha256:3ddd369eb33068f3cfcc3fc055cccbdb9b5dcf373f700f618b711eaf773f5915

Observation 07a508e4-5330-4201-aa48-0450b0c8fce8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.301341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.301341Z digest=sha256:e471d475e7f68d1214dbaca8037ee1e6e79c838bdd38788300dc0fba6f675dc8

Observation 11fd7396-2961-4a85-864e-47a3fd4a0165 · outbound

This paper cites Decoupled Weight Decay Regularization.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Decoupled Weight Decay Regularization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.306545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.306545Z digest=sha256:262f1ea4d6c90e9132f9296d5eaf8326ccd09b040f6783286837941accf6b391

Observation 2c966cb9-8a2a-48eb-94a5-7cbeec8f84cc · outbound

This paper cites Tracking by natural language specification with long short-term context decoupling.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Tracking by natural language specification with long short-term context decoupling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.710708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.311253Z digest=sha256:f5132747a58e723b5395e839af7943a1913d40f9ea57ed1d834800cbd1564c9f

Observation 6f8b7738-ada1-479a-9ec6-523f38a3787b · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.316004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.316004Z digest=sha256:578a21e3fddfd2ad2973f9d481f4e6143e0b8586cb05aa967c8a66f906fa6974

Observation d5830c71-77d9-45e3-baba-7c2d3e72a7ea · outbound

This paper cites Unifying visual and vision-language tracking via contrastive learning.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Unifying visual and vision-language tracking via contrastive learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.692889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.320855Z digest=sha256:c58fd2be068e84a3cc78becccaa2ec61add78600eb611e6143866261ddd70d8b

Observation 0779b1d5-73b4-46c5-8549-6df17a2bbcf1 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Generation and comprehension of unambiguous object descriptions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.675570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.325482Z digest=sha256:10048cbbb7a33e7f5012d5e94a097e856ebf896262975b1dc5a2156ab2fbd8ae

Observation 79976080-0973-49b3-a477-4f80fd4cab12 · outbound

This paper cites Textual tokens classification for multi-modal alignment in vision-language tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Textual tokens classification for multi-modal alignment in vision-language tracking

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.656749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.330545Z digest=sha256:e624ccb252bc6c66c46124f1b8dc76469f72eab1726e033dc9fcd0ebbbd7fbf5

Observation fe42269f-44fb-45ca-ac5c-ffc338187d76 · outbound

This paper cites Learning target candidate association to keep track of what not to track.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Learning target candidate association to keep track of what not to track

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.639875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.336176Z digest=sha256:c9e699a6c89fbb316fb46f32452310353d2b9db61b6c4d6403898ebbc2f81fdd

Observation 4449fdf6-1f5f-4cbe-b041-ec613717c58d · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.624769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.341584Z digest=sha256:d921b61f67c7eef141c36dadff7605f1f6f4a8f773831b04afdf06ec972bc72d

Observation 79a5574f-e0ac-48cf-aab0-dda64beee430 · outbound

This paper cites Vast- track: Vast category visual object tracking.Advances in Neu- ral Information Processing Systems , 37:130797–130818,.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Vast- track: Vast category visual object tracking.Advances in Neu- ral Information Processing Systems , 37:130797–130818,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.609754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.346874Z digest=sha256:7e663efd6642025654d6e259f7e66a3f9f10e2080b70e5d3e3fc9025a3c0556f

Observation c5af3050-d178-473f-89b7-5e73428fbd17 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.593658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.352561Z digest=sha256:8b956751bf8399706811e2d91a814b5ae14affa930bf498b76c69d5057e234a1

Observation b391ec34-bb7a-4b16-bc38-97bd873bf123 · outbound

This paper cites Generalized in- tersection over union: A metric and a loss for bounding box regression.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Generalized in- tersection over union: A metric and a loss for bounding box regression

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.357235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.357235Z digest=sha256:926795e26cabe42bd0b9396c94182bb615fd9e161e1644e80a7eaad03aa11a89

Observation bc8af11e-9e20-48ea-968c-7be50078ca59 · outbound

This paper cites Generating semantically precise scene graphs from textual descriptions for improved image retrieval.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Generating semantically precise scene graphs from textual descriptions for improved image retrieval

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.566625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.362928Z digest=sha256:44d921b3e4713abe202693f17c64b392de5c0bc2d64bd5a188d25358e1f854ca

Observation 5c1f058e-025a-4884-80c9-a491e4c32882 · outbound

This paper cites Context-Aware Integration of Language and Visual References for Natural Language Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Context-Aware Integration of Language and Visual References for Natural Language Tracking

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.367754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.367754Z digest=sha256:9e952649bc394392336d841e6f403ab135b23126fc31a8da83aefb641a22d9c7

Observation e6bf2372-4fc2-4361-bd35-f84ed6e21c25 · outbound

This paper cites Explicit Visual Prompts for Visual Object Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Explicit Visual Prompts for Visual Object Tracking

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.373163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.373163Z digest=sha256:770b862690955cca1ba8b780295036383d8e948c69408387abae8e1fab70f486

Observation e7c45bfe-6790-4790-9e82-67358170e76c · outbound

This paper cites Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Chat- tracker: Enhancing visual tracking performance via chatting with multimodal large language model

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.550322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.378218Z digest=sha256:7b88d95f00815064efb5819abecaaac6805af246d8a06d84b7d95a2fec71c564

Observation 14d4b30a-8a65-4337-9b36-d4f5e94fa0e5 · outbound

This paper cites What makes for good views for contrastive learning? Advances in neural informa- tion processing systems, 33:6827–6839, 2020.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking What makes for good views for contrastive learning? Advances in neural informa- tion processing systems, 33:6827–6839, 2020

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.535026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.382961Z digest=sha256:4755bb1967c741b66f806cf64c5374335fca48f475733f80b90856602ba27be8

Observation bbbed7b5-4ef0-438f-a903-a55d946ff226 · outbound

This paper cites Fast-itpn: Integrally pre- trained transformer pyramid network with token migration.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Fast-itpn: Integrally pre- trained transformer pyramid network with token migration

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.518967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.387723Z digest=sha256:8c00c0a4c3aa30435e20a4dd911c5ad4c9da83d9d934a3fd3ea197a2f5ac787e

Observation dbd7e95f-68fd-44d6-b339-179c6817665e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking LLaMA: Open and Efficient Foundation Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.392906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.392906Z digest=sha256:54390f422bee9a69cf1a1dd4d5c86072860ac4c4c9f2975f061a84cb90a3bb0d

Observation 8bcf2dbb-950a-4d23-8bdf-2636f4cc4953 · outbound

This paper cites Attention is all you need.Proceedings of the Ad- vances in Neural Information Processing Systems, 30, 2017.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Attention is all you need.Proceedings of the Ad- vances in Neural Information Processing Systems, 30, 2017

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.503791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.398258Z digest=sha256:88a9cab73eb9b2bcdcdb5bb83c3d547539501560a4bc14c3302b05b9a8736890

Observation 97147d00-61c6-4eb4-b74e-8395c75fa157 · outbound

This paper cites Temporal adaptive rgbt tracking with modality prompt.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Temporal adaptive rgbt tracking with modality prompt

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.488450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.403665Z digest=sha256:de7229abd4cc814a5e6f05e702480cf7006639ed555aa595ce01337508214951

Observation 387deadd-f8dd-4f07-87bb-129b691ca116 · outbound

This paper cites Transformer meets tracker: Exploiting temporal context for robust visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Transformer meets tracker: Exploiting temporal context for robust visual tracking

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.472658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.408214Z digest=sha256:c2470562f725ad23150a37cef197b67cc19485155c7f4498931a702343237aae

Observation 72d0a56c-f80d-4747-a621-668380a7ca51 · outbound

This paper cites Unified transformer with isomorphic branches for natural language tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Unified transformer with isomorphic branches for natural language tracking

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.457089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.413111Z digest=sha256:84d5b9c9c2f3ceba2836c765b482e84cd4b3c75b0a14fce1aff3ce7b80f50f42

Observation 9560beba-56ea-42d6-9f95-9808deaafd35 · outbound

This paper cites Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.417726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.417726Z digest=sha256:d5b183c062dcf921b7f5abd6aede0dabd2d79c25f1e7574ca2aae1a4e99078ef

Observation 67692cc2-0fe9-4b8c-bd99-3655fb1e7132 · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.441493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.422744Z digest=sha256:523209e79a5a54d1a7ecd7fd081a589b3f7254d1275ea2c55604204889ce864a

Observation c2f82135-bd43-4b95-a14a-0743dd896159 · outbound

This paper cites Autoregressive visual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Autoregressive visual tracking

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.425576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.427440Z digest=sha256:753433e3d7cb0b12648aea6e2d9fe95b9b3da54ad7e7b57348262f50a57ce3f6

Observation 8573aae1-3df1-4911-aebe-395c538b691b · outbound

This paper cites Dropmae: Masked autoen- coders with spatial-attention dropout for tracking tasks.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Dropmae: Masked autoen- coders with spatial-attention dropout for tracking tasks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.411229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.432096Z digest=sha256:fa62aa1cccff1091845aea19a8bd7d07ef60faec9ffc70d33a9425623b522a06

Observation 5ec8bf32-ce83-46d2-ad31-8bfc46afd404 · outbound

This paper cites Object track- ing benchmark.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Object track- ing benchmark

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.395947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.436818Z digest=sha256:6cde00080189c8712353cba94f6935338bd9dcea1e917163886dbaf2f6cde79b

Observation f56365bd-fdf5-45de-87c2-ec0588407938 · outbound

This paper cites Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.441550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.441550Z digest=sha256:d18500672544b87d8acd51af8c08f3a99ecb4692bb75aa47d06b0ac393c93904

Observation 08292d3a-6347-4310-8ec3-3d7c5a70afef · outbound

This paper cites Less is more: Token context-aware learning for object tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Less is more: Token context-aware learning for object tracking

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.380626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.446417Z digest=sha256:39333e32c76e0063a42d102130aa241a63eedd99db84e598eae9fe611108e3f9

Observation f42128ac-9417-4ea7-a6bc-e4398784bd5c · outbound

This paper cites Learning spatio-temporal transformer for vi- sual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Learning spatio-temporal transformer for vi- sual tracking

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.364736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.451324Z digest=sha256:a25422ffb244103f043b550b23b93287c5fa4843918e1c287e8de4ecb7bbeabe

Observation 04bb95d1-433f-459c-9f04-82d0e2e659da · outbound

This paper cites Learning spatio-temporal transformer for vi- sual tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Learning spatio-temporal transformer for vi- sual tracking

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.348240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.455733Z digest=sha256:d5cac0a519ced29f56b464cba3f0ddd10b5a3a08b055ee6ef7bc7be8055b5894

Observation 0b4b19a0-230e-43dd-b401-15f565d5fbdd · outbound

This paper cites Foreground-background distribution mod- eling transformer for visual object tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Foreground-background distribution mod- eling transformer for visual object tracking

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.333030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.460077Z digest=sha256:618800b425d8edf84cb1ffea728f598e7bffc47e40df2925596b5d0663017bae

Observation 7d2b5daf-8b45-4459-b1e3-7513d6eb3080 · outbound

This paper cites Grounding-tracking-integration.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Grounding-tracking-integration

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.318160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.464646Z digest=sha256:80739e5ca3524834c5897e16430b9ddb13f93bbb3e734c72c2a69fc61e004a52

Observation 014d38a9-b19e-4f0a-9053-f91b3f049817 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.302570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.469358Z digest=sha256:d1c470aa3d6a7bf8e2b87c9a44f4e45d8984840a2079fde3f5cc632743199fa2

Observation e748ad6f-bcd2-49f6-bea1-79a96bb6bf03 · outbound

This paper cites All in one: Exploring uni- fied vision-language tracking with multi-modal alignment.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking All in one: Exploring uni- fied vision-language tracking with multi-modal alignment

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.287233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.473993Z digest=sha256:44f2ba69d9ddd11617e2ed03b407806daee79db56e0475fea7ae2f21f36da558

Observation d0b53a5a-ebf2-45c5-925c-08c7161b930e · outbound

This paper cites Beyond accuracy: Tracking more like human via visual search.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Beyond accuracy: Tracking more like human via visual search

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.270441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.478672Z digest=sha256:47e746dfaa2ca9c5607838bd05d646a7f37e5b1561f108db519a6c15a9607452

Observation 3124d896-a416-495f-b9a4-0e8185b6b4d8 · outbound

This paper cites One-stream stepwise decreas- ing for vision-language tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking One-stream stepwise decreas- ing for vision-language tracking

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.255495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.483392Z digest=sha256:87778728396ce78defb380dbf2de5d07cb78fe81087321ebbd45805fa0159472

Observation ef1aafe4-0de8-4373-8969-a42c6682dad8 · outbound

This paper cites Hivit: A simpler and more efficient design of hierarchical vision transformer.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Hivit: A simpler and more efficient design of hierarchical vision transformer

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.240458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.488278Z digest=sha256:967371b70142bb02b3727e04f30352a0e907af8fa9448b7f5bed3a37fd473f0f

Observation 57cc62c4-9c13-4106-8d12-6940d28778dd · outbound

This paper cites Transformer vision-language tracking via proxy token guided cross-modal fusion.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Transformer vision-language tracking via proxy token guided cross-modal fusion

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.225498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.492826Z digest=sha256:e93e9da215578cce5fd89eb90a325e1050896389a18c0fff86c06ea644b7012b

Observation 6194473d-e370-4145-ad3e-3d187e5d9f80 · outbound

This paper cites Biodrone: A bionic drone-based single object tracking benchmark for robust vision.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Biodrone: A bionic drone-based single object tracking benchmark for robust vision

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.210101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.497585Z digest=sha256:019085d1dc2cc30f6fb935ba4f38cd0d2987e254dcd5fc0ab1ec521ee99661d9

Observation c4381f3a-f5b6-4ff2-9db2-39b0618eefa4 · outbound

This paper cites Leveraging local and global cues for visual tracking via parallel interaction network.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Leveraging local and global cues for visual tracking via parallel interaction network

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.502187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.502187Z digest=sha256:c9c86f55ca7209531eed148c45fcdddc7c3993e66edc7e597c336050ece4c5ad

Observation d76db66e-8426-42c0-a158-608c8bf3fb62 · outbound

This paper cites Towards unified token learn- ing for vision-language tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Towards unified token learn- ing for vision-language tracking

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.183016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.506485Z digest=sha256:ab5730d26adb9513dc19c02096d3cac32c5f1d45b75dde69554cddc2df5d37b7

Observation 389d8f8b-d52d-450d-99a9-6e2cc63643c2 · outbound

This paper cites ODTrack: Online Dense Temporal Token Learning for Visual Tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking ODTrack: Online Dense Temporal Token Learning for Visual Tracking

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.511264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.511264Z digest=sha256:87baa52f794bc05a476a80054ea19bc796fa74500c36f6a8a9918ee866f20a96

Observation b38f2ec4-7338-43a7-9a17-e75e1816ef5f · outbound

This paper cites Decoupled spatio-temporal consistency learning for self-supervised tracking.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Decoupled spatio-temporal consistency learning for self-supervised tracking

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.165065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.516080Z digest=sha256:78cff1344ca46362863008fe33fd00f046b9d204dcefb93027d47d7a041254de

Observation e37c8ada-10ed-47b5-8fdd-2a2c58282a59 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Conditional prompt learning for vision-language mod- els

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.520680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.520680Z digest=sha256:93fcb4d2ab8b9aaf5c57dd6b68fc675f09250b925b9bc2c304aee2849bc8717e

Observation 2f21eb56-ac82-445b-b26e-0b05d4223b36 · outbound

This paper cites Joint vi- sual grounding and tracking with natural language specifica- tion.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking Joint vi- sual grounding and tracking with natural language specifica- tion

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:25.525240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:25.525240Z digest=sha256:783c09f94a5601fab4b48c379bf82f103a45856dc079b11a0c2498f1583ed7fc

Observation 218710c5-93e8-4ae6-aa17-dbbad206feee · outbound

This paper cites the ironman in red flying in the sky.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking the ironman in red flying in the sky

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.112817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.536019Z digest=sha256:bae2afa7eb4c79e085c1ae979a7b58fb019efa60eb1622e1ec27b52c676aafbf

Observation d98f2fbd-536a-4b66-a8b2-298adeaa7136 · outbound

This paper cites plane”. In the corresponding Attl heatmap, the target word “plane.

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking plane”. In the corresponding Attl heatmap, the target word “plane

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:00:26.128639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:00:25.530220Z digest=sha256:acca145a95db2d901f38add8348e33ed7f93bff7d6b87e7b9eb82a1c5cc3b892

Pith citing papers

No inbound Pith citation observations are available.