Pith. sign in

Paper Citation Record · LEDGER

Adaptive Perception for Unified Visual Multi-modal Object Tracking

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2502.06583.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06583 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:01:38.486514Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:20:07.965249Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:20:13.646436Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact6
  • verified fuzzy44
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae2b2bb3-95cb-4da3-8082-80449fcb96cb · outbound

This paper cites Backbone is all your need: A simplified architecture for visual object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Backbone is all your need: A simplified architecture for visual object tracking,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.195544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.256848Z digest=sha256:309910618ed6bce6fb47d7d35251f5de1ba97d99ae8d0d80a2b0755acf01baa6

Observation 6d37118a-2ee4-416a-bb42-55bddf870ca2 · outbound

This paper cites Mixformer: End-to-end tracking with iterative mixed attention,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Mixformer: End-to-end tracking with iterative mixed attention,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.186036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.261223Z digest=sha256:7fe7c0db35463f4867e943b3adafa4dba72453d3b451d1f3dfe459901143da8e

Observation 774a935a-5458-49a9-9a30-c8e726e2e8e1 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Joint feature learning and relation modeling for tracking: A one-stream framework,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.176512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.264760Z digest=sha256:84fe734c42192c51b56a1e970f42f813bca94acdcea76242885da00d831b0f04

Observation c616e41c-04eb-485c-b31e-aaf76c4797ad · outbound

This paper cites Seqtrack: Sequence to sequence learning for visual object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Seqtrack: Sequence to sequence learning for visual object tracking,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.166811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.268681Z digest=sha256:a0eb1cb62b7a8b2547b05874886ea914ad109d572644f4d8e8216b4299e91a30

Observation c4057e42-c032-4133-9eea-38776be93d97 · outbound

This paper cites Swintrack: A simple and strong baseline for transformer tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Swintrack: A simple and strong baseline for transformer tracking,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.156677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.272577Z digest=sha256:513a619a651b4e6a4c01b682616d29daeb4d1cd63052debf9e2bcefee19a6831

Observation f78f198c-5177-40cb-8af0-a877820e7249 · outbound

This paper cites Autoregressive visual tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Autoregressive visual tracking,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.146390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.276091Z digest=sha256:8e74725f29c0f54c00133159abb7e1705cfbf08c8f242da04cd3b9f17ddcf9d2

Observation 8108e1d3-9c7a-4079-b285-da49a945e1df · outbound

This paper cites Visual prompt multi- modal tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Visual prompt multi- modal tracking,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.136277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.280162Z digest=sha256:da6e7300f5c43e26de6d031279cc3d2bef0ce1dc78c1aaf31ab08df99ee09165

Observation f1e2e9a1-1af7-4cef-ac7c-22d2c45659a9 · outbound

This paper cites Single-Model and Any-Modality for Video Object Tracking.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Single-Model and Any-Modality for Video Object Tracking

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.664936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.284479Z digest=sha256:72c170ba1e00f40f2a9183e0e6afd98a5c55af7188fb4c40b315781e7bb4340c

Observation e60ec162-1789-4f98-bd54-6dded71c3a08 · outbound

This paper cites Robust Tracking via Mamba-based Context-aware Token Learning.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Robust Tracking via Mamba-based Context-aware Token Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.289301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.289301Z digest=sha256:6cac43b634a9f0417c82b62fb5ec2df76605c4d53c2866f9b258eedb22a87fc1

Observation dd84f229-f4bd-42ef-aead-49efbf9852fd · outbound

This paper cites Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.293522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.293522Z digest=sha256:0a9ce68211025c960a8955cc508e7870e33aeb3a130bd121f4d2ea3977759fac

Observation 58b423eb-40aa-447b-b708-d7faf8ac53cf · outbound

This paper cites Curricular contrastive regularization for physics-aware single image dehazing,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Curricular contrastive regularization for physics-aware single image dehazing,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.125360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.297829Z digest=sha256:c5128d2f1ab93aabbc6564f8336e2dc1fe3884da52a5140c9b2afce3f967d550

Observation 3127b424-4be4-4df1-a08f-b3163a397d14 · outbound

This paper cites Dynamic group difference coding based on thermal infrared face image for fever screening,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Dynamic group difference coding based on thermal infrared face image for fever screening,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.114649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.301951Z digest=sha256:352a6345c31aa0e2fcca8004cbc7fb9529d6d259ca39d1ebf30c3acb744745c8

Observation 3abafafd-98ca-473b-8504-fccd1ef64521 · outbound

This paper cites Few-shot learning with long- tailed labels,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Few-shot learning with long- tailed labels,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.104340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.305685Z digest=sha256:b16593cf88d433a181692c179f7eefee94ed474466432b8ad23071aa35f00ae0

Observation a59b6c1b-ce2c-4e4d-8279-c96f6390d2eb · outbound

This paper cites SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.093728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.309654Z digest=sha256:87b15b8df43ec2e95b5b49fff1add445fa2c41ec56b5af34a2110f919358f9e1

Observation ec604d4e-9319-411b-95c1-ff689e4564bf · outbound

This paper cites An emotion recognition method based on eye movement and audiovisual features in mooc learning environment,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking An emotion recognition method based on eye movement and audiovisual features in mooc learning environment,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.083502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.313467Z digest=sha256:c65e6970d13f0fe2119d2343528fd8a0e46312c361ca065dc3ea8f0007ab843a

Observation 8c5c0319-ada0-4e34-926c-275af4dc4392 · outbound

This paper cites 3d- guided multi-feature semantic enhancement network for person re-id,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking 3d- guided multi-feature semantic enhancement network for person re-id,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.072896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.317387Z digest=sha256:9c35d764d4e2e532507da16ec85687e02df53d5c077d5d2940dbdf5add8b64bd

Observation 8bc50acc-e8fb-43fc-99a0-b31dba64f507 · outbound

This paper cites Multi-branch enhanced discriminative network for vehicle re-identification,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Multi-branch enhanced discriminative network for vehicle re-identification,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.061548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.321291Z digest=sha256:bd86d29fcaeeab5f705fd48bca54fbd3f63626bec2271fdd2eabb74600e7f117

Observation 75e1fd3c-61ea-4eff-8542-49ef97f0067d · outbound

This paper cites Guided Real Image Dehazing using YCbCr Color Space.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Guided Real Image Dehazing using YCbCr Color Space

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.626079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.325098Z digest=sha256:2b7d0b8841f9d0926c23e57df53ed6d7b406a6ac428dfddddcc6ad56dae33e1d

Observation b0a3ff3b-19c4-438f-97ed-c9fe611ac37c · outbound

This paper cites OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning.

Adaptive Perception for Unified Visual Multi-modal Object Tracking OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.612183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.329265Z digest=sha256:6f2a6972059c2cec4c2286486209c85344a5fe49eb8df1fca13a427ad083ba3b

Observation b4afefca-df2a-48da-983e-0f1579b8a5fa · outbound

This paper cites Prompting for multi-modal tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Prompting for multi-modal tracking,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.050806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.333793Z digest=sha256:04a84116f81c7c315fa61e13ad3d6b235ebbb7a900e02af2e823b7305c6ccb5b

Observation 586c5bbc-acf0-469d-a51b-dce848553d4e · outbound

This paper cites SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking.

Adaptive Perception for Unified Visual Multi-modal Object Tracking SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.596444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.338768Z digest=sha256:a7518ad1f685764a1b8e2d6a9fb3b0f62fbe5102e022e427f0d80a8aed9a85ff

Observation 2bc63b7a-a1ee-4a72-9de5-d6e6fbe98bda · outbound

This paper cites Bridging search region interaction with template for RGB-T tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Bridging search region interaction with template for RGB-T tracking,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.343304Z digest=sha256:68d19ea37247805e2f3859c9cca754f6a4ec224044fec483d1849ceb10ccb330

Observation 796e2777-6c9d-4570-bfd0-15090e63e4a1 · outbound

This paper cites Bi-directional adapter for multi- modal tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Bi-directional adapter for multi- modal tracking,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.029978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.347554Z digest=sha256:a39b10cc0a042e3f4d552d010b4cb9960354938bef4bec540d78bad43bacb6a7

Observation 3ea5e07d-5d7e-4700-92e8-a06ce244ce02 · outbound

This paper cites Spiking transformers for event-based single object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Spiking transformers for event-based single object tracking,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.020710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.351534Z digest=sha256:f1c1dcfd6f87d6bcf682942f370234625fb84d352fecc4e48a87738abc66b298

Observation 8b1f559c-8ee2-4d90-bea0-50eae7c83025 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Lasot: A high-quality benchmark for large-scale single object tracking,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:39.009603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.355559Z digest=sha256:26205bc97d4f6eafd6365b09d20547cc9f5cec1dc260d644b83239818cb5eb20

Observation c3d2d9ec-e421-4119-b0ba-26dbeb35d237 · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Got-10k: A large high-diversity benchmark for generic object tracking in the wild,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.998252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.359442Z digest=sha256:17bab86ae463c39c5879a7a4de69060d452240ccfcd3f4607d2ec86d69eb1a24

Observation 194a8d21-c048-4cd1-8f62-cf8c9c9d92f2 · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Trackingnet: A large-scale dataset and benchmark for object tracking in the wild,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.986870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.363352Z digest=sha256:94d1bb137ead4e771d0b3d4c62e3a5a4a10065ccf5877099754868f1d62ef794

Observation a73088f9-0cc1-41ce-8626-4b00a8f9b49c · outbound

This paper cites Lasher: A large-scale high-diversity benchmark for RGBT tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Lasher: A large-scale high-diversity benchmark for RGBT tracking,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.973857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.367242Z digest=sha256:12e344ae030ec914234ef35205c59e0551e172fb1bbaca8b716944e354d38b47

Observation 96f5ea52-3207-43f2-8308-3c3de07a75ce · outbound

This paper cites RGB-T object tracking: Benchmark and baseline,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking RGB-T object tracking: Benchmark and baseline,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.961719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.371104Z digest=sha256:aac4043b5c25cb27d941e8e7fe1be4ee54a5241cfd0ff01d4e1b05e55ad3fb88

Observation ab07304f-7740-4ad3-ab8b-7ecb33720098 · outbound

This paper cites Depthtrack: Unveiling the power of rgbd tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Depthtrack: Unveiling the power of rgbd tracking,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.949221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.374858Z digest=sha256:a253043b3f64a85cadb34e7ac4f446750cb43810ad90e2263c8378764b7d73d4

Observation 3aa9dbad-8a55-40b9-93fd-6f09d8839568 · outbound

This paper cites The visual object tracking vot2015 challenge results,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking The visual object tracking vot2015 challenge results,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.936593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.378540Z digest=sha256:644f0984b68444a6fd405ba69c027a007ad2148cee9135271b35b79b64f0db14

Observation 88f7be3a-626a-4d16-8af3-b4b9962248e6 · outbound

This paper cites Visevent: Reliable object tracking via collaboration of frame and event flows,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Visevent: Reliable object tracking via collaboration of frame and event flows,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.924562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.382631Z digest=sha256:1dfdd002fa51d3d1b526cf9ff4a80b16ff995d664309d4e69bc85d1f02974a9d

Observation 6115b44c-0ead-413f-9f68-3ca71ac655d3 · outbound

This paper cites Unified-io: A unified model for vision, language, and multi-modal tasks,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Unified-io: A unified model for vision, language, and multi-modal tasks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.386671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.386671Z digest=sha256:9a9f089ea9329c130d97975fec8aebc105b0a00991d319e2d117c94837cb2500

Observation 59d77813-4c96-4129-8b07-9a6faae8009e · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Imagebind: One embedding space to bind them all,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.904614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.390455Z digest=sha256:4e6bae41e5017ce47eb63e40b720c8baafd57da89a91357f5383f2dd090e647c

Observation a63fec7a-da80-4305-abd7-d572bb5a9d0d · outbound

This paper cites MUTEX: Learning Unified Policies from Multimodal Task Specifications.

Adaptive Perception for Unified Visual Multi-modal Object Tracking MUTEX: Learning Unified Policies from Multimodal Task Specifications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.393426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.393426Z digest=sha256:acddc198466ba5af0f3585ecaf654c58d86601b96992920657a07eb6ac291827

Observation 766d9f0a-ffd3-4334-8fdf-9469ea52f920 · outbound

This paper cites Siamese Vision Transformers are Scalable Audio-visual Learners.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Siamese Vision Transformers are Scalable Audio-visual Learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.397539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.397539Z digest=sha256:053b240a09b4130c079a2c4d5249aac76f1224ca122abe0a86b6247685d354f1

Observation f2bfc63a-4f65-4813-bcf2-2e19670f1d95 · outbound

This paper cites A unified audio-visual learning framework for localization, separation, and recognition,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking A unified audio-visual learning framework for localization, separation, and recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.891981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.401184Z digest=sha256:194317b0774859b41d00017d129323e93ce5f7fd7afd5a43d9454d15d6d0a2f2

Observation 0c1a4fa6-a92c-4773-85e1-692bb4cdda91 · outbound

This paper cites Learning visual representation from modality-shared contrastive language-image pre-training,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Learning visual representation from modality-shared contrastive language-image pre-training,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.878140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.404589Z digest=sha256:19dfff51c36b9ad1a52f5ccb77bfc05da30906917486a95bf3db70a2396a5d04

Observation cdfd3f0b-d730-4075-b0d0-e834ec74cfb6 · outbound

This paper cites Decoupled weight decay regularization,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Decoupled weight decay regularization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.866174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.407603Z digest=sha256:f3af216ba270a0dc71e20769b5042691b7c12569ab248917c38c9ce4179d8bac

Observation 5942d894-1489-408f-988d-916e45a4d3db · outbound

This paper cites Weighted sparse representa- tion regularized graph learning for rgb-t object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Weighted sparse representa- tion regularized graph learning for rgb-t object tracking,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.855104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.410684Z digest=sha256:5e8889a490da27683461878ae511308e95fb660f651c44a681d6be6f77979a34

Observation 0371c9fe-b5e7-4fc6-856f-357ab1008aa2 · outbound

This paper cites Generative-based fusion mechanism for multi-modal tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Generative-based fusion mechanism for multi-modal tracking,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.845254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.413706Z digest=sha256:35ab92cd7152873b384c58b4a9b55c7b8280c368843a22c99196a74c3256d37c

Observation 59dbba3c-c935-4ac3-b145-53800b2d829f · outbound

This paper cites Transformer tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Transformer tracking,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.835252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.416861Z digest=sha256:b51a16f73f588b62f6ac7a29182d410f0092f4291412546f24f9c4bc050a0d28

Observation e0c0914f-3e9b-45e1-8846-ba04d4acff3b · outbound

This paper cites Learning spatio-temporal transformer for visual tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Learning spatio-temporal transformer for visual tracking,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.420525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.420525Z digest=sha256:e088b73084297fb13d4e9ad5dd36a0e194912abae66c287b2fbda954c94af3cb

Observation e4ff1443-2a31-4fbf-be4a-4dacc6dfc7e1 · outbound

This paper cites Aiatrack: Attention in attention for transformer visual tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Aiatrack: Attention in attention for transformer visual tracking,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.424619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.424619Z digest=sha256:f31819bd8c7932f6d1c208f628ec37b55f2a2c2e031db4b91490b9fee7933568

Observation 8e89baee-71c9-431c-8944-3a5acd28a9be · outbound

This paper cites Rgbd1k: A large-scale dataset and benchmark for rgb-d object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Rgbd1k: A large-scale dataset and benchmark for rgb-d object tracking,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.812467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.428469Z digest=sha256:c63b18b7c8146ef41942a56790e0634610b57e89ee2426165cedb1aab0de4d6f

Observation b0677cfa-d238-45bc-aacf-7fea2b8c6926 · outbound

This paper cites Focal loss for dense object detection,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Focal loss for dense object detection,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.432244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.432244Z digest=sha256:31618c5ddd9612662477ca8402d77474e3d17cb71c37d6b14399714f832191d6

Observation 823456f3-4507-4856-9a7e-96e38f8ab8fb · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Generalized intersection over union: A metric and a loss for bounding box regression,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.794074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.436214Z digest=sha256:a03a5012b9d21064cf4decd89c56c7501b114cc01bf8e5211891ad9a59425108

Observation daa29079-8457-4bd8-9f6c-50806d158431 · outbound

This paper cites Depthtrack: Unveiling the power of RGBD tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Depthtrack: Unveiling the power of RGBD tracking,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.782160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.440102Z digest=sha256:ad489fa7961a7d6e248ff0658efebf5695f53aa8943fa3d89afe3664c89530be

Observation d5690a09-d456-4f20-ab09-aecd738e4914 · outbound

This paper cites Transformer tracking via frequency fusion,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Transformer tracking via frequency fusion,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.770507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.443660Z digest=sha256:7de31f70aa8fc7e90bbd99cb626be1b0a7d29c880b8860d2dd3e8acfb34fe265

Observation 33dce51f-e2f7-4125-a47a-ad004815f6a1 · outbound

This paper cites Multiple source domain adaptation for multiple object tracking in satellite video,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Multiple source domain adaptation for multiple object tracking in satellite video,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.758488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.447390Z digest=sha256:4b280e97eeaf6b03c889deb3a1a97c5be837607e2bb7406ccead50ab46f18d06

Observation 7dd02a68-a0b1-451a-9eea-7850cf970597 · outbound

This paper cites Explicit visual prompts for visual object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Explicit visual prompts for visual object tracking,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.747050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.451209Z digest=sha256:3ba0c2a16c57e395068839d4f86e150e59f7f5fd2d012cbd2e4a783607d54e9f

Observation 6b9fe4db-d580-487f-96d2-73ff40e0dd65 · outbound

This paper cites Autoregressive queries for adaptive tracking with spatio-temporal trans- formers,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Autoregressive queries for adaptive tracking with spatio-temporal trans- formers,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.455226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.455226Z digest=sha256:ad9af7b296fe4abaad43f2e7753c79207c0f06983a49194ba0c2be572dfd78a8

Observation fcdd4def-e275-4da9-9e85-5a9eeb715089 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Learning transferable visual models from natural language supervision,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.726920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.458954Z digest=sha256:a51a1472c9e778dfd7b9202f52b872e6014c4ba2c79940a84aba3fb68dbfc939

Observation 83c21fd7-bdef-41bb-8554-d16961611ed2 · outbound

This paper cites Towards modalities correlation for rgb-t tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Towards modalities correlation for rgb-t tracking,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.713885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.462662Z digest=sha256:0ef39380a20dace9970981f64ff5e2963ee193dd34e6076a4b1e45b70a298756

Observation d9553625-d158-4c93-a018-9920580c311d · outbound

This paper cites RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.702132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.466373Z digest=sha256:0fa7e722e9a6e932c899a9eb14d1571b1c3fffbf63022b6f526115acfb292514

Observation 64227ba3-8b07-4445-96dc-68280b42ce2f · outbound

This paper cites Siamban: Target-aware tracking with siamese box adaptive network,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Siamban: Target-aware tracking with siamese box adaptive network,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.689883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.470170Z digest=sha256:b54a233ca5160b6559452ea9b244f043c707e8280cf66a7600d1d55a5818a890

Observation f0f4888b-998b-4d60-ac2b-34faa6667c2a · outbound

This paper cites Reliable Object Tracking by Multimodal Hybrid Feature Extraction and Transformer-Based Fusion.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Reliable Object Tracking by Multimodal Hybrid Feature Extraction and Transformer-Based Fusion

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.557849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.473979Z digest=sha256:2817150010ccc27fadd613a2726d7f3ce21a6dc8d947c28e2239e2be86ec5067

Observation 63703f02-2ee1-466c-a648-9804ce6f755a · outbound

This paper cites Depthrefiner: Adapting rgb trackers to rgbd scenes via depth-fused refinement,.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Depthrefiner: Adapting rgb trackers to rgbd scenes via depth-fused refinement,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:01:38.677307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.478408Z digest=sha256:bc967e10357574affdff203746d9540ca10d061ab2444636ab3cef52532cfacc

Observation 1ab43ecb-3c4d-4d1e-b0cb-799a3e344e0e · outbound

This paper cites Cross-modulated Attention Transformer for RGBT Tracking.

Adaptive Perception for Unified Visual Multi-modal Object Tracking Cross-modulated Attention Transformer for RGBT Tracking

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:01:38.539120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:01:38.482277Z digest=sha256:00b8c57065e529c4b1f666ee066665524423ef9e2390fbf382a78fe0c1f4f58b

Observation 876085fb-394e-4348-8dda-c6647c7db7cc · outbound

This paper cites RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba.

Adaptive Perception for Unified Visual Multi-modal Object Tracking RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:38.486514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:38.486514Z digest=sha256:f1c0b2cccac3ffcdd949a75c7b21930af561c78123af065552eabca67c75e9ca

Pith citing papers

Observation 16c70ea3-faff-4ac1-8bd3-9e1cdca7e96a · inbound

Explicit Context Reasoning with Supervision for Visual Tracking cites this paper.

Explicit Context Reasoning with Supervision for Visual Tracking Adaptive Perception for Unified Visual Multi-modal Object Tracking

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:20:13.728895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:20:07.965249Z digest=sha256:0cdf316c07ac1e0ffd34be2700b9cb8b9c2189e19089f745729f2b40c6b140de