Pith. sign in

Paper Citation Record · LEDGER

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2411.15459.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15459 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:20:01.720631Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a6581c63-6da7-4362-99d1-cb7d782d3f99 · outbound

This paper cites Longformer: The Long-Document Transformer.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Longformer: The Long-Document Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.544122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.544122Z digest=sha256:5e0eed56dea018beaf4dc8b9b0db47a66459ff4cc35e1e9ea9042e6b96dd2fde

Observation 9dd2fb2c-8096-4ff5-8b7e-265da37e7d88 · outbound

This paper cites Fully-convolutional siamese networks for object tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Fully-convolutional siamese networks for object tracking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.181808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.548262Z digest=sha256:430c62d8190264ab7c685ed339405b3f7410797c3b1b07d22c5176dc604849c7

Observation f30c6986-6c06-491d-bba9-9abdf184baf0 · outbound

This paper cites Learning discriminative model prediction for track- ing.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Learning discriminative model prediction for track- ing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.173259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.551394Z digest=sha256:4d594f599f2ec4981d36870dbb3c73bc78b41e229f1099dfefa4dd0687357a1a

Observation f49a0f57-6910-4181-98a0-74b4532cca80 · outbound

This paper cites Transformer tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Transformer tracking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.162868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.554839Z digest=sha256:fb82bb0b450ff495d52830c41a49e87b3044d175d649141f7fa033b77e777b6b

Observation 4fa9e11a-d20c-4404-9ed7-5cb8ec8c3479 · outbound

This paper cites Mixformer: End-to-end tracking with iterative mixed atten- tion.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Mixformer: End-to-end tracking with iterative mixed atten- tion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.558719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.558719Z digest=sha256:e69be84e1f05717047c936ff1b22af79e180ca4bdb523bc8cea91871a1c3b068

Observation 8bf97ae4-53ba-42ea-9ff9-6ee8b80dcc12 · outbound

This paper cites Eco: Efficient convolution operators for tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Eco: Efficient convolution operators for tracking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.150042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.562494Z digest=sha256:563d3c8135a2c6f2917b716e17c380a32fce0de798e1b2c7e2e04ca6aa372b73

Observation 8181d726-7c28-4ac5-ba72-9e7dc704893d · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.565544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.565544Z digest=sha256:0b2799e29418936976813078efe91895f9d43ff4d780d8b29873ecec3261c68e

Observation ae09d6b5-30dd-4372-b1ab-ebb5e54f4606 · outbound

This paper cites Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.568536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.568536Z digest=sha256:8baea222a8214d4e5cdef2e48ee42dcbe465aa8832b73ad558a8d8e52ae50641

Observation 57470215-477e-40d7-8267-55067c18dacb · outbound

This paper cites Real-time visual object tracking with natural lan- guage description.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Real-time visual object tracking with natural lan- guage description

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.134905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.572867Z digest=sha256:a43d4fa4123a32179359c91203c951fbe7be83cbe1509d36a8d93c0e339e9ec2

Observation 1d6ea254-fd70-4f8e-a343-e45693357589 · outbound

This paper cites Siamese natural language tracker: Tracking by natural lan- guage descriptions with siamese trackers.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Siamese natural language tracker: Tracking by natural lan- guage descriptions with siamese trackers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.125997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.576484Z digest=sha256:890e369359eef48598f0c94878663bdce7dba24c6434d5c31f29a864b8a314bd

Observation c14e1255-552f-4572-9c2e-41e5d40f1dc2 · outbound

This paper cites Hungry hungry hippos: To- wards language modeling with state space models.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Hungry hungry hippos: To- wards language modeling with state space models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.117325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.580176Z digest=sha256:f2823d5f949e7590ce68263db9db9bf3adc98e963495bd2c385261f0114ccbaf

Observation 52ba6913-936f-46bb-a3c9-ababb534a261 · outbound

This paper cites Stmtrack: Template-free visual tracking with space-time memory networks.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Stmtrack: Template-free visual tracking with space-time memory networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.108752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.583728Z digest=sha256:a02482df8ec449a0735647eb05bef8ad038e25a4f2368c47fc0dcccaad14bd96

Observation d65da1bc-3c72-46a5-91ba-83adef6e77eb · outbound

This paper cites Generalized relation modeling for transformer tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Generalized relation modeling for transformer tracking

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.100081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.586708Z digest=sha256:0bb0aa656afc3fec1eed19027ff3c84a164839302d7872df5d1f0920123c7322

Observation 6c3e1075-1a27-40e2-98df-385da687db40 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.589615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.589615Z digest=sha256:76d9b8e535aaae08656347c5c699272f4045260bcca82d07e0aa766841099519

Observation b2a870ee-c7f5-4644-89be-513fbf5718b2 · outbound

This paper cites Efficiently mod- eling long sequences with structured state spaces.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Efficiently mod- eling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.092058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.592732Z digest=sha256:19d21c8e47753492d8a83426b03c3226d1300996c76afe48cfb17b187a84ade4

Observation 0609f4d3-4151-49de-90ff-1cbec8a71231 · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.595970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.595970Z digest=sha256:4fbbd11eeda3e07d1f5ac7eea3b5208521fe5afc53541cc1676250ca8fc9e765

Observation af42b50c-e536-4b5c-87b0-7f2cc8fa3c63 · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking On the parameterization and initialization of diagonal state space models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.598955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.598955Z digest=sha256:6b72d3c2e4f3e78fe3b241ec8682ca9c4079a040cf22de9961a114d11f9f6d11

Observation fdfa493b-397c-484d-987b-d830f387c830 · outbound

This paper cites Mambair: A simple baseline for image restoration with state-space model.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Mambair: A simple baseline for image restoration with state-space model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.601871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.601871Z digest=sha256:9307569cd619a3a2febd2f74d1b5c888007f203b1e64433acfe7dbee0584c6ff

Observation 0d24d2ca-49b5-4a07-b55b-8a9bd4a5197d · outbound

This paper cites Divert more attention to vision-language tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Divert more attention to vision-language tracking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.068064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.604784Z digest=sha256:08f366396d654e5c4f70a713317e63394a90080fbc8faff2040588172acb8997

Observation 300d66d5-3d69-4d8d-b254-b7218fe9d833 · outbound

This paper cites Demystify Mamba in Vision: A Linear Attention Perspective.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Demystify Mamba in Vision: A Linear Attention Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.608104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.608104Z digest=sha256:319641f5c0e983b48146e1ac7427dbcf1384e314b6d3516c4868f49b72a49178

Observation 61d363e2-8650-4f9c-95e8-a155e8b296c8 · outbound

This paper cites MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.611489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.611489Z digest=sha256:2ec9f168b4d8f2228d23e27ac8edba7162336f5e6d19ec2c949eafce7b94f0a4

Observation 028492ef-d463-4afc-ba58-9a768c3584ca · outbound

This paper cites High-speed tracking with kernelized correlation fil- ters.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking High-speed tracking with kernelized correlation fil- ters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.059423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.614864Z digest=sha256:f677a93a12acb92fc1b6962b2c75129357c82cf72487ae6ee2806e74384f71b1

Observation a9d49dbb-20d5-4d17-a7dd-1e83c68e2e12 · outbound

This paper cites A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking A multi-modal global instance tracking benchmark (mgit): Better locating target in complex spatio-temporal and causal relationship

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.050178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.617841Z digest=sha256:91dd28752d48fb93938081657465d0d8f49e9818ce4818d9cb331ef6649ebb08

Observation e66ed6a9-f12e-40a1-a5f3-72b75b7224a4 · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.040139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.620570Z digest=sha256:98c0e54411d9623e1f27b4c20e9f3542d1dde3ba20e54f97c8f8eeb5c761c5cc

Observation abcc008d-4975-4f85-9c1e-529f71ce5104 · outbound

This paper cites High performance visual tracking with siamese region pro- posal network.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking High performance visual tracking with siamese region pro- posal network

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.623287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.623287Z digest=sha256:67be3ce1e0d56ed8079c24c0f55d69d689fe1fa7f65de57d54a5c84bae09b3ad

Observation 5c9d99d7-0374-4249-81a0-f8ac8f4ecec4 · outbound

This paper cites Siamrpn++: Evolution of siamese vi- sual tracking with very deep networks.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Siamrpn++: Evolution of siamese vi- sual tracking with very deep networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.025986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.627029Z digest=sha256:043f0306a7a6c6be06be438b04345eb3dd91fcb990b87e06fc84a2138bad3fe2

Observation 5adcd7d7-0fd3-4c44-b089-585491d537f2 · outbound

This paper cites VideoMamba: State Space Model for Efficient Video Understanding.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking VideoMamba: State Space Model for Efficient Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.630113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.630113Z digest=sha256:aad9ddcddcbc61a923f629be32c0eb7897e693ff1db4a75936eab9479a85b7cc

Observation 6099526f-d93d-4be3-9c6d-dfc03239752e · outbound

This paper cites Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.633620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.633620Z digest=sha256:c1a15b8827009f9aa98146e6ac2d06eee98653e34286365aa16e8d67729dcecc

Observation 728fedc6-303c-4b87-8af9-30e0d4c51114 · outbound

This paper cites Cross- modal target retrieval for tracking by natural language.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Cross- modal target retrieval for tracking by natural language

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.015505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.636982Z digest=sha256:5874f39d699e6edbdfc364c2348f9d74c3649a540735395922841da732a98941

Observation 98df1291-0654-4c4d-8b92-b1f54915966c · outbound

This paper cites Tracking by natural language specification.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Tracking by natural language specification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.639871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.639871Z digest=sha256:04578331468e732dfd2db5c0fdc3a86a571ca53afd146fae4aa0df23ae6445ee

Observation 74d970e1-2326-4893-9a36-9d851db40bcc · outbound

This paper cites MTMamba: Enhancing multi-task dense scene understanding by mamba-based de- coders.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking MTMamba: Enhancing multi-task dense scene understanding by mamba-based de- coders

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:02.001461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.643134Z digest=sha256:c81fc1a425a613f5a0661d40a052cf30741fc713433ffa04858725f9299a7902

Observation 48b519f6-27a5-4f43-9076-c2abea571e20 · outbound

This paper cites Swintrack: A simple and strong baseline for trans- former tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Swintrack: A simple and strong baseline for trans- former tracking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.990328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.645985Z digest=sha256:0ac28a87510f11e5a71160e1dbd54fda5c2f83fcfa3bf31d3258a771ae16fd11

Observation 355e5ee0-c2b1-4e87-bbc8-c4bc297017de · outbound

This paper cites VMamba: Visual State Space Model.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking VMamba: Visual State Space Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.649090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.649090Z digest=sha256:66d7c7824e30c95dbd18041c331fc99b5e08aa33ccaa3116d75215efccf6855d

Observation e81b9aad-2ddc-401f-8b03-9df689d56cad · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.652546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.652546Z digest=sha256:980df64b29660605e3f1d0b14cb0ad4aaae192f6a000446802973c9ae9461e41

Observation deeb7364-25fe-410e-acb3-bb29b85d01db · outbound

This paper cites Unifying visual and vision-language tracking via contrastive learning.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Unifying visual and vision-language tracking via contrastive learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.974955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.655452Z digest=sha256:f0e4caa789e01e37a725c897356bb0e76a7e3bcd3e01e39d20e99c74d3fa766d

Observation f99182f9-a5e9-4931-ac73-4f59d7fab7da · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Generation and comprehension of unambiguous object descriptions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.658639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.658639Z digest=sha256:39c98e169a69ff9c57f8521ca3699ac6b0cdeb8459a70039f11c4053bef77aaf

Observation dd85b160-767a-490e-8d47-3020cd2b1f7e · outbound

This paper cites Long Range Language Modeling via Gated State Spaces.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Long Range Language Modeling via Gated State Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.662622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.662622Z digest=sha256:b97a3c389835f7c3bdb88b28417cb52d7f137e8a0a32f6259b225220abdb5093

Observation 4df67e8a-c9ff-485c-84fa-4dbe85e2776f · outbound

This paper cites Learning multi-domain convolutional neural networks for visual tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Learning multi-domain convolutional neural networks for visual tracking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.960683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.666664Z digest=sha256:508bec4c17fce795c0babb420cfd4b258687aef1f6149143a406e48d6adf4124

Observation fd97c461-71e7-4c1e-9049-546836d36abb · outbound

This paper cites Context-aware integration of lan- guage and visual references for natural language tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Context-aware integration of lan- guage and visual references for natural language tracking

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.951243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.669880Z digest=sha256:020c70496674b0f703928cda2e18792950aeaa3892d190ecaeb9bef21779b2e8

Observation c7707fe5-cfa6-46e8-8e09-f698e646a1e9 · outbound

This paper cites Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.673803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.673803Z digest=sha256:67ff6b52652fb13dd616cfcac4930e7973c0d0434bc8df54507fd0ed06584376

Observation bb2cd656-518e-4d52-a50a-1f91343ffbc7 · outbound

This paper cites Simplified state space layers for sequence model- ing.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Simplified state space layers for sequence model- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.943019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.677927Z digest=sha256:0a99b93a8e938cd636690eac3a136f9a3abbd796dec7a836b970ea5feb62f58a

Observation 8163d3bf-69a2-445d-b2aa-6a7fca3820d6 · outbound

This paper cites Fast template matching and update for video object tracking and segmentation.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Fast template matching and update for video object tracking and segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.934590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.681010Z digest=sha256:5e0bf25cc7b68467c5cdd9c9cd9b039318e0f0eb2f950940510f4093341b36df

Observation 9ab72c60-ba61-4856-b2c4-d80099f52fe2 · outbound

This paper cites Transformer meets tracker: Exploiting temporal context for robust visual tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Transformer meets tracker: Exploiting temporal context for robust visual tracking

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.925530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.684047Z digest=sha256:09353a5de6315a30439b4242baae700fd3109f08f830d6d2adf996027dce00b0

Observation fae2b7b6-ed6c-4e61-8012-98bcb3e3d0c4 · outbound

This paper cites Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Towards more flexible and accurate object tracking with natural language: Algo- rithms and benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.917419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.687386Z digest=sha256:e92675f03a9c193418d4976f0b9440612fe6f44be3b03f6cc06f68d98bd63d84

Observation 4bb6b9a7-8050-43c3-ae9a-b1f873567280 · outbound

This paper cites MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State Space.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State Space

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.690659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.690659Z digest=sha256:b6ffc920eb3e52b68b074aa48dc8a6464792874cb225c35c6f0dd147a0e3fc37

Observation 9f4d9c95-edb4-4d2e-a23e-7baea397b7cd · outbound

This paper cites Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.908175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.694007Z digest=sha256:44360a4ae4102ab89f84011263aad01ae177a61fefe05436daeca09cee0003bd

Observation dc204745-316b-43b0-a1c7-3439150e22c1 · outbound

This paper cites Learning dynamic mem- ory networks for object tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Learning dynamic mem- ory networks for object tracking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.899060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.696798Z digest=sha256:47ae5428a3ecf75337eb099381bc865b665031baf026e7ea805c9a74aa63b212

Observation 40b3d93b-bb3d-4665-8ac0-c45c2f0b7d2c · outbound

This paper cites Grounding-tracking-integration.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Grounding-tracking-integration

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.890476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.699953Z digest=sha256:b1510a9ff1c604d26c80d0f3e5e6c467a08b3537531435332c11b8727af61a06

Observation 4db2eb79-5d11-4566-9d97-bfd2b92375f5 · outbound

This paper cites Joint feature learning and relation modeling for tracking: A one-stream framework.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Joint feature learning and relation modeling for tracking: A one-stream framework

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.881755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.703010Z digest=sha256:09ae0e33b8322687cc858909549ad5b5fac8f9c11980a4ce5e48b5cdf346baff

Observation 7c4f2322-1be1-4234-8f7c-a0e7b4b8c8c2 · outbound

This paper cites Object track- ing: A survey.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Object track- ing: A survey

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.872687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.705945Z digest=sha256:3564e98154178d059c7e7657f41b8663bef50bb7f4a005857672877708cd0120

Observation f97e2aee-3c2f-4b43-b45f-d1a3edcd877b · outbound

This paper cites VFIMamba: Video Frame Interpolation with State Space Models.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking VFIMamba: Video Frame Interpolation with State Space Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.709076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.709076Z digest=sha256:a23f8ce44f3418748f90a0005a1810fe1421f0923161fbac17bea88b4be30049

Observation d9af8241-1c4c-4c1b-be10-4e786bb577cd · outbound

This paper cites Learn to match: Automatic matching network design 10 for visual tracking.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Learn to match: Automatic matching network design 10 for visual tracking

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.863568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.712704Z digest=sha256:3d26cb1e4f8ba5ce9f1e46c82149f6c9c2eba5be414037e918a4de008bce7c25

Observation f064469c-6e1b-4b67-9fad-8a97627fa7df · outbound

This paper cites Joint visual grounding and tracking with natural language specifi- cation.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Joint visual grounding and tracking with natural language specifi- cation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.854384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.716871Z digest=sha256:7fd4e01fc79a1e36be430d600a6942269fe7483632dc2b0dfadf2d32ce772883

Observation 4f811161-bb32-43ac-906e-2f40a87e6f66 · outbound

This paper cites Vision mamba: Efficient visual representation learning with bidirectional state space model.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking Vision mamba: Efficient visual representation learning with bidirectional state space model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:20:01.845324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:20:01.720631Z digest=sha256:e144f6dde8b66e55056f82b768fd72ec9a32ae71c41e55ee687dfc6128da14f5

Pith citing papers

No inbound Pith citation observations are available.