Pith. sign in

Paper Citation Record · LEDGER

VideoMamba: State Space Model for Efficient Video Understanding

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.06977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.06977 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:20:01.630113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.123917Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc3e6eb1-35bb-48ac-b4bf-c0883b8af3c4 · inbound

A Survey of Mamba cites this paper.

A Survey of Mamba VideoMamba: State Space Model for Efficient Video Understanding

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:13:30.939914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T22:09:19.917854Z digest=sha256:224ae96a8de067b4e80b311ca3967e99dcd6a681a6fe74903dad84d898cd27c9

Observation 5adcd7d7-0fd3-4c44-b089-585491d537f2 · inbound

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking cites this paper.

MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking VideoMamba: State Space Model for Efficient Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:01.630113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:01.630113Z digest=sha256:17d2a7e6c8e1ea5f014a87a3be27f3f192497a5dda96b3faebcb4c07da197182

Observation 91f95b12-c9b6-4d7f-9e96-d60d91046bf7 · inbound

Deformable Mamba for Wide Field of View Segmentation cites this paper.

Deformable Mamba for Wide Field of View Segmentation VideoMamba: State Space Model for Efficient Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:06:35.337073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:06:35.337073Z digest=sha256:bbebc8acc4fa7bc1457c431cc23e83a4a2f338126f1e227bc4e3dea089c91c21

Observation ee44ba62-509e-41d2-b805-86f5ea6d8118 · inbound

MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning cites this paper.

MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning VideoMamba: State Space Model for Efficient Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:58.560536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:58.560536Z digest=sha256:8c4067cbb234bc3f22c83c11b174cec00bf8b1231d4812663eeaebefd31b2ed3

Observation 27182142-0746-4282-b62c-f014952ecb1d · inbound

MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection cites this paper.

MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection VideoMamba: State Space Model for Efficient Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:11.380629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:27:11.380629Z digest=sha256:af4ef2198b0893d72bf39b23a499cec48e816de880b89367f2030327424fe906

Observation 25ac5823-7017-48d0-a5aa-8f0bdc729f79 · inbound

MambaLCT: Boosting Tracking via Long-term Context State Space Model cites this paper.

MambaLCT: Boosting Tracking via Long-term Context State Space Model VideoMamba: State Space Model for Efficient Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:00:30.625789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:00:30.625789Z digest=sha256:26d9bb066c11c8a512617e970a5b69ed81da22ab4e6f9791ba81dc5b508adc15

Observation 7d18cdbb-f0af-447a-9a99-39726de1a73f · inbound

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking cites this paper.

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking VideoMamba: State Space Model for Efficient Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:12:52.333311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:12:52.333311Z digest=sha256:d1e40b580f094b9cf8faa6411e98e3600b71f7d7271b6e9ff2845979a39f4b7d

Observation f8ce410b-9b52-42f1-bfad-48c2aae2d9c7 · inbound

V"Mean"ba: Visual State Space Models only need 1 hidden dimension cites this paper.

V"Mean"ba: Visual State Space Models only need 1 hidden dimension VideoMamba: State Space Model for Efficient Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:37.003688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:37.003688Z digest=sha256:ee54fcfc0fb5ab054c1d9fc3bc9a9fbf28cbb0958bc1d0883d04268e0c285517

Observation c574a8b3-4361-48b2-8816-b02ceddbf2c2 · inbound

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing cites this paper.

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing VideoMamba: State Space Model for Efficient Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:07.883499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:07.883499Z digest=sha256:03a88b47fbfd78dd1db81ebe2b31a8376d63fb411b04a0e66d64203df36f53f1

Observation 9521212c-fc34-4fb2-979e-440ae27c7e8f · inbound

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving cites this paper.

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving VideoMamba: State Space Model for Efficient Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:11.569111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:11.569111Z digest=sha256:0226d46914b4dc80f84365c04021d10a619187da740e5bfbdf34e439bc3cc290

Observation 2c0f48e6-09af-4641-b416-3321b1f3bd60 · inbound

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation cites this paper.

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation VideoMamba: State Space Model for Efficient Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:38:56.483125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:38:56.483125Z digest=sha256:566b8e7b51c4484494cceb39b2c23b8839914847645c97ead33dfc44a4a55219

Observation de4aa581-649f-4391-b74f-f759561fe1ec · inbound

MV-GMN: State Space Model for Multi-View Action Recognition cites this paper.

MV-GMN: State Space Model for Multi-View Action Recognition VideoMamba: State Space Model for Efficient Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:30.555080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:37:30.555080Z digest=sha256:1cb45307ecb33f8d632cf79a4a053a387856a0dea1991799cd33094f94c4900d

Observation 3faf318b-c026-4671-8a37-75a93ea98b8e · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks VideoMamba: State Space Model for Efficient Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.758010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.758010Z digest=sha256:2457be8391c40123be5e8104a8b506ffed67c8abdd7bbd95a0e5e54013534079

Observation d2465311-0db7-4e7e-8740-81f8cbadd9eb · inbound

Mamba Drafters for Speculative Decoding cites this paper.

Mamba Drafters for Speculative Decoding VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.919222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:12.919222Z digest=sha256:bef3961027021a663631f68965c42fde3a3120c379dcfed7c00b36949cced3c2

Observation 03a5c6c0-ad48-4a67-be9d-51d07e94ec7f · inbound

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos cites this paper.

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos VideoMamba: State Space Model for Efficient Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:23.682312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:23.682312Z digest=sha256:d6b003e05aab206d3170a2eda673b023c947e90ba83e77e125e330c80b75bf12

Observation 60fd24f9-7f8e-4d8d-bade-a4d8a55cc9a8 · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization VideoMamba: State Space Model for Efficient Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.396361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.396361Z digest=sha256:7a45cc08d11bb852f6af9df9f0f86d2b70a5d41cc3995598c45baf2c7c6b6960

Observation 421ad7d2-17c9-475e-b20f-1894381d7cff · inbound

QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models cites this paper.

QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models VideoMamba: State Space Model for Efficient Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:28.795259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:59:28.795259Z digest=sha256:b576587755e2348c630e055dccf1cab5ca1f8efe90663fbc05065526ea64201d

Observation ad9977a5-0e2f-4aa0-909f-dee0562fc13e · inbound

Few-Shot Object Detection via Spatial-Channel State Space Model cites this paper.

Few-Shot Object Detection via Spatial-Channel State Space Model VideoMamba: State Space Model for Efficient Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:16.242741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:39:16.242741Z digest=sha256:cd743bce8c26c8c372063d9afb6a11a5455a52c61052ffa7caf34a67b68af699

Observation cf99f8eb-e72c-4ddc-9c7e-33afe00406fe · inbound

HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning cites this paper.

HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning VideoMamba: State Space Model for Efficient Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:41.966729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:41.966729Z digest=sha256:e0a9051891bbf1dec2101134b729c818a0fb3787c0ca0a49cb91d82407e31392

Observation 8897b41a-c3b7-4c25-9cc7-e293b2f6b83a · inbound

Straightforward Bayesian A/B testing with Dirichlet posteriors cites this paper.

Straightforward Bayesian A/B testing with Dirichlet posteriors VideoMamba: State Space Model for Efficient Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T21:43:17.884122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:43:17.884122Z digest=sha256:cfbc08f5bc09fa2f047a884e4f990569e9606c0a9d34f37c8c944f1fd092a715

Observation c77fc7dd-61f4-410d-aa78-68d13ba08b1e · inbound

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds cites this paper.

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds VideoMamba: State Space Model for Efficient Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T21:09:10.359648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:09:10.359648Z digest=sha256:8fafea4b37ed0a17a44075533a36ca4340d491cba8d282a35a3fbece4024c2ae

Observation 1a643685-9f82-42b7-a214-1ecb92e22736 · inbound

Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos cites this paper.

Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos VideoMamba: State Space Model for Efficient Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:44.196052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:44.196052Z digest=sha256:0454454aa75a89f503045bbaf08d18372f004843ae0f9c8adabe1eea01d6ca8d

Observation cbb424ea-83e3-4e0a-95eb-569bb96cbb12 · inbound

Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression cites this paper.

Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:33.351148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:13:33.351148Z digest=sha256:967ae7bdffedfd088d8a7927a66fc66c30d8ce0803b1c9a3d59a836727c1e19c

Observation 8b02b6aa-54c2-4cd9-9c75-b4f3705b80aa · inbound

Time-Scaling State-Space Models for Dense Video Captioning cites this paper.

Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.179915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.179915Z digest=sha256:5c04d51b16345100202fbe274d9ed27ec906b0072d554d43b2aba4e8e2bbfe93

Observation 36de079a-1ff4-4fe6-a456-9424b8091984 · inbound

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection cites this paper.

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection VideoMamba: State Space Model for Efficient Video Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.636582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T09:26:51.456652Z digest=sha256:24fb6d119c1bcec95d64770ca7490d68f526597bf3d8bef270cc248fad99de38

Observation f26f67de-cf65-4323-9ad8-1fd3fdad76a3 · inbound

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models cites this paper.

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models VideoMamba: State Space Model for Efficient Video Understanding

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.126013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:20:21.295192Z digest=sha256:ebd051a8650a0fa6caabf518b0311ec12dcd3fe29e2a50d5307a5b4bf779716f

Observation eec9768f-e5f3-4c09-90a0-481b673ef511 · inbound

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing cites this paper.

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing VideoMamba: State Space Model for Efficient Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T09:28:19.511968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:28:19.511968Z digest=sha256:1393d4bad23abd32f8bedf88eff264e98c89607e78118967b6cca58be0a95632

Observation 4869c109-36eb-443f-a9e6-7d320f6d3e54 · inbound

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding cites this paper.

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding VideoMamba: State Space Model for Efficient Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:52.300928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:52.300928Z digest=sha256:7fecc977d7e63833f2e0998414765965b5a42e6aba18a7b0114e3d228a6fa042