Pith. sign in

Paper Citation Record · LEDGER

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2506.23283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23283 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:59.930418Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact4
  • verified fuzzy43
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0566584-8b1a-4e89-bc30-b1d8a7a49a6e · outbound

This paper cites Vivit: A video vision transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.381794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.381794Z digest=sha256:e3768cd56d7b3f700858ff9d111aca2f1e9f7f99a9c87e1d6364be5cb2de91fd

Observation 18a8fd3e-b682-462a-913c-6f3ea3a4878d · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition BEiT: BERT Pre-Training of Image Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.473450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.473450Z digest=sha256:8bb7210a7e71e46f4d7154388e0abb39e6dfd6c258a3bead023eafe540279fcf

Observation 97822a62-148e-4a60-ad27-4c0290eedf20 · outbound

This paper cites Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.566560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.566560Z digest=sha256:c2d4881cf811211c4c31586439b8bbbe135b5709c97cacdb14a2547632bcb6d5

Observation 0dd5bef8-715c-41db-a180-87c6bea1eb99 · outbound

This paper cites Coyo-700m: Image-text pair dataset, 2022.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coyo-700m: Image-text pair dataset, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.684318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.684318Z digest=sha256:f4ac9a887cb5ba451f2284925fff9568fa58c7587c24f8f06c26a087dfb79dad

Observation 066d303e-31f0-495a-a9d5-e53f30ab3fcf · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Quo vadis, ac- tion recognition? a new model and the kinetics dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.823473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.823473Z digest=sha256:7bcf13d973ae54a4e59f7b7fdea723717923ba9bee60cde200f8acab996bcab4

Observation 6b1b048b-180e-4829-9f5b-7141258a6a96 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.963310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.963310Z digest=sha256:a085cdc81ec8e69e4dc93862899d92e10924c6040158befee7886c5013beb186

Observation 745bc955-cd20-424a-bb12-2e2c8673636e · outbound

This paper cites A simple framework for con- trastive learning of visual representations, 2020.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition A simple framework for con- trastive learning of visual representations, 2020

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.088191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.088191Z digest=sha256:213c3c4e3bcb24d7a6b95f7dd63a4177d00ccf398522a79a1d299afe707f78f5

Observation 35f85598-e66d-4960-b93e-a9b3d5b956d3 · outbound

This paper cites Feature-wise transformations.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Feature-wise transformations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.215965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.215965Z digest=sha256:c3f448b4e5ef356b413e3d84d64cc2dc92a2d5764aad972cc01774010e9c6b8f

Observation ae2ebe95-7d61-4550-8588-d8c48b4c95d4 · outbound

This paper cites Multiscale vision transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiscale vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.324371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.324371Z digest=sha256:bc27c593a3ee795529d8e6809d13fd4fcb887e2f9405f2961bfb4bd1cf0f99a7

Observation 56401364-3da8-488c-ad07-e74ededb08f6 · outbound

This paper cites X3d: Expanding architec- tures for efficient video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition X3d: Expanding architec- tures for efficient video recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.451686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.451686Z digest=sha256:fd9da1e48cfeb7c84d36330e41243a1cef074fdcdb4e71c3f257a6702ba22afd

Observation 2cfc5cce-ae53-43a3-9dfa-93799361f043 · outbound

This paper cites Masked Autoencoders As Spatiotemporal Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Masked Autoencoders As Spatiotemporal Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.583049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.583049Z digest=sha256:ed1b2ca33796771fd37990cf15ec03b12041de4b2ceaec9db3996dfa626344a0

Observation c259ede9-b86a-4879-9f91-ae3ec10321ff · outbound

This paper cites Slowfast networks for video recog- nition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Slowfast networks for video recog- nition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.697762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.697762Z digest=sha256:916bd743aa060ba81ec73f350d263c71837c6d4fd5278d8d5d2206ea7a87384b

Observation 5da590e6-82da-4603-8c2e-6b59e9ac8884 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The” something something” video database for learning and evaluating visual common sense

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.327642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:51.830505Z digest=sha256:3e168dfa7e67f9d1ad67ba181a014d470b2322986b56d4920d475799ff695ef8

Observation 68690598-06ab-4175-870e-f52779ee66e8 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.926035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.926035Z digest=sha256:c385a5dd24e3341dedc5fe2ca2e874c0bd56ad54956f989aef177a21e0e1ed9a

Observation 39b35f9b-124f-4e9d-90f8-3ea5f91027d2 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Efficiently Modeling Long Sequences with Structured State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.045734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.045734Z digest=sha256:44732c8c6c6988e16a1c19efcf74c9a677eb3828d4a8d1e3652a19bd5a0d2919

Observation 70f67fe2-2691-497b-a62c-b808bbe5a4b0 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary object detection via vision and language knowledge distillation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.279929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.127801Z digest=sha256:3d1c2464540dcc1d9250d1dda7f785ac13205df5b672a65824c0bd54d09fe02d

Observation d3fe1164-933b-4afb-abab-edf756748487 · outbound

This paper cites Trustworthy machine learning: From data to models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Trustworthy machine learning: From data to models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.195299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.256779Z digest=sha256:c6284b2e016e52edee49434cbece06fec1f5149f86f63fa2df8dbb5732ca6f35

Observation 245cfa19-8e54-4bed-a6b5-431633cdeba3 · outbound

This paper cites Turbo training with token dropout.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Turbo training with token dropout

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.084240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.390034Z digest=sha256:99913f64a206455d479963ec37d456b68014bebe70bfac8a517c34537639181b

Observation 75dbef3c-98ee-4f7e-9e9e-9b8c55438318 · outbound

This paper cites Learning spatio-temporal features with 3d residual networks for action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning spatio-temporal features with 3d residual networks for action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.978319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.549772Z digest=sha256:6b0764921b0c5cc3e640b8126745330408f71356b5be1699a0b9e32db70b21ba

Observation fd048668-ac8d-4cbe-87be-bcb5b743aee7 · outbound

This paper cites MambaVision: A Hybrid Mamba-Transformer Vision Backbone.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.633842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.633842Z digest=sha256:153fa737969520abd84eb7fc7bdfe67c4f92f1d5c71d3f62823630bfda4c2656

Observation a4995629-e6d5-4ab9-a680-a5d3da6ba861 · outbound

This paper cites Clipscore: A reference- free evaluation metric for image captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Clipscore: A reference- free evaluation metric for image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.875378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.745648Z digest=sha256:60d3963de8c360f5d871ba2ab1b8c90db0faa7e5d15e9ac5fc9b3a0721d49e7a

Observation afba095f-b5d7-4a4d-a56f-5b12f2cf32f8 · outbound

This paper cites Arbitrary style trans- fer in real-time with adaptive instance normalization.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Arbitrary style trans- fer in real-time with adaptive instance normalization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.750886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:52.831996Z digest=sha256:ac96e2518e0dd7ed899aaca26e12ef79368f29b00e232df3378d420c8774dd3e

Observation 250059a2-427d-4720-9901-4c9cb5b8384e · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.943841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.943841Z digest=sha256:0a58a132fda8230c1e73458f9ec1e5a53acd5af7970c9a39b2c633660029751d

Observation 75807e7a-5368-46d2-982f-815c6503e810 · outbound

This paper cites Smeulders.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Smeulders

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.626186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.022920Z digest=sha256:f5315e64f7c978ae7b3ff2fe01e1abcc5ed1f1aaf74f745f4e59a9004075bca9

Observation 5c017a9e-706f-40cf-af86-3ef5a43b02f3 · outbound

This paper cites Long movie clip classification with state-space video mod- els.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Long movie clip classification with state-space video mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.504852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.133192Z digest=sha256:b428f5dca14140c71f624ade894003a456f19b18ee2716e7fd3fe36cc2dbb610

Observation 034efcb6-bfc0-4177-9cbf-e9149da6fe12 · outbound

This paper cites Scaling up visual and vision- language representation learning with noisy text su- pervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling up visual and vision- language representation learning with noisy text su- pervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.396778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.286576Z digest=sha256:efa5f7e12f5194bab820b93e5a2e17c7f77377b273e0f1926398569f8d93ae14

Observation 06acb9ba-58e8-401a-b5ea-eddbfbff387e · outbound

This paper cites Laine, and Timo Aila.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Laine, and Timo Aila

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.285784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.420848Z digest=sha256:cfdde949cc167d57eb6ae7855df86dfafe60eaeada3002df523cf44c109991d3

Observation c231f89a-87ef-43fe-b2bd-7cd483c811f4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The Kinetics Human Action Video Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.493512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.493512Z digest=sha256:93063a956e5e2e88e3504207b6e2fd2cbdc95883009b3695377896eb4a753d91

Observation 4e4c4956-8e4e-498e-9b53-449f2cedceab · outbound

This paper cites Kuehne, H.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Kuehne, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.170200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.561748Z digest=sha256:9206520456c57fbff34dcca504ec6c80ce62e0b0a2467440d2aa13372931f7ca

Observation 8821f97d-59ee-4773-9bee-8fc6a466d9c6 · outbound

This paper cites The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.036537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.693704Z digest=sha256:7b5f211f993ef2d0b7c6508e3a8826c055280fe9ee72b043d55a152c907351d6

Observation 11c6d7db-7c34-4bcf-985a-a185bbf26bb1 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.789728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.789728Z digest=sha256:4794178708ffe2f6a819405f7437fa386c50b01a228a723fd377286380ede839

Observation 965818a5-d127-4beb-b212-05a651f0ad89 · outbound

This paper cites Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:01.162745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:53.897777Z digest=sha256:7112451711669c1535424710bfe5737a2d17c7c358b8c8f09140ac7ea609b2a3

Observation 6c983faf-b075-426a-beb7-3bd8197b733e · outbound

This paper cites Videomamba: State space model for efficient video understanding, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Videomamba: State space model for efficient video understanding, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.830280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.013868Z digest=sha256:9687935b68d1bfa4e5dd9f8492a8b8adeb9d9aab46686f9c2ad3ed9a53676d69

Observation 02fcb1cf-32a1-4da0-b55e-7c84cb0018dc · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unmasked teacher: Towards training-efficient video foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.653523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.127628Z digest=sha256:4598685fc8a52c95fb21b2f7d23ccf00d3e906e58732ee5e2f06f010bf78c96b

Observation 27f5fb1b-f812-4cce-97ce-0a366490a1c7 · outbound

This paper cites Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.465539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.238037Z digest=sha256:46072e864eda3c8fe9a7c0d91c3fd8daf129e78efebea8ccd32b1a114ac8cf74

Observation 26abb7e7-33e8-4293-b6cf-5b079c84c844 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.330276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.367511Z digest=sha256:a5e6670153c19019686021fcdbf0a55c196c1f89fd9a8ff86fee878fb4c0424a

Observation 1141e522-dbbe-4950-b689-f4c2bf95f6ed · outbound

This paper cites Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.141385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.431653Z digest=sha256:86b0074c7907c89bf323bf6eb31aa01349e709f739d3107ac651bcca0dc27957

Observation 83f7f27f-ffb3-4433-80db-6750c1eb9ce2 · outbound

This paper cites Pointmamba: A simple state space model for point cloud analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Pointmamba: A simple state space model for point cloud analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.960882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.524479Z digest=sha256:d334de6dde8d13c3c9af86cd4e5d304aeb6ed6652f7c2c53ed6615b2d60437a9

Observation 096c446c-94c6-459f-8729-50371225bc84 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Jamba: A Hybrid Transformer-Mamba Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.612897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.612897Z digest=sha256:b1c36086843ff515a6cf7aa2167618753eda39aff58b75bc2a2bcb41067f9df6

Observation e29d25ed-702c-46a4-ac21-893c111fc511 · outbound

This paper cites Learning to recognize procedural activities with dis- tant supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning to recognize procedural activities with dis- tant supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.822263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:54.696571Z digest=sha256:b03e5ec0a0dcc848924c0aad9c7a0322958300a8cf2f4b194000b8085ae13a63

Observation 82ff3b0f-f212-4e6f-88d6-b2b3490609bf · outbound

This paper cites Frozen CLIP Models are Efficient Video Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Frozen CLIP Models are Efficient Video Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.788646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.788646Z digest=sha256:9e107ed4611263999d328fbac97c84bddd8b8da110d21ef860e688ccf412e399

Observation 9898fcdc-ced0-4e8e-9241-e170e65e6353 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Annotation-free Audio-Visual Segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.889200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.889200Z digest=sha256:789849fea4ff1c2cdc7efc54b82022afc4b94e4f1956e234f14227278095acd1

Observation b070c1c1-bf0e-4edd-b688-d0623e3d3f36 · outbound

This paper cites VMamba: Visual State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VMamba: Visual State Space Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.973660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.973660Z digest=sha256:7e9aab5335d29b2999dc55793a8a92550806be9db8e22f52cbfbb8bdf76d42d6

Observation 3b47763b-25ec-4a8e-bbc0-c867acec771c · outbound

This paper cites Video swin transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video swin transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.670936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.081355Z digest=sha256:8b4f73decc851a62b2404e1d0c215b38b72b03cdb0522c7015da33cbd9c68973

Observation 6170fa3d-63f0-4201-b1f0-b538fecb1aa3 · outbound

This paper cites Freesegdiff: Annotation-free saliency segmentation with diffusion models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Freesegdiff: Annotation-free saliency segmentation with diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.524175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.165046Z digest=sha256:0f1120f3eabf7b89e57d7c74b46ce799efabfc4a6019e4e2b398d873acea00e4

Observation 03621015-443a-4b08-92ce-7b5cb1f70656 · outbound

This paper cites DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.235358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.235358Z digest=sha256:d9f90b589bf7bcacedeb5e2cb2dc830af3afd86ee8db35095346225a524a65ce

Observation e52b8b84-b3f9-4342-bf72-92861215bdd6 · outbound

This paper cites Open-vocabulary semantic segmenta- tion with frozen vision-language models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary semantic segmenta- tion with frozen vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.336866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.327089Z digest=sha256:00cbcd299349fbac12c36ecb65433c019dee44583d677500dc8e637d5d06ff43

Observation 69c49ab9-c49e-4727-87ca-5df1356d7e02 · outbound

This paper cites Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.120762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.420685Z digest=sha256:967ee3bd4fcd2b183102681c73c0859650ad6f0065e57e641708c7f24216a002

Observation 2573ff8e-3f54-4079-b207-f54ad18e4c47 · outbound

This paper cites Scaling open-vocabulary object detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:07.890235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.482406Z digest=sha256:1eb10914a3912f4b76ddd29f75ab83edaad649a1c1267a0b8e772947c1c0dcae

Observation 46177889-2581-457d-bd97-0e86f0a016ed · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:6b4e8a7e39c59d7d68c4590a482f56b08bf797b7c923ae9d59e9fb3099efc4ac

Observation 8bad67b6-8264-424e-98ae-4d4ea9723705 · outbound

This paper cites Expanding Language-Image Pretrained Models for General Video Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Expanding Language-Image Pretrained Models for General Video Recognition

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.859981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.663136Z digest=sha256:4c28b29aa7c647770431a1c379377124da0934d7cc96df85c102ada52045fa7a

Observation 00a9a893-7320-4bb5-884a-67d9c6c1167c · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.746646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.746646Z digest=sha256:b99a230abd09dbe0d498370776939d640e2f675d1a1cafe75e8c4c487893a4d4

Observation 1fceb07c-de6f-4b34-9b3b-72af3549bc11 · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:07.162702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.855563Z digest=sha256:320fa09dd69c4ed59e26afefabc4263c252925dc5b8bd334427f870fc78dfe7c

Observation 600a5af7-0e65-45fa-9ef4-dc11436ed1f6 · outbound

This paper cites ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.638728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:55.932723Z digest=sha256:fe2b3262b973e60efed10cbc54482959d1e64730aed3ebeb8f883400b15eba22

Observation 47ad84a8-cfe1-489d-8460-5ab7d35ae11f · outbound

This paper cites MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.425970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.065967Z digest=sha256:742fc35a0e5e37f9ac5a1c32db234cd5d21c173aa001e022db34811c7dc41e3d

Observation 01124fa0-a327-49ef-ac19-31cb7f5beac9 · outbound

This paper cites Dual-path adaptation from image to video transform- ers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dual-path adaptation from image to video transform- ers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.981489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.170231Z digest=sha256:f44b5e3ce9261cc820d948006764273b85e27ea4257fab32124c324fdb3eff20

Observation 56679b2e-1c87-4056-a665-fe6bc03ccdb6 · outbound

This paper cites Peebles and Saining Xie.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Peebles and Saining Xie

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.766762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.291434Z digest=sha256:bcf21263afe3543b71a1857041579851e41c6d91b26b9e9a9f995da413db52cd

Observation e3362da8-a59f-4667-bb1e-fe3b44df7dbd · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:06.501535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.418109Z digest=sha256:bb91f9abbc37ae848f30c0839c599fb96a48c415c3bf436b4c0fe0b7f81ed074

Observation 09d3aab8-921b-417a-ba0d-7e92b0719a3c · outbound

This paper cites Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.295700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.546195Z digest=sha256:c54e392780dbbc14d62ed7b582c032f46a65bd965443c549523c77c18b39807b

Observation 4c805022-7f59-4adb-be59-353adf757540 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning transferable visual models from natural lan- guage supervision

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.119792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.659418Z digest=sha256:377ab086ba92cd70a401796ae2c308b83b6e2e55234c0d32909ae1fcacb8a69b

Observation ead28b01-3e89-46be-8e2f-ef446ded4382 · outbound

This paper cites Token- learner: Adaptive space-time tokenization for videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Token- learner: Adaptive space-time tokenization for videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.975865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.759866Z digest=sha256:36aad76e4b2bcc8a236690ea71726dcaa6bdc390155ef77def5242667e4c2ebc

Observation 7c17898b-796b-46d0-afdd-8836a651f2b1 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:56.863163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:56.863163Z digest=sha256:3278798db61459824b3d225baa395f301d2163a5f57b34aba8f1f5641b666dbf

Observation 8b354a31-49bd-4357-9a09-10a6cd7e5e34 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Only time can tell: Discovering temporal data for temporal modeling

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.790063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:56.962660Z digest=sha256:df804c7ed13afdd36a1bc3de4c3658905d1769c1763511eb8de39cd1ee3259df

Observation 8c0e1250-f422-4fe9-9601-1a08ccdc984d · outbound

This paper cites Darf: Depth- aware generalizable neural radiance field.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Darf: Depth- aware generalizable neural radiance field

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.620347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:57.082523Z digest=sha256:48d687904059014054cdd7e7415553cc6ea6ce545dec552f185cd45cae5b5522

Observation 28917599-655e-4971-8114-058987c3fdcb · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.234563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.234563Z digest=sha256:ab1d41e1aae3ba5cd2b9cd5b67296efb09c1354b67a59c8e9fc782d380014e42

Observation 779064c5-4c2f-41b4-a898-307d77d7c2ca · outbound

This paper cites Coin: A large-scale dataset for comprehensive instruc- tional video analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coin: A large-scale dataset for comprehensive instruc- tional video analysis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.432865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:57.360492Z digest=sha256:1af3c6eef0cf7ff83f00ef475e796f2c662acbf77921b9d978bd0c9dbee696b1

Observation fbdc6a50-7eec-48f7-8c45-b34de633b22f · outbound

This paper cites Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.938654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:57.464332Z digest=sha256:baa7879661b8bbafd9b0f97cc5b22bdc0645de5719a9161c2747936cc6568441

Observation af449d3e-434b-4f99-ada6-63a5585c1387 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.583666Z digest=sha256:01ddf4bf56aa7153794d943bae96644d6474cf511b88127a025a0a88d6e7f6d8

Observation 0c40be8d-6286-4410-9506-79a9f25f2837 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition An Empirical Study of Mamba-based Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.712328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.712328Z digest=sha256:97b1b9a1590d64603ae8e99454642ee45d27be62babe7d6afc4e20098348e11a

Observation a3092b77-db40-4209-bc01-513dabb0a114 · outbound

This paper cites Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.852368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.852368Z digest=sha256:f45077fad097b6ddf3bbf4e251773aa69bc0e5c1e52a2db4a5b4bddbf96f01c2

Observation fefb174e-7ad2-4264-838b-62483d589634 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.921969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.921969Z digest=sha256:229a5a45fc223e09f33f5af8b9b5645efc8c5a4087183467f3fb24b662ef97f3

Observation 897aba94-bd40-4570-a160-c81813c23320 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.007202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.007202Z digest=sha256:45ac38d5785cfb6a09675516b28afaad4f304562a3687985ab30c15299d973b2

Observation 6cdd6186-d78a-4fba-8888-cf116c13b917 · outbound

This paper cites PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.110217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.110217Z digest=sha256:8fa6e7e2e6990116deea0196b2b49101cbb0af881b75b8c7eb5d5d9e2f00083f

Observation 94f0fbf7-8e97-43b8-a1e3-e50cb9b47a84 · outbound

This paper cites DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.193805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.193805Z digest=sha256:67b52d6f0b46a91393f8cdca446c90fd4d2e1d4d69aad953eff0fb19b5122766

Observation 7d526405-4e77-4079-a62e-68b17786e9ee · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.251948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.251948Z digest=sha256:7495a98ce838b45eb7cd4cc870f1647b104ae66c3247e768640c2997dbbd8945

Observation 90e734f1-29e5-4dc2-8f66-3a7369f14b76 · outbound

This paper cites Multiview transformers for video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiview transformers for video recognition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.304085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:58.324442Z digest=sha256:c39f2904b98ccdfb72fbcca4de5bb2c17d254cd2fba24bf2e93a5634659c62aa

Observation 68feddfa-c382-4fae-bb7d-49084f20ac09 · outbound

This paper cites MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.389943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.389943Z digest=sha256:2d57d2a657528af05e36234d23832dda6f825513c6f97c99c6ebc99365af00dd

Observation f2ed0136-122a-4a41-902c-67b430d24944 · outbound

This paper cites AIM: Adapting image mod- els for efficient video action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition AIM: Adapting image mod- els for efficient video action recognition

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.146644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:58.460734Z digest=sha256:08ec4654b6e49d926fae8d2839fbc7ec6f215b62cb82164a561b87c0e08c12e6

Observation 8c6e1585-036c-4a36-987c-68a3ef9dc655 · outbound

This paper cites Multi- modal prototypes for open-world semantic segmen- tation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multi- modal prototypes for open-world semantic segmen- tation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.959704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:58.591097Z digest=sha256:72c677f216071a39613fba9ce4fd6b0671b5c38016bb91db197f7b51216e0bf2

Observation f07b56a4-579a-47ed-ad1b-744a7fafac25 · outbound

This paper cites Remamber: Referring image segmentation with mamba twister.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Remamber: Referring image segmentation with mamba twister

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.772877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:58.763767Z digest=sha256:eb9973485e9ae44e38c506720847653c802acabb1256bf18054d7cf232a352b2

Observation c229ec09-62d6-4d04-b8e3-a67848f1fee9 · outbound

This paper cites Learning with multi- class auc: Theory and algorithms.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning with multi- class auc: Theory and algorithms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.585717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:58.912661Z digest=sha256:65d398211434ec7da3f1f1279be78abd3a303155c814e744ff440ad5b307d508

Observation 9a409808-e05f-4f5c-a661-b69f11de62a6 · outbound

This paper cites Optimizing two-way partial auc with an end-to-end framework.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Optimizing two-way partial auc with an end-to-end framework

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.417014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:59.152808Z digest=sha256:ef30e6e7ea2a5b73f66f75643d6f5d6bbe36993fbdbd5bc3a20246c334317ec0

Observation 7cbfadaf-943e-4673-8f37-e4d2aad83fb1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigmoid Loss for Language Image Pre-Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.342925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.342925Z digest=sha256:bdb9f0f9308366899b12eb079adcb056687051f700dee34f17c63a702b242608

Observation c448c7b1-0717-4d7e-baff-928868910d1e · outbound

This paper cites Point Cloud Mamba: Point Cloud Learning via State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Point Cloud Mamba: Point Cloud Learning via State Space Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.480981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.480981Z digest=sha256:a315abcd0a8705d26c38cbbf749669a81ff5da6ff694d0a3e5cc4a96fbd61bad

Observation 613a4a46-2840-47e1-a6a4-eb0402e8d530 · outbound

This paper cites G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.564079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.564079Z digest=sha256:16006d35a23b8d05b0530eee810bee614c17f2cfa41e1f07194ae61a0d129f6d

Observation c101d2c5-237a-4331-8e65-f43b9565ab7e · outbound

This paper cites Vidtr: Video transformer without convo- lutions.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vidtr: Video transformer without convo- lutions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.118114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:59.662077Z digest=sha256:1e2f8edd540347a8ce9ae740653c0f77e2ed75ec73fe7c8a5940317d8d3326b2

Observation 8014a4aa-3a2c-4def-8347-74f7b1beac90 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Regionclip: Region-based language-image pretraining

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.837632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:59.733742Z digest=sha256:d17d00c15bf7f5caf88e0006d9ceaa90e8d89870580a83e72a31fd579e1547bd

Observation 6deeac2c-3a55-4d5d-8cb7-97eb4bc75e57 · outbound

This paper cites Graph-based high-order relation modeling for long-term action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Graph-based high-order relation modeling for long-term action recognition

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.571641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:51:59.857627Z digest=sha256:e78a54f7a569b7f1a00b28e5b3ffded51bc8170695900673327b7c594ee2c080

Observation ce255339-f313-4b06-8e78-b21533f4e672 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.930418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.930418Z digest=sha256:1aa9c79fd8be41948343580e785e758452f52bcfbdaf63e4d57837099d5b583c

Pith citing papers

No inbound Pith citation observations are available.